20260606152506

This commit is contained in:
oneao committed 2026-06-06 15:25:06 +08:00
1 parent 8dc2775cf1
commit 9947193b6d
8 files changed
+71 -53

No files matched your search

@@ -1,18 +1,21 @@
# PaddlePaddle(Python)
# 集装箱号识别(PaddleOCR)
[飞桨(PaddlePaddle)](https://www.paddlepaddle.org.cn/)以百度多年的深度学习技术研究和业务应用为基础,集深度学习核心训练和推理框架、基础模型库、端到端开发套件、丰富的工具组件于一体,是中国首个自主研发、功能丰富、开源开放的产业级深度学习平台。 飞桨于2016 年正式开源,是主流深度学习框架中一款完全国产化的产品。
## PaddlePaddle
### 1 安装
参考网站:https://www.paddlepaddle.org.cn/install/quick?docurl=/documentation/docs/zh/develop/install/pip/linux-pip.html
安装 GPU 版本:
```bash
python -m pip install paddlepaddle-gpu==3.3.0 -i https://www.paddlepaddle.org.cn/packages/stable/cu118/
```
安装 CPU 版本
```bash
python -m pip install paddlepaddle==3.3.0 -i https://www.paddlepaddle.org.cn/packages/stable/cu118/
```
@@ -20,7 +23,7 @@ python -m pip install paddlepaddle==3.3.0 -i https://www.paddlepaddle.org.cn/pac
如果是 **GPU版本** 需要注意自己的 `CUDA` 版本,不要过高即可,查看命令:
```bash
nvidia-smi
nvidia-smi
```
输出如下:
@@ -37,7 +40,7 @@ nvidia-smi
| N/A 27C P8 8W / 70W | 4MiB / 15360MiB | 0% Default |
| | | N/A |
+-------------------------------+----------------------+----------------------+
+-----------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
@@ -65,7 +68,7 @@ python -c "import paddle; paddle.utils.run_check()"
(ocr) root@VM-0-80-ubuntu:/workspace# python -c "import paddle; paddle.utils.run_check()"
/root/miniforge3/envs/ocr/lib/python3.11/site-packages/paddle/utils/cpp_extension/extension_utils.py:715: UserWarning: No ccache found. Please be aware that recompiling all source files may be required. You can download and install ccache from: https://github.com/ccache/ccache/blob/master/doc/INSTALL.md
warnings.warn(warning_message)
Running verify PaddlePaddle program ...
Running verify PaddlePaddle program ...
I0710 06:19:32.810492 2651 pir_interpreter.cc:1524] New Executor is Running ...
W0710 06:19:32.813128 2651 gpu_resources.cc:114] Please NOTE: device: 0, GPU Compute Capability: 7.5, Driver API Version: 12.0, Runtime API Version: 11.7
I0710 06:19:35.382279 2651 pir_interpreter.cc:1547] pir interpreter is running by multi-thread mode ...
@@ -78,6 +81,7 @@ PaddlePaddle is installed successfully! Let's start deep learning with PaddlePad
> 项目背景:识别集装箱号
### 1 环境配置
当前使用的是虚拟环境:
- python:3.11
@@ -95,7 +99,7 @@ PaddlePaddle is installed successfully! Let's start deep learning with PaddlePad
下载后解压
```bash
unzip PaddleOCR-3.1.0.zip
unzip PaddleOCR-3.1.0.zip
```
安装依赖
@@ -118,17 +122,18 @@ python tools/infer/predict_system.py --image_dir="/workspace/img/" --use_angle_c
ppocr INFO: not find det model file path None
```
### 2 准备数据集
#### 2.1 标注工具
1. [**PPOcrLabel**](https://github.com/PFCCLab/PPOCRLabel),专门为 PPOCR 制作的标注工具,**推荐。**
> 存在问题:EXE 安装包运行会报错。但是该工具在本地运行标注的话还是比较推荐的。
> 存在问题:EXE 安装包运行会报错。但是该工具在本地运行标注的话还是比较推荐的。
2. [**X-AnyLabeling**](https://github.com/CVHub520/X-AnyLabeling),支持多种导出格式,其中就支持 PPOCR 的格式。
#### 2.2 标注完毕后的数据集
```bash
conno/1b4e2845-e9ba-4233-9268-bf07801ace0e.jpg [{"transcription": "RKSU5020243", "points": [[233, 317], [299, 317], [258, 833], [181, 839]], "difficult": false}]
conno/6ffd244b-1d27-4084-998b-e4d7ac076c5f.jpg [{"transcription": "ZGXU6173701", "points": [[717, 448], [781, 419], [835, 980], [751, 996]], "difficult": false}]
@@ -144,6 +149,7 @@ conno/9751b358-0b48-4006-8a73-89e816d37b67.jpg [{"transcription": "CNIU2452605",
```
### 3 准备预训练模型
如果已经有预训练模型的话,可以忽略这一步。
下载地址:https://github.com/PaddlePaddle/PaddleOCR/blob/release/3.1/docs/version3.x/model_list.md
@@ -151,23 +157,26 @@ conno/9751b358-0b48-4006-8a73-89e816d37b67.jpg [{"transcription": "CNIU2452605",
如果报 404 的话,根据 版本 + 路径 进行查找即可。
根据情况选择 mobile 和 server 模型
- mobile: 训练速度快,精度较低,识别速度快
- server: 训练速度慢,精度较高,识别速度慢
#### 3.1 文本检测训练模型
![alt text](assets/paddlepaddle/1776995543202.png)
![alt text](assets/boxocr/1776995543202.png)
#### 3.2 文本识别训练模型
![alt text](assets/paddlepaddle/1776995552047.png)
![alt text](assets/boxocr/1776995552047.png)
### 错误情况
#### ImportError: libGL.so.1: cannot open shared object file: No such file or directory
原因:缺少 `libGL.so.1` 库
解决方案:
```bash
sudo apt update
sudo apt install libgl1-mesa-glx
@@ -179,6 +188,7 @@ sudo apt install libgl1-mesa-glx
解决方案:
将 batch_size_per_card 设置为 1
```yaml
Eval:
loader:
@@ -200,17 +210,19 @@ cp /workspace/PaddleOCR-3.1.0/doc/fonts/simfang.ttf /usr/share/fonts/
[PaddleX 3.0](https://paddlepaddle.github.io/PaddleX/latest/index.html) 是基于飞桨框架构建的低代码开发工具,它集成了众多**开箱即用的预训练模型**,可以实现模型从训练到推理的**全流程开发**,支持国内外**多款主流硬件**,助力AI 开发者进行产业实践。
### 1 环境配置
当前使用的是虚拟环境:
- python:3.11
- paddlepaddle:3.1
- paddleocr:3.1.0
**安装paddlex**
```bash
pip install "paddlex[base]"
```
> 确保已经安装 PaddlePaddle
### 2 部署项目
@@ -218,12 +230,15 @@ pip install "paddlex[base]"
参考文档:https://paddlepaddle.github.io/PaddleX/latest/pipeline_deploy/serving.html#12
#### 2.1 安装服务化部署插件
方便 java 等语言接口的调用
```bash
paddlex --install serving
```
#### 2.2 生成配置文件
使用相关命令,其中 OCR 当前为通用 OCR 产线,需要不同的产线就修改不同的名称即可。
```bash
@@ -274,12 +289,12 @@ SubModules:
unclip_ratio: 1.5
TextLineOrientation:
module_name: textline_orientation
model_name: PP-LCNet_x1_0_textline_ori
model_name: PP-LCNet_x1_0_textline_ori
model_dir: null
batch_size: 6
batch_size: 6
TextRecognition:
module_name: text_recognition
model_name: PP-OCRv5_server_rec
model_name: PP-OCRv5_server_rec
model_dir: null
batch_size: 6
score_thresh: 0.0
@@ -287,7 +302,7 @@ SubModules:
**修改配置文件**
注意将训练好的模型路径配置 `./pipeline/rec_inference` 和 `./pipeline/det_inference`
注意将训练好的模型路径配置 `./pipeline/rec_inference` 和 `./pipeline/det_inference`
```yaml
# 整体 pipeline 名称,用于识别产线名称
@@ -308,9 +323,9 @@ SubModules:
# 文本检测模块(通常是基于 DB 的检测器)
TextDetection:
module_name: text_detection
model_name: PP-OCRv5_mobile_det # 使用的是 mobile 版大模型
model_dir: ./det_inference # 本地模型文件夹路径(需包含 model.pdmodel 等)
# 调大输入尺寸,关注细节
model_name: PP-OCRv5_mobile_det # 使用的是 mobile 版大模型
model_dir: ./det_inference # 本地模型文件夹路径(需包含 model.pdmodel 等)
# 调大输入尺寸,关注细节
limit_side_len: 960
limit_type: min
max_side_limit: 4000
@@ -323,10 +338,10 @@ SubModules:
# 文本识别模块(通常是 CRNN + CTC 或 SVTR 模型)
TextRecognition:
module_name: text_recognition
model_name: PP-OCRv5_mobile_rec # 同样使用的是 mobile 版识别模型
model_name: PP-OCRv5_mobile_rec # 同样使用的是 mobile 版识别模型
model_dir: ./rec_inference
batch_size: 6 # 一次识别图块的数量,适当调大可提高 GPU 利用率
score_thresh: 0.0 # 识别结果置信度下限,低于此不输出
batch_size: 6 # 一次识别图块的数量,适当调大可提高 GPU 利用率
score_thresh: 0.0 # 识别结果置信度下限,低于此不输出
```
#### 2.3 运行服务器
@@ -387,39 +402,37 @@ POST /ocr
- 请求体的属性如下:
| 名称 | 类型 | 含义 | 是否必填 |
| :-------------------------- | :-------- | :----------------------------------------------------------- | :----------------------------------------------------------- |
| `file` | `string` | 服务器可访问的图像文件或PDF文件的URL,或上述类型文件内容的Base64编码结果。默认对于超过10页的PDF文件,只有前10页的内容会被处理。 要解除页数限制,请在产线配置文件中添加以下配置:`Serving: extra: max_num_input_imgs: null ` | 是 |
| `fileType` | `integer` | `null` | 文件类型。`0`表示PDF文件,`1`表示图像文件。若请求体无此属性,则将根据URL推断文件类型。 |
| `visualize` | `boolean` | `null` | 是否返回可视化结果图以及处理过程中的中间图像等。传入 `true`:返回图像。传入 `false`:不返回图像。若请求体中未提供该参数或传入 `null`:遵循产线配置文件`Serving.visualize` 的设置。 例如,在产线配置文件中添加如下字段: `Serving: visualize: False `将默认不返回图像,通过请求体中的`visualize`参数可以覆盖默认行为。如果请求体和配置文件中均未设置(或请求体传入`null`、配置文件中未设置),则默认返回图像。 |
| `useDocOrientationClassify` | `boolean` | `null` | 请参阅产线对象中 `predict` 方法的 `use_doc_orientation_classify` 参数相关说明。 |
| `useDocUnwarping` | `boolean` | `null` | 请参阅产线对象中 `predict` 方法的 `use_doc_unwarping` 参数相关说明。 |
| | | | |
| `useTextlineOrientation` | `boolean` | `null` | 请参阅产线对象中 `predict` 方法的 `use_textline_orientation` 参数相关说明。 |
| `textDetLimitSideLen` | `integer` | `null` | 请参阅产线对象中 `predict` 方法的 `text_det_limit_side_len` 参数相关说明。 |
| `textDetLimitType` | `string` | `null` | 请参阅产线对象中 `predict` 方法的 `text_det_limit_type` 参数相关说明。 |
| `textDetThresh` | `number` | `null` | 请参阅产线对象中 `predict` 方法的 `text_det_thresh` 参数相关说明。 |
| `textDetBoxThresh` | `number` | `null` | 请参阅产线对象中 `predict` 方法的 `text_det_box_thresh` 参数相关说明。 |
| `textDetUnclipRatio` | `number` | `null` | 请参阅产线对象中 `predict` 方法的 `text_det_unclip_ratio` 参数相关说明。 |
| `textRecScoreThresh` | `number` | `null` | 请参阅产线对象中 `predict` 方法的 `text_rec_score_thresh` 参数相关说明。 |
| 名称 | 类型 | 含义 | 是否必填 |
| :-------------------------- | :-------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `file` | `string` | 服务器可访问的图像文件或PDF文件的URL,或上述类型文件内容的Base64编码结果。默认对于超过10页的PDF文件,只有前10页的内容会被处理。 要解除页数限制,请在产线配置文件中添加以下配置:`Serving: extra: max_num_input_imgs: null ` | 是 |
| `fileType` | `integer` | `null` | 文件类型。`0`表示PDF文件,`1`表示图像文件。若请求体无此属性,则将根据URL推断文件类型。 |
| `visualize` | `boolean` | `null` | 是否返回可视化结果图以及处理过程中的中间图像等。传入 `true`:返回图像。传入 `false`:不返回图像。若请求体中未提供该参数或传入 `null`:遵循产线配置文件`Serving.visualize` 的设置。 例如,在产线配置文件中添加如下字段: `Serving: visualize: False `将默认不返回图像,通过请求体中的`visualize`参数可以覆盖默认行为。如果请求体和配置文件中均未设置(或请求体传入`null`、配置文件中未设置),则默认返回图像。 |
| `useDocOrientationClassify` | `boolean` | `null` | 请参阅产线对象中 `predict` 方法的 `use_doc_orientation_classify` 参数相关说明。 |
| `useDocUnwarping` | `boolean` | `null` | 请参阅产线对象中 `predict` 方法的 `use_doc_unwarping` 参数相关说明。 |
| | | | |
| `useTextlineOrientation` | `boolean` | `null` | 请参阅产线对象中 `predict` 方法的 `use_textline_orientation` 参数相关说明。 |
| `textDetLimitSideLen` | `integer` | `null` | 请参阅产线对象中 `predict` 方法的 `text_det_limit_side_len` 参数相关说明。 |
| `textDetLimitType` | `string` | `null` | 请参阅产线对象中 `predict` 方法的 `text_det_limit_type` 参数相关说明。 |
| `textDetThresh` | `number` | `null` | 请参阅产线对象中 `predict` 方法的 `text_det_thresh` 参数相关说明。 |
| `textDetBoxThresh` | `number` | `null` | 请参阅产线对象中 `predict` 方法的 `text_det_box_thresh` 参数相关说明。 |
| `textDetUnclipRatio` | `number` | `null` | 请参阅产线对象中 `predict` 方法的 `text_det_unclip_ratio` 参数相关说明。 |
| `textRecScoreThresh` | `number` | `null` | 请参阅产线对象中 `predict` 方法的 `text_rec_score_thresh` 参数相关说明。 |
- 请求处理成功时,响应体的`result`具有如下属性:
| 名称 | 类型 | 含义 |
| :----------- | :------- | :----------------------------------------------------------- |
| 名称 | 类型 | 含义 |
| :----------- | :------- | :---------------------------------------------------------------------------------------------------------------------------------------------- |
| `ocrResults` | `object` | OCR结果。数组长度为1(对于图像输入)或实际处理的文档页数(对于PDF输入)。对于PDF输入,数组中的每个元素依次表示PDF文件中实际处理的每一页的结果。 |
| `dataInfo` | `object` | 输入数据信息。 |
| `dataInfo` | `object` | 输入数据信息。 |
`ocrResults`中的每个元素为一个`object`,具有如下属性:
| 名称 | 类型 | 含义 |
| :---------------------- | :------- | :----------------------------------------------------------- |
| 名称 | 类型 | 含义 |
| :---------------------- | :------- | :------------------------------------------------------------------------------------------------------------------- |
| `prunedResult` | `object` | 产线对象的 `predict` 方法生成结果的 JSON 表示中 `res` 字段的简化版本,其中去除了 `input_path` 和 `page_index` 字段。 |
| `ocrImage` | `string` | `null` |
| `docPreprocessingImage` | `string` | `null` |
| `inputImage` | `string` | `null` |
| `ocrImage` | `string` | `null` |
| `docPreprocessingImage` | `string` | `null` |
| `inputImage` | `string` | `null` |
**Java调用实例**
@@ -562,15 +575,16 @@ public class Test {
}
```
### 错误情况
#### Please use PaddlePaddle with GPU version.
原因:当前 PaddlePaddle 版本不是GPU的,而是CPU的。
解决方法:安装GPU版本的PaddlePaddle
#### ImportError:DLL load failed while importing cv2:找不到指定的模块
当前环境:Windows Server 2012
**安装 Microsoft Visual C++ Redistributable**
@@ -581,36 +595,39 @@ public class Test {
- [Visual C++ Redistributable for Visual Studio 2015, 2017 and 2019(32 位系统)](https://aka.ms/vs/16/release/vc_redist.x86.exe)
**检查 Python 和 OpenCV 的位数是否匹配**
- 确保你安装的 Python 版本(32 位或 64 位)与 OpenCV 的位数一致。
- 你可以通过以下命令检查 Python 的位数:
```bash
python -c "import struct; print(struct.calcsize('P') * 8)"
```
- 如果不匹配,卸载并重新安装正确位数的 Python 和 OpenCV。
安装完成后,重启计算机。
安装完成后,重启计算机。
**一定要打开桌面实验**
1. 打开服务器面板,选择 **添加角色和功能**
![alt text](assets/paddlepaddle/1776995933823.png)
![alt text](assets/boxocr/1776995933823.png)
2. 在功能处开启 **桌面体验**
![alt text](assets/paddlepaddle/1776995956696.png)
![alt text](assets/boxocr/1776995956696.png)
如果开启 桌面体验 报错:**尚未开启 WinRM 服务**, 这个时候需要开启 `Windows Remote Management (WS-Management) `
![alt text](assets/paddlepaddle/1776995962380.png)
![alt text](assets/boxocr/1776995962380.png)
如果在开启过程中报错:**错误1068:依存服务或组无法启动**
点击属性,找到 **依存关系**,确保依赖的服务已开启
![alt text](assets/paddlepaddle/1776995970106.png)
![alt text](assets/boxocr/1776995970106.png)
检查 `HTTP Service` 服务是否启动
```bash
sc query http
```
如果出现 `STATE : 1 STOPPED` ,那么就代表该服务已被禁用,解决方案如下:
1. 启用 HTTP 服务
```bash
sc config http start= auto
@@ -627,4 +644,4 @@ sc query http
```bash
winrm invoke Restore winrm/Config @{}
winrm quickconfig -q
```
```
@@ -0,0 +1 @@
# 单证智能识别(PP-StructureV3,LLM)