ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

Claude Code 100个真实案例 - 用AI搭建OCR文字识别服务(身份证+发票+车牌)

Claude Code 100个真实案例 - 用AI搭建OCR文字识别服务(身份证+发票+车牌) 1. 从零搭建 OCR 服务为什么我选 Claude Code PaddleOCR 这套组合OCR 文字识别服务说白了就是把图片里的文字变成能编辑、能入库、能检索的文本。身份证、发票、车牌这三类图像是业务系统里出现频率最高的三种结构化文档银行开户要读身份证财务报销要读发票停车场抬杆要读车牌。它们的共同点是版式相对固定、字段可枚举所以特别适合用 OCR 加规则提取的方式做成一个可复用的服务。我这次要分享的是用 Claude Code 从零把一个多场景 OCR 服务搭起来。技术栈是 PaddleOCR 做核心识别、OpenCV 做图像预处理、FastAPI 做接口封装。为什么是这三样PaddleOCR 对中文和英文混排的识别效果在开源方案里属于第一梯队而且自带方向分类器处理身份证这种可能拍歪的图很省心OpenCV 负责去噪、增强、倾斜矫正把脏图变成干净图再喂给识别引擎准确率能明显往上走FastAPI 负责把整条链路包成一个 HTTP 服务前端、后端、脚本都能调。适合谁看如果你正在做证件识别、票据录入、车辆管理这类需求又不想一上来就买商业 OCR 的额度那这套自建方案很值得跑一遍。整篇文章我会给出可复制的目录结构、依赖清单、PaddleOCR 初始化参数、OpenCV 增强代码和 FastAPI 路由配置最后用 curl 把三类图片的识别输出都验证一遍。你跟着做能跑通端到端流程。在动手写代码之前先解决一个容易被忽略但很关键的问题Claude Code 这类编码助手在长会话里会频繁调用模型如果每次都走官方直连成本和稳定性都不太可控。我的做法是给它配一个统一的模型接入层把 Base URL、Key、Model ID 三件套固定下来这样无论是 Claude Code 还是后面要接的其他工具都走同一套配置。下面先说这块前置准备。2. 前置准备给 Claude Code 配好模型接入层再谈写代码很多人搭 OCR 服务卡住不是卡在 PaddleOCR 装不上而是卡在编码助手本身跑不起来——会话一长就断、请求一多就报错。所以第一步不是写 config.py而是先把 Claude Code 的模型接入配置理顺。我用的方式是走一个统一的 API 网关把模型调用收敛到一个入口。这样做的直接好处是Base URL 只配一次Key 只存一份Model ID 只改一处后面不管换模型还是加工具都不用满项目找配置。TaoToken 就是干这个的它提供兼容主流协议的统一接入Claude Code、Cline、Codex 这些工具都能接。具体配置三件套是这样的Base URLhttps://taotoken.net/apiAPI Key在控制台的 API Keys 页面生成形如sk-开头的一串Model ID按你实际要用的模型填比如claude-sonnet-4-5这类如果你用的是 Claude Code 的 settings 配置方式可以在项目根目录或用户目录下建一个 settings 文件把接入信息写进去。下面是一个可复制的 JSON 片段路径按你本地的实际位置放{ env: { ANTHROPIC_BASE_URL: https://taotoken.net/api, ANTHROPIC_API_KEY: sk-你的Key, ANTHROPIC_MODEL: claude-sonnet-4-5 } }注意这里 Base URL 用的是https://taotoken.net/api不要多加路径后缀工具会自己拼接。Key 一定要从控制台生成后复制别手敲容易漏字符。Model ID 要和你在控制台里开通的模型一致填错了会直接报模型不存在。如果你用的是 Codex 那套auth.json的配置方式结构类似把 base_url、api_key、model 三个字段对应填上就行。Cline 的 MCP 配置也是同一个思路Base URL 指向网关Key 填生成的Model ID 选你要的。三件套对齐了工具才能正常发请求。配好之后先别急着写 OCR 代码用一句最简单的对话验证一下接入是否通。打开 Claude Code问它一个简单问题比如帮我列一下 Python 读取图片的三种方式。如果能正常返回说明接入层没问题可以进入下一步。如果报 401多半是 Key 错了或没生效如果报连接失败检查 Base URL 有没有写错、网络是否可达。这一步看起来和 OCR 无关但它是后面所有代码生成的前提。接入层稳了Claude Code 才能在你写 config、写预处理、写路由的时候持续给你补全和纠错。下面进入正题先把项目骨架和依赖定下来。3. 项目骨架与依赖config.py、models.py 和可复制的 pyproject.toml搭服务我习惯先把目录和配置定死再往里填逻辑。这样 Claude Code 生成代码时也有明确的落点不会东一榔头西一棒子。项目目录结构如下ocr-service/ ├── config.py # 全局配置 ├── models.py # 数据模型 ├── preprocessor.py # 图像预处理 ├── ocr_engine.py # OCR 识别引擎 ├── extractors.py # 结构化信息提取器 ├── app.py # FastAPI 服务 ├── data/ │ └── temp/ # 临时文件目录 └── pyproject.toml # 项目依赖依赖清单我用 uv 管理pyproject.toml里核心依赖是这样[project] name ocr-service version 0.1.0 requires-python 3.11 dependencies [ paddlepaddle2.6.0, paddleocr2.7.0, opencv-python-headless4.9.0, fastapi0.115.0, uvicorn0.30.0, python-multipart0.0.9, pillow10.0.0, numpy1.26.0, ]初始化命令就两行uv init ocr-service cd ocr-service uv add paddlepaddle paddleocr opencv-python-headless fastapi uvicorn python-multipart pillow numpyconfig.py用 dataclass 把配置集中管理PaddleOCR 的初始化参数、预处理开关、上传限制都放这里。关键参数我解释一下langch表示中英文通用模型身份证和发票都是中文为主这个够用use_gpuFalse是 CPU 模式本地跑没显卡也能用有卡再改 Trueconfidence_threshold0.6是识别置信度阈值低于这个值的文本块直接丢掉能过滤掉不少噪声use_angle_clsTrue开启方向分类处理拍歪的证件很关键。from dataclasses import dataclass, field from typing import Optional dataclass class OCRConfig: lang: str ch use_gpu: bool False confidence_threshold: float 0.6 use_angle_cls: bool True det_model_dir: Optional[str] None rec_model_dir: Optional[str] None max_image_size: int 4096 dataclass class PreprocessConfig: denoise: bool True denoise_kernel: int 3 binarize: bool False binary_threshold: int 127 enhance_contrast: bool True contrast_alpha: float 1.5 contrast_beta: int 10 dataclass class AppConfig: ocr: OCRConfig field(default_factoryOCRConfig) preprocess: PreprocessConfig field(default_factoryPreprocessConfig) max_file_size: int 10 * 1024 * 1024 allowed_extensions: list field( default_factorylambda: [.jpg, .jpeg, .png, .bmp, .tiff, .webp] ) host: str 0.0.0.0 port: int 8002 temp_dir: str ./data/temp config AppConfig()models.py定义请求和响应的数据结构用 Pydantic 做校验。三类文档各有自己的结果模型身份证有姓名、性别、民族、出生日期、住址、身份证号发票有发票代码、号码、开票日期、购买方、销售方、金额、税额车牌有车牌号、颜色、类型、置信度。统一响应模型OCRResponse把 success、ocr_type、result、elapsed_ms、message 包在一起前端处理起来很省事。from typing import Optional from pydantic import BaseModel, Field class OCRTextBlock(BaseModel): text: str confidence: float position: list[list[int]] class GeneralOCRResult(BaseModel): text_blocks: list[OCRTextBlock] Field(default_factorylist) full_text: str block_count: int 0 avg_confidence: float 0.0 class IDCardResult(BaseModel): side: str front name: Optional[str] None gender: Optional[str] None ethnicity: Optional[str] None birth_date: Optional[str] None address: Optional[str] None id_number: Optional[str] None issuing_authority: Optional[str] None valid_period: Optional[str] None raw_text: list[str] Field(default_factorylist) class InvoiceResult(BaseModel): invoice_type: Optional[str] None invoice_code: Optional[str] None invoice_number: Optional[str] None invoice_date: Optional[str] None buyer_name: Optional[str] None buyer_tax_id: Optional[str] None seller_name: Optional[str] None total_amount: Optional[str] None tax_amount: Optional[str] None total_with_tax: Optional[str] None raw_text: list[str] Field(default_factorylist) class LicensePlateResult(BaseModel): plate_number: Optional[str] None plate_color: Optional[str] None plate_type: Optional[str] None confidence: float 0.0 position: list[list[int]] Field(default_factorylist) raw_text: list[str] Field(default_factorylist) class OCRResponse(BaseModel): success: bool True ocr_type: str general result: dict Field(default_factorydict) elapsed_ms: int 0 message: str 配置和模型定好之后先跑一句验证配置能加载python -c from config import config; print(fOCR配置: 语言{config.ocr.lang}, GPU{config.ocr.use_gpu})输出OCR配置: 语言ch, GPUFalse就说明骨架没问题。这一步别跳过配置错了后面全白搭。接下来写图像预处理这是提升识别率的关键一环。4. 图像预处理与识别引擎OpenCV 增强 PaddleOCR 初始化参数图像预处理这块我的原则是能增强就增强但别过度处理。身份证、发票、车牌这三类图常见问题是拍得暗、有噪点、轻微倾斜。OpenCV 的GaussianBlur去噪、convertScaleAbs增强对比度、HoughLinesP检测倾斜角这三招组合起来效果最稳。preprocessor.py的核心是process方法按尺寸标准化 → 对比度增强 → 去噪 → 倾斜矫正 → 二值化的顺序走。顺序很重要先增强再二值化比反过来效果好因为二值化会丢信息。倾斜矫正默认关闭因为霍夫变换有计算开销只有明显歪的图才开。import math from typing import Optional import cv2 import numpy as np from config import config class ImagePreprocessor: def __init__(self): self.config config.preprocess self.max_size config.ocr.max_image_size def process(self, image, denoiseNone, binarizeNone, enhanceNone, deskewFalse): image self.resize_image(image) if enhance or (enhance is None and self.config.enhance_contrast): image self.enhance_contrast(image) if denoise or (denoise is None and self.config.denoise): image self.denoise(image) if deskew: image self.deskew(image) if binarize or (binarize is None and self.config.binarize): image self.binarize(image) return image def resize_image(self, image): h, w image.shape[:2] max_dim max(h, w) if max_dim self.max_size: scale self.max_size / max_dim image cv2.resize(image, (int(w * scale), int(h * scale)), interpolationcv2.INTER_AREA) return image def to_grayscale(self, image): if len(image.shape) 3: return cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) return image def denoise(self, image): kernel_size self.config.denoise_kernel if kernel_size % 2 0: kernel_size 1 return cv2.GaussianBlur(image, (kernel_size, kernel_size), 0) def binarize(self, image): gray self.to_grayscale(image) _, binary cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY cv2.THRESH_OTSU) return cv2.cvtColor(binary, cv2.COLOR_GRAY2BGR) def enhance_contrast(self, image): return cv2.convertScaleAbs(image, alphaself.config.contrast_alpha, betaself.config.contrast_beta) def deskew(self, image): gray self.to_grayscale(image) edges cv2.Canny(gray, 50, 150, apertureSize3) lines cv2.HoughLinesP(edges, 1, np.pi / 180, threshold100, minLineLength100, maxLineGap10) if lines is None: return image angles [] for line in lines: x1, y1, x2, y2 line[0] angle math.degrees(math.atan2(y2 - y1, x2 - x1)) if abs(angle) 30: angles.append(angle) if not angles: return image median_angle np.median(angles) if abs(median_angle) 0.5: return image h, w image.shape[:2] center (w // 2, h // 2) matrix cv2.getRotationMatrix2D(center, median_angle, 1.0) return cv2.warpAffine(image, matrix, (w, h), flagscv2.INTER_CUBIC, borderModecv2.BORDER_REPLICATE)ocr_engine.py负责初始化 PaddleOCR 并封装识别方法。初始化参数里use_angle_cls、lang、use_gpu都从 config 读show_logFalse关掉 PaddleOCR 自己的日志输出避免和 FastAPI 的日志混在一起。识别结果里每个文本块带坐标和置信度低于阈值的直接过滤。import os import cv2 import numpy as np from paddleocr import PaddleOCR from config import config from models import GeneralOCRResult, OCRTextBlock from preprocessor import ImagePreprocessor class OCREngine: def __init__(self): print(正在初始化PaddleOCR引擎...) self.ocr PaddleOCR( use_angle_clsconfig.ocr.use_angle_cls, langconfig.ocr.lang, use_gpuconfig.ocr.use_gpu, show_logFalse, ) self.preprocessor ImagePreprocessor() self.confidence_threshold config.ocr.confidence_threshold print(PaddleOCR引擎初始化完成) def recognize(self, image, preprocessTrue, deskewFalse): if preprocess: image self.preprocessor.process(image, deskewdeskew) result self.ocr.ocr(image, clsconfig.ocr.use_angle_cls) text_blocks [] all_texts [] if result and result[0]: for line in result[0]: position line[0] text line[1][0] confidence line[1][1] if confidence self.confidence_threshold: continue pos [[int(p[0]), int(p[1])] for p in position] text_blocks.append(OCRTextBlock(texttext, confidenceround(confidence, 4), positionpos)) all_texts.append(text) avg_conf 0.0 if text_blocks: avg_conf round(sum(b.confidence for b in text_blocks) / len(text_blocks), 4) return GeneralOCRResult( text_blockstext_blocks, full_text\n.join(all_texts), block_countlen(text_blocks), avg_confidenceavg_conf, ) def recognize_from_bytes(self, image_bytes, **kwargs): nparr np.frombuffer(image_bytes, np.uint8) image cv2.imdecode(nparr, cv2.IMREAD_COLOR) if image is None: raise ValueError(无法解码图片数据) return self.recognize(image, **kwargs)验证预处理和引擎能跑起来python -c import numpy as np from preprocessor import ImagePreprocessor img np.random.randint(0, 255, (600, 800, 3), dtypenp.uint8) p ImagePreprocessor() out p.process(img) print(f原图: {img.shape}, 处理后: {out.shape}) 输出尺寸一致就说明预处理链路通了。这里有个坑要提醒opencv-python-headless和opencv-python不能同时装会冲突服务器环境用 headless 版本就行不需要 GUI。接下来写三类文档的结构化提取器。5. 三类文档结构化提取身份证、发票、车牌的正则与字段映射OCR 识别出来是一堆文本块业务要的是结构化字段。这一步靠正则和关键词匹配来做。三类文档各有各的套路我分开说。身份证提取的关键是判断正反面。正面有姓名性别民族出生住址公民身份号码反面有签发机关有效期限。判断逻辑就是看全文里有没有签发机关或有效期限有就是反面。身份证号用 18 位正则匹配出生日期用年月日格式匹配住址可能跨行需要把下一行拼上来。import re from models import GeneralOCRResult, IDCardResult, InvoiceResult, LicensePlateResult class IDCardExtractor: ID_PATTERN re.compile(r(\d{6}(?:19|20)\d{2}(?:0[1-9]|1[0-2])(?:0[1-9]|[12]\d|3[01])\d{3}[\dXx])) BIRTH_PATTERN re.compile(r(\d{4})\s*年\s*(\d{1,2})\s*月\s*(\d{1,2})\s*日) VALID_PATTERN re.compile(r(\d{4}\.\d{2}\.\d{2})\s*[-—]\s*(\d{4}\.\d{2}\.\d{2}|长期)) def extract(self, ocr_result: GeneralOCRResult) - IDCardResult: texts [b.text for b in ocr_result.text_blocks] full_text .join(texts) result IDCardResult(raw_texttexts) if 签发机关 in full_text or 有效期限 in full_text: result.side back for text in texts: if 签发机关 in text: result.issuing_authority text.replace(签发机关, ).strip() m self.VALID_PATTERN.search(full_text) if m: result.valid_period f{m.group(1)} - {m.group(2)} else: result.side front for i, text in enumerate(texts): if 姓名 in text: name text.replace(姓名, ).strip() if not name and i 1 len(texts): name texts[i 1].strip() result.name name if 男 in text: result.gender 男 elif 女 in text: result.gender 女 m self.BIRTH_PATTERN.search(text) if m: result.birth_date f{m.group(1)}年{m.group(2)}月{m.group(3)}日 if 住址 in text: addr text.replace(住址, ).strip() if i 1 len(texts) and not any(k in texts[i1] for k in [姓名, 性别, 民族, 出生, 号码]): addr texts[i 1].strip() result.address addr m self.ID_PATTERN.search(full_text) if m: result.id_number m.group(1).upper() return result发票提取的重点是发票代码、号码、日期、金额。发票代码是 10 到 12 位数字号码是 8 位数字金额前面通常有 ¥ 符号。购买方和销售方名称一般在购买方或销售方关键词的下一行。class InvoiceExtractor: CODE_PATTERN re.compile(r发票代码[:\s]*(\d{10,12})) NUMBER_PATTERN re.compile(r发票号码[:\s]*(\d{8})) AMOUNT_PATTERN re.compile(r[¥]\s*([\d,]\.?\d*)) DATE_PATTERN re.compile(r(\d{4})\s*年\s*(\d{1,2})\s*月\s*(\d{1,2})\s*日) def extract(self, ocr_result: GeneralOCRResult) - InvoiceResult: texts [b.text for b in ocr_result.text_blocks] full_text .join(texts) result InvoiceResult(raw_texttexts) if 增值税专用发票 in full_text: result.invoice_type 增值税专用发票 elif 增值税普通发票 in full_text: result.invoice_type 增值税普通发票 elif 电子发票 in full_text: result.invoice_type 电子发票 m self.CODE_PATTERN.search(full_text) if m: result.invoice_code m.group(1) m self.NUMBER_PATTERN.search(full_text) if m: result.invoice_number m.group(1) m self.DATE_PATTERN.search(full_text) if m: result.invoice_date f{m.group(1)}年{m.group(2)}月{m.group(3)}日 for i, text in enumerate(texts): if 购买方 in text and i 1 len(texts): result.buyer_name texts[i 1].replace(名称, ).replace(, ).replace(:, ).strip() if 销售方 in text and i 1 len(texts): result.seller_name texts[i 1].replace(名称, ).replace(, ).replace(:, ).strip() if 合计 in text and 金额 in text: m self.AMOUNT_PATTERN.search(text) if m: result.total_amount m.group(1) if 税额 in text and 合计 in text: m self.AMOUNT_PATTERN.search(text) if m: result.tax_amount m.group(1) if 价税合计 in text: m self.AMOUNT_PATTERN.search(text) if m: result.total_with_tax m.group(1) return result车牌提取靠正则匹配省份简称加字母数字的组合。普通车牌 7 位新能源车牌 8 位警用车牌以警结尾教练车牌以学结尾。匹配到之后按长度和结尾字符判断类型和颜色。class LicensePlateExtractor: PLATE_PATTERN re.compile( r([京津冀沪渝冀豫云辽黑湘皖鲁新苏浙赣鄂桂甘晋蒙陕吉闽贵粤川青藏琼宁]) r[A-HJ-NP-Z][A-HJ-NP-Z0-9]{4,5}[A-HJ-NP-Z0-9挂学警港澳] ) def extract(self, ocr_result: GeneralOCRResult) - LicensePlateResult: texts [b.text for b in ocr_result.text_blocks] result LicensePlateResult(raw_texttexts) best_match None best_conf 0.0 best_pos [] for block in ocr_result.text_blocks: clean block.text.replace( , ).replace(-, ).upper() m self.PLATE_PATTERN.search(clean) if m and block.confidence best_conf: best_match m.group(0) best_conf block.confidence best_pos block.position if best_match: result.plate_number best_match result.confidence best_conf result.position best_pos if len(best_match) 7: result.plate_type 普通车牌 result.plate_color 蓝色 elif len(best_match) 8: result.plate_type 新能源车牌 result.plate_color 绿色 if best_match.endswith(警): result.plate_type 警用车辆 result.plate_color 白色 elif best_match.endswith(学): result.plate_type 教练车牌 result.plate_color 黄色 return result class ExtractorFactory: _extractors { id_card: IDCardExtractor, invoice: InvoiceExtractor, license_plate: LicensePlateExtractor, } classmethod def get_extractor(cls, doc_type: str): extractor_class cls._extractors.get(doc_type) if extractor_class is None: raise ValueError(f不支持的类型: {doc_type}支持: {list(cls._extractors.keys())}) return extractor_class()验证车牌正则python -c from extractors import LicensePlateExtractor e LicensePlateExtractor() for t in [京A12345, 粤BD12345, 沪C88888]: m e.PLATE_PATTERN.search(t) print(f{t} - {m.group(0) if m else \未匹配\}) 输出三行匹配结果就说明提取器正常。这里有个经验车牌正则里的省份简称列表要写全漏一个省份就会导致那个地方的车牌识别不出来。另外新能源车牌是 8 位正则里{4,5}要能覆盖到。接下来把这些模块用 FastAPI 串起来。6. FastAPI 路由封装与 curl 验证三类图片端到端跑通app.py把预处理、识别、提取串成 HTTP 接口。用lifespan在服务启动时初始化 OCREngine避免每次请求都重新加载模型。路由分五个通用识别、身份证、发票、车牌、批量。每个路由都先校验文件格式再读字节流调引擎识别最后按类型走对应的提取器。import logging import os import time from contextlib import asynccontextmanager from typing import Optional from fastapi import FastAPI, File, HTTPException, Query, UploadFile from fastapi.middleware.cors import CORSMiddleware from config import config from extractors import ExtractorFactory from models import OCRResponse from ocr_engine import OCREngine logging.basicConfig(levellogging.INFO, format%(asctime)s [%(levelname)s] %(message)s) logger logging.getLogger(ocr-service) ocr_engine: Optional[OCREngine] None asynccontextmanager async def lifespan(app: FastAPI): global ocr_engine logger.info(正在初始化OCR服务...) os.makedirs(config.temp_dir, exist_okTrue) ocr_engine OCREngine() logger.info(OCR服务初始化完成) yield logger.info(OCR服务已关闭) app FastAPI(titleOCR文字识别服务, version1.0.0, lifespanlifespan) app.add_middleware(CORSMiddleware, allow_origins[*], allow_methods[*], allow_headers[*]) def _validate_image(file: UploadFile): if not file.filename: raise HTTPException(status_code400, detail缺少文件名) ext os.path.splitext(file.filename)[1].lower() if ext not in config.allowed_extensions: raise HTTPException(status_code400, detailf不支持的格式: {ext}) app.post(/api/ocr/general, response_modelOCRResponse) async def general_ocr(file: UploadFile File(...), preprocess: bool Query(True), deskew: bool Query(False)): _validate_image(file) start time.time() try: image_bytes await file.read() result ocr_engine.recognize_from_bytes(image_bytes, preprocesspreprocess, deskewdeskew) elapsed int((time.time() - start) * 1000) return OCRResponse(successTrue, ocr_typegeneral, resultresult.model_dump(), elapsed_mselapsed, messagef识别到 {result.block_count} 个文本块) except Exception as e: logger.error(f通用OCR失败: {e}) raise HTTPException(status_code500, detailstr(e)) app.post(/api/ocr/id_card, response_modelOCRResponse) async def id_card_ocr(file: UploadFile File(...)): _validate_image(file) start time.time() try: image_bytes await file.read() ocr_result ocr_engine.recognize_from_bytes(image_bytes) extractor ExtractorFactory.get_extractor(id_card) id_result extractor.extract(ocr_result) elapsed int((time.time() - start) * 1000) return OCRResponse(successTrue, ocr_typeid_card, resultid_result.model_dump(), elapsed_mselapsed, messagef身份证{id_result.side}面识别完成) except Exception as e: logger.error(f身份证OCR失败: {e}) raise HTTPException(status_code500, detailstr(e)) app.post(/api/ocr/invoice, response_modelOCRResponse) async def invoice_ocr(file: UploadFile File(...)): _validate_image(file) start time.time() try: image_bytes await file.read() ocr_result ocr_engine.recognize_from_bytes(image_bytes) extractor ExtractorFactory.get_extractor(invoice) inv_result extractor.extract(ocr_result) elapsed int((time.time() - start) * 1000) return OCRResponse(successTrue, ocr_typeinvoice, resultinv_result.model_dump(), elapsed_mselapsed, messagef发票识别完成: {inv_result.invoice_type or 未知类型}) except Exception as e: logger.error(f发票OCR失败: {e}) raise HTTPException(status_code500, detailstr(e)) app.post(/api/ocr/license_plate, response_modelOCRResponse) async def license_plate_ocr(file: UploadFile File(...)): _validate_image(file) start time.time() try: image_bytes await file.read() ocr_result ocr_engine.recognize_from_bytes(image_bytes) extractor ExtractorFactory.get_extractor(license_plate) plate_result extractor.extract(ocr_result) elapsed int((time.time() - start) * 1000) return OCRResponse(successTrue, ocr_typelicense_plate, resultplate_result.model_dump(), elapsed_mselapsed, messagef车牌识别: {plate_result.plate_number or 未检测到车牌}) except Exception as e: logger.error(f车牌OCR失败: {e}) raise HTTPException(status_code500, detailstr(e)) app.get(/api/ocr/types) async def list_types(): return { types: [ {name: general, description: 通用文字识别}, {name: id_card, description: 身份证识别}, {name: invoice, description: 增值税发票识别}, {name: license_plate, description: 车牌号识别}, ], supported_formats: config.allowed_extensions, max_file_size: f{config.max_file_size / 1024 / 1024:.0f}MB, } if __name__ __main__: import uvicorn uvicorn.run(app:app, hostconfig.host, portconfig.port, reloadTrue)启动服务uv run python app.py看到Uvicorn running on http://0.0.0.0:8002就说明服务起来了。然后用 curl 验证三类图片。先测通用识别curl -X POST http://localhost:8002/api/ocr/general -F filetest_image.png返回结构里text_blocks是文本块数组每个带 text、confidence、positionfull_text是拼接后的全文。再测身份证curl -X POST http://localhost:8002/api/ocr/id_card -F fileidcard_front.jpg返回的result里应该有 name、gender、id_number 这些字段。发票和车牌同理把 URL 换成/api/ocr/invoice和/api/ocr/license_plate就行。批量接口用-F filesa.jpg -F filesb.jpg传多个文件。如果识别结果里字段是 null先看raw_text里有没有对应的文本。有文本但没提取出来说明正则要调没文本说明 OCR 没识别到要回头看预处理参数。这套流程跑通之后你可以把服务部署到内网前端直接调接口就行。7. 常见报错排查401、连接失败、识别为空、字段缺失怎么定位搭这套服务报错基本集中在四类我按出现频率排一下。第一类是 401 未授权。这个几乎都出在模型接入层不是 OCR 代码的问题。表现是 Claude Code 或调用 API 时返回 401提示 invalid api key 或 unauthorized。排查顺序先确认 Key 是不是从控制台复制的完整字符串有没有多空格再确认 Base URL 是不是https://taotoken.net/api有没有多加/v1之类的后缀最后确认 Model ID 是不是控制台里开通的那个。三件套里任何一个不对都会 401。改完配置记得重启工具有些工具会缓存配置。第二类是连接失败或超时。表现是请求发不出去报 connection refused 或 timeout。先确认服务本身有没有起来curl http://localhost:8002/api/ocr/types能不能返回。如果服务正常但外部调不通检查防火墙和端口映射。如果是模型接入层连不上检查网络是否可达Base URL 有没有写错。这类问题九成是地址写错或端口没开。第三类是识别结果为空。表现是接口返回 200但text_blocks是空数组block_count为 0。原因通常是图片质量太差或预处理过度。排查方法先把preprocess参数设为 false看原图能不能识别出东西。如果能说明是预处理把文字弄没了把binarize关掉、contrast_alpha调小试试。如果原图也识别不出检查图片是不是太小、太模糊或者文字方向完全颠倒。PaddleOCR 的use_angle_cls能处理 180 度旋转但极端情况还是要先人工摆正。第四类是字段缺失。表现是 OCR 识别出了文本但结构化字段是 null。比如身份证识别出了姓名张三但name字段是空的。这通常是正则或关键词匹配的问题。排查方法先看raw_text里文本长什么样是不是有空格、换行、特殊符号干扰。比如姓 名中间有空格姓名 in text就匹配不到。这时候要么在预处理里去掉空格要么把匹配逻辑改成更宽松的正则。发票的金额字段也容易出问题因为金额前面可能是 ¥ 也可能是 正则里两个都要覆盖。还有一类是依赖冲突。表现是import paddleocr报错或者cv2找不到。最常见的是opencv-python和opencv-python-headless同时装了卸载一个就行。PaddleOCR 对 numpy 版本也有要求太新的 numpy 可能不兼容按 pyproject.toml 里锁的版本装。如果装 paddlepaddle 时报平台不支持确认 Python 版本是不是 3.11 以上以及系统架构是不是 x86_64。排查的时候有个通用技巧把日志级别调到 INFO看每一步的输出。app.py里已经配了 logging识别失败会打 error 日志。如果日志里看到PaddleOCR引擎初始化完成但请求还是失败问题就在请求处理逻辑如果连初始化都没完成问题在依赖或模型下载。PaddleOCR 首次运行会下载模型文件网络不好的话会卡住可以提前把模型下好放到指定目录。8. 把服务用起来接入文档、模型对话和长期编码的三条路径服务跑通之后怎么把它用起来有三条路径可以走。第一条是接入文档。FastAPI 自带 Swagger UI启动服务后访问http://localhost:8002/docs就能看到所有接口的在线文档可以直接在页面上传图片测试。如果你要把服务给前端或第三方调把这份文档地址给他们就行。更完整的接入说明和参数细节可以看接入文档里面有 Base URL、鉴权方式、请求示例的完整说明。第二条是模型对话验证。如果你不确定某个模型对中文证件的识别效果可以先用模型对话快速试几张图对比不同模型的表现再决定生产环境用哪个。这样不用改代码就能做模型选型。第三条是长期编码和 Agent 场景。如果你打算把 OCR 服务作为长期项目维护或者要接进更大的 Agent 工作流建议用 Coding Plan它更适合持续性的编码任务和工具链集成配置一次就能长期用。三条路径对应不同的使用强度偶尔测一下用模型对话正式接入用 API Keys 加接入文档长期开发用 Coding Plan。按你的实际场景选就行。最后说个实操经验这套服务我建议先在内网跑一段时间用真实业务图片压一压把正则和预处理参数调到稳定再对外提供服务。OCR 这东西实验室效果和真实场景效果差距往往在图片质量上多收集一些脏图来测比调参更重要。
返回列表