Token导航 LogoToken导航TokenDH.com
研究检索敏感数据unknown未标认证来源可访问许可证需确认审计未展示

paddleocr-text-recognitionpaddleocr 文本识别

Agent Skill

paddleocr-text-recognition 用于查找、检索和筛选相关信息,适合在 Local Agent 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

212

周安装

9

下载量

74
Local Agent

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:paddleocr-text-recognition(paddleocr 文本识别)
来源仓库:https://skills.volces.com
仓库路径:paddleocr-text-recognition
安装命令:
Scripts declare their dependencies inline ( PEP 723 ). No separate install step is needed — uv resolves dependencies automatically:
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 对应宿主 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.sh安装方式未标明
Scripts declare their dependencies inline ( PEP 723 ). No separate install step is needed — uv resolves dependencies automatically:

简介

paddleocr-text-recognition 用于图像中文本识别与提取,支持 OCR 任务处理。

  • 适用于扫描件、截图或 PDF 中的文字内容转换场景。
  • 通过上传图片调用,输出识别后的文本与位置信息。
  • 使用前需确认图片清晰度与语言类型是否符合模型训练范围。
  • 建议查阅原始文档了解支持的字体与排版复杂度限制。

SKILL.md

PaddleOCR Text Recognition Skill

When to Use This Skill

Trigger keywords (routing): Bilingual trigger terms (Chinese and English) are listed in the YAML description above—use that field for discovery and routing.

Use this skill for:

  • Extract text from images (screenshots, photos, scans)
  • Extract text from PDFs or document images when the goal is line/box-level text, not recovering table grids, formulas, or full reading-order layout
  • Extract text from URLs or local files that point to images/PDFs

Do not use for:

  • Plain text files, code files, or markdown documents that can be read directly as text
  • Documents with tables, formulas, charts, or complex layouts — use Document Parsing instead
  • Tasks that do not involve image-to-text conversion

Installation

Scripts declare their dependencies inline (PEP 723). No separate install step is needed — uv resolves dependencies automatically:

uv run scripts/ocr_caller.py --help

How to Use This Skill

Working directory: All uv run scripts/... commands below should be run from this skill's root directory (the directory containing this SKILL.md file).

Basic Workflow

  1. Identify the input source:

- User provides URL: Use the --file-url parameter - User provides local file path: Use the --file-path parameter

  1. Execute OCR: uv run scripts/ocr_caller.py --file-url "URL provided by user" --pretty Or for local files: uv run scripts/ocr_caller.py --file-path "file path" --pretty Performance note: Parsing time scales with document complexity. Single-page images typically complete in 1-3 seconds; large PDFs (50+ pages) may take several minutes. Allow adequate time before assuming a timeout. Default behavior: save raw JSON to a temp file:

- If --output is omitted, the script saves automatically under the system temp directory - Default path pattern: <system-temp>/paddleocr/text-recognition/results/result_<timestamp>_<id>.json - If --output is provided, it overrides the default temp-file destination - If --stdout is provided, JSON is printed to stdout and no file is saved - In save mode, the script prints the absolute saved path on stderr: Result saved to: /absolute/path/... - In default/custom save mode, read and parse the saved JSON file before responding - Use --stdout only when you explicitly want to skip file persistence

  1. Parse JSON response:

- In default/custom save mode, load JSON from the saved file path shown by the script - Check the ok field: true means success, false means error - Extract text: text field contains all recognized text - If --stdout is used, parse the stdout JSON directly - Handle errors: If ok is false, display error.message

  1. Present results to user:

- Display extracted text in a readable format - If the text is empty, the image may contain no text - In save mode, always tell the user the saved file path and that full raw JSON is available there

What to Do After Extraction

Common next steps once you have the recognized text:

  • Save to file: Write the text field to a .txt or .md file
  • Search the content: Search the saved output file for keywords
  • Feed to another pipeline: The text field is clean plain text, ready for downstream processing
  • Poor results: See "Tips for Better Results" below before retrying

Complete Output Display

Always display the COMPLETE recognized text to the user. The user typically needs the full content for downstream use — truncation silently loses data they may not notice is missing.

  • Display the entire text field, no matter how long
  • Do not use phrases like "Here's a summary" or "The text begins with..."
  • Do not truncate with "..." unless the text truly exceeds reasonable display limits (>10,000 chars)

Example - Correct:

User: "Extract the text from this image"
Agent: I've extracted the text from the image. Here's the complete content:

[Display the entire text here]

Example - Incorrect:

User: "Extract the text from this image"
Agent: I found some text in the image. Here's a preview:
"The quick brown fox..." (truncated)

Understanding the Output

The script returns a JSON envelope with ok, text, result, and error fields. Use text for the recognized content; result contains the raw API response for debugging.

For the full schema and field-level details, see references/output_schema.md.

Raw result location (default): the temp-file path printed by the script on stderr

Usage Examples

Example 1: URL OCR

uv run scripts/ocr_caller.py --file-url "https://example.com/invoice.jpg" --pretty

Example 2: Local File OCR

uv run scripts/ocr_caller.py --file-path "./document.pdf" --pretty

Example 3: OCR With Explicit File Type

uv run scripts/ocr_caller.py --file-url "https://example.com/input" --file-type 1 --pretty
  • --file-type 0: PDF
  • --file-type 1: image
  • If omitted, the type is auto-detected from the file extension. For local files, a recognized extension (.pdf, .png, .jpg, .jpeg, .bmp, .tiff, .tif, .webp) is required; otherwise pass --file-type explicitly. For URLs with unrecognized extensions, the service attempts inference.

Example 4: Print JSON Without Saving

uv run scripts/ocr_caller.py --file-url "https://example.com/input" --stdout --pretty

First-Time Configuration

When API is not configured, the script outputs:

{
  "ok": false,
  "text": "",
  "result": null,
  "error": {
    "code": "CONFIG_ERROR",
    "message": "PADDLEOCR_OCR_API_URL not configured. Get your API at: https://paddleocr.com"
  }
}

Configuration workflow:

  1. Show the exact error message to the user.
  2. Guide the user to obtain credentials: Visit the PaddleOCR website, click API, select the PP-OCRv5 model, select the language, then copy the API_URL and Token. They map to these environment variables: Optionally configure PADDLEOCR_OCR_TIMEOUT for request timeout. Recommend using the host application's standard configuration method rather than pasting credentials in chat.

- PADDLEOCR_OCR_API_URL — full endpoint URL ending with /ocr - PADDLEOCR_ACCESS_TOKEN — 40-character alphanumeric string

  1. Apply credentials — one of:

- User configured via the host UI: ask the user to confirm, then retry. - User pastes credentials in chat: warn that they may be stored in conversation history, help the user persist them using the host's standard configuration method, then retry.

Error Handling

All errors return JSON with ok: false. Show the error message and stop — do not fall back to your own vision capabilities. Identify the issue from error.code and error.message:

Authentication failed (403)error.message contains "Authentication failed"

  • Token is invalid, reconfigure with correct credentials

Quota exceeded (429)error.message contains "API rate limit exceeded"

  • Daily API quota exhausted, inform user to wait or upgrade

Unsupported formaterror.message contains "Unsupported file format"

  • File format not supported, convert to PDF/PNG/JPG

No text detected:

  • text field is empty
  • Image may be blank, corrupted, or contain no text

Tips for Better Results

If recognition quality is poor:

  • Low resolution: Provide a higher resolution image (≥300 DPI works well for most printed text)
  • Noisy background: A cleaner scan or screenshot typically yields better results than a phone photo
  • Check confidence: The raw JSON (result.result.ocrResults[n].prunedResult.rec_scores) shows per-line confidence scores — low values identify uncertain regions worth reviewing

Reference Documentation

  • references/output_schema.md — Full output schema, field descriptions, and command examples
Note: Model version, capabilities, and supported file formats are determined by your API endpoint (PADDLEOCR_OCR_API_URL) and its official API documentation.

Testing the Skill

To verify the skill is working properly:

uv run scripts/smoke_test.py
uv run scripts/smoke_test.py --skip-api-test
uv run scripts/smoke_test.py --test-url "https://..."

The first form tests configuration and API connectivity. --skip-api-test checks configuration only. --test-url overrides the default sample image URL.

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Local Agent

83.57%
按下载量换算62

安全审计

暂无安全审计结果可展示。

权限和风险

敏感数据

该 Skill 可能接触密钥、Token、环境变量或敏感配置,应进入高风险复核队列,默认不自动发布。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills