Token导航 LogoToken导航TokenDH.com
研究检索需要联网clawhub未标认证来源可访问clear审计通过

paddleocr-doc-parsingpaddleocr 文档解析

Agent Skill

用于辅助文档、README、Markdown、说明文和内容稿件的整理与改写。它适合让 Agent 提炼结构、补齐章节、统一术语、检查链接或把零散材料整理成可读文档。使用时应保留项目已有事实、命令和路径,不要把未确认的信息写成确定结论;涉及对外文案时,还需要控制语气,避免过度营销或夸大能力。

总安装

297,523

周安装

11,920

GitHub Stars

43

下载量

96,314
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:paddleocr-doc-parsing(paddleocr 文档解析)
来源仓库:https://github.com/bobholamovic/paddleocr-doc-parsing
安装命令:
openclaw skills install paddleocr-doc-parsing
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install paddleocr-doc-parsing

简介

从PDF和文档图像提取结构化Markdown/JSON数据。

  • 支持单元格级精度表格和LaTeX公式识别。
  • 可识别印章、图表等特殊文档元素。paddleocr-doc-parsing 属于研究检索类 Skill,可作为该场景下的辅助能力补充。
  • 保持原始文档结构和布局信息完整。
  • 建议预处理图像质量以获得更好识别效果。

SKILL.md

name
paddleocr-doc-parsing
description
>-
compatibility
Requires Python 3.9+, uv, and internet access.
metadata
openclaw
requires
env
bins
primaryEnv
PADDLEOCR_ACCESS_TOKEN
emoji
📄
homepage
https://github.com/PaddlePaddle/PaddleOCR/tree/main/skills/paddleocr-doc-parsing

PaddleOCR Document Parsing Skill

When to Use This Skill

Trigger keywords (routing): Bilingual trigger terms (Chinese and English) are listed in the YAML description above—use that field for discovery and routing.

Use this skill for:

  • Documents with tables (invoices, financial reports, spreadsheets)
  • Documents with mathematical formulas (academic papers, scientific documents)
  • Documents with charts and diagrams
  • Multi-column layouts (newspapers, magazines, brochures)
  • Complex document structures requiring layout analysis
  • Any document requiring structured understanding

Do not use for:

  • Simple text-only extraction
  • Quick OCR tasks where speed is critical
  • Screenshots or simple images with clear text

Installation

Scripts declare their dependencies inline (PEP 723). No separate install step is needed — uv resolves dependencies automatically:

uv run scripts/layout_caller.py --help

How to Use This Skill

Working directory: All uv run scripts/... commands below should be run from this skill's root directory (the directory containing this SKILL.md file).

Basic Workflow

  1. Identify the input source:

- User provides URL: Use the --file-url parameter - User provides local file path: Use the --file-path parameter

  1. Execute document parsing:
   uv run scripts/layout_caller.py --file-url "URL provided by user" --pretty

Or for local files:

   uv run scripts/layout_caller.py --file-path "file path" --pretty

Optional: explicitly set file type:

   uv run scripts/layout_caller.py --file-url "URL provided by user" --file-type 0 --pretty

- --file-type 0: PDF - --file-type 1: image - If omitted, the type is auto-detected from the file extension. For local files, a recognized extension (.pdf, .png, .jpg, .jpeg, .bmp, .tiff, .tif, .webp) is required; otherwise pass --file-type explicitly. For URLs with unrecognized extensions, the service attempts inference.

> Performance note: Parsing time scales with document complexity. Single-page images typically complete in 1-5 seconds; large PDFs (50+ pages) may take several minutes. Allow adequate time before assuming a timeout.

Default behavior: save raw JSON to a temp file: - If --output is omitted, the script saves automatically under the system temp directory - Default path pattern: <system-temp>/paddleocr/doc-parsing/results/result_<timestamp>_<id>.json - If --output is provided, it overrides the default temp-file destination - If --stdout is provided, JSON is printed to stdout and no file is saved - In save mode, the script prints the absolute saved path on stderr: Result saved to: /absolute/path/... - In default/custom save mode, read and parse the saved JSON file before responding - Use --stdout only when you explicitly want to skip file persistence

  1. Parse JSON response:

- Check the ok field: true means success, false means error - The output contains complete document data: text, tables, formulas (LaTeX), figures, seals, headers/footers, and reading order - Use the appropriate field based on what the user needs: - text — full document text across all pages - result.result.layoutParsingResults[n].markdown.text — page-level markdown - result.result.layoutParsingResults[n].prunedResult — structured layout data with positions and confidence - Handle errors: If ok is false, display error.message

  1. Present results to user:

- Display content based on what the user requested (see "Complete Output Display" below) - If the content is empty, the document may contain no extractable text - In save mode, always tell the user the saved file path and that full raw JSON is available there

What to Do After Parsing

Common next steps once you have the structured output:

  • Save as Markdown: Write the text field to a .md file — tables, headings, and formulas are preserved
  • Extract specific tables: Navigate result.result.layoutParsingResults[n].prunedResult to access individual layout elements with position and confidence data
  • Feed to RAG / search pipeline: The text field is structured markdown, ready for chunking and indexing
  • Poor results: See "Tips for Better Results" below before retrying

Complete Output Display

Display the COMPLETE extracted content based on what the user asked for. The parsed output is only useful if the user receives all of it — truncation silently drops data.

  • If user asks for "all text", show the entire text field
  • If user asks for "tables", show ALL tables in the document
  • If user asks for "main content", filter out headers/footers but show ALL body text
  • Do not truncate with "..." unless content is excessively long (>10,000 chars)
  • Do not say "Here's a preview" when user expects complete output

Example - Correct:

User: "Extract all the text from this document"
Agent: I've parsed the complete document. Here's all the extracted text:

[Display entire text field or concatenated regions in reading order]

Document Statistics:
- Total regions: 25
- Text blocks: 15
- Tables: 3
- Formulas: 2
Quality: Excellent (confidence: 0.92)

Example - Incorrect:

User: "Extract all the text"
Agent: "I found a document with multiple sections. Here's the beginning:
'Introduction...' (content truncated for brevity)"

Understanding the Output

The script returns an envelope with ok, text, result, and error. Use text for the full document content; navigate result.result.layoutParsingResults[n] for per-page structured data.

For the complete schema and field-level details, see references/output_schema.md.

Raw result location (default): the temp-file path printed by the script on stderr

Usage Examples

Example 1: Extract Full Document Text

uv run scripts/layout_caller.py \
  --file-url "https://example.com/paper.pdf" \
  --pretty

Then use:

  • Top-level text for quick full-text output
  • result.result.layoutParsingResults[n].markdown when page-level output is needed

Example 2: Extract Structured Page Data

uv run scripts/layout_caller.py \
  --file-path "./financial_report.pdf" \
  --pretty

Then use:

  • result.result.layoutParsingResults[n].prunedResult for structured parsing data (layout/content/confidence)

Example 3: Print JSON to stdout (without saving to file)

uv run scripts/layout_caller.py \
  --file-url "URL" \
  --stdout \
  --pretty

By default the script writes JSON to a temp file and prints the path to stderr. Add --stdout to print the full JSON directly to stdout instead. Use this when you need to inspect the result inline or pipe it to another tool.

First-Time Configuration

When API is not configured, the script outputs:

{
  "ok": false,
  "text": "",
  "result": null,
  "error": {
    "code": "CONFIG_ERROR",
    "message": "PADDLEOCR_DOC_PARSING_API_URL not configured. Get your API at: https://paddleocr.com"
  }
}

Configuration workflow:

  1. Show the exact error message to the user.
  1. Guide the user to obtain credentials: Visit the PaddleOCR website, click API, select a model (PP-StructureV3, PaddleOCR-VL, or PaddleOCR-VL-1.5), then copy the API_URL and Token. They map to these environment variables:

- PADDLEOCR_DOC_PARSING_API_URL — full endpoint URL ending with /layout-parsing - PADDLEOCR_ACCESS_TOKEN — 40-character alphanumeric string

Optionally configure PADDLEOCR_DOC_PARSING_TIMEOUT for request timeout. Recommend using the host application's standard configuration method rather than pasting credentials in chat.

  1. Apply credentials — one of:

- User configured via the host UI: ask the user to confirm, then retry. - User pastes credentials in chat: warn that they may be stored in conversation history, help the user persist them using the host's standard configuration method, then retry.

Handling Large Files

For PDFs, the maximum is 100 pages per request.

Optimize Large Images Before Parsing

For large image files, compress before uploading — this reduces upload time and can improve processing stability:

uv run scripts/optimize_file.py input.png output.jpg --quality 85
uv run scripts/layout_caller.py --file-path "output.jpg" --pretty

--quality controls JPEG/WebP lossy compression (1-100, default 85); it has no effect on PNG output. Use --target-size (in MB, default 20) to set the max file size — the script iteratively downscales until the target is met.

Use URL for Large Local Files (Recommended)

For very large local files, prefer --file-url over --file-path to avoid base64 encoding overhead:

uv run scripts/layout_caller.py --file-url "https://your-server.com/large_file.pdf"

Process Specific Pages (PDF Only)

If you only need certain pages from a large PDF, extract them first:

# Extract pages 1-5
uv run scripts/split_pdf.py large.pdf pages_1_5.pdf --pages "1-5"

# Mixed ranges are supported
uv run scripts/split_pdf.py large.pdf selected_pages.pdf --pages "1-5,8,10-12"

# Then process the smaller file
uv run scripts/layout_caller.py --file-path "pages_1_5.pdf"

Error Handling

All errors return JSON with ok: false. Show the error message and stop — do not fall back to your own vision capabilities. Identify the issue from error.code and error.message:

Authentication failed (403)error.message contains "Authentication failed"

  • Token is invalid, reconfigure with correct credentials

Quota exceeded (429)error.message contains "API rate limit exceeded"

  • Daily API quota exhausted, inform user to wait or upgrade

Unsupported formaterror.message contains "Unsupported file format"

  • File format not supported, convert to PDF/PNG/JPG

No content detected:

  • text field is empty
  • Document may be blank, image-only, or contain no extractable text

Tips for Better Results

If parsing quality is poor:

  • Large or high-resolution images: Compress with optimize_file.py before parsing — oversized inputs can degrade layout detection:
  uv run scripts/optimize_file.py input.png optimized.jpg --quality 85
  • Check confidence: result.result.layoutParsingResults[n].prunedResult includes confidence scores per layout element — low values indicate regions worth reviewing

Reference Documentation

  • references/output_schema.md — Full output schema, field descriptions, and command examples
Note: Model version and capabilities are determined by your API endpoint (PADDLEOCR_DOC_PARSING_API_URL).

Testing the Skill

To verify the skill is working properly:

uv run scripts/smoke_test.py
uv run scripts/smoke_test.py --skip-api-test
uv run scripts/smoke_test.py --test-url "https://..."

The first form tests configuration and API connectivity. --skip-api-test checks configuration only. --test-url overrides the default sample document URL.

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

81.66%
按下载量换算78,650

安全审计

VirusTotal

通过

ClawScan

通过

Static analysis

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills