Skip to content

feat(parse): support task-based MinerU APIs and the online batch service - #4725

Open
now-ing wants to merge 1 commit into
volcengine:mainfrom
now-ing:feat/iss-4708-mineru-async-api
Open

now-ing wants to merge 1 commit into
volcengine:mainfrom
now-ing:feat/iss-4708-mineru-async-api

Conversation

@now-ing

@now-ing now-ing commented Sep 6, 2026

Copy link
Copy Markdown
Contributor

Fixes #4708

Problem

The PDF parser only implemented the legacy self-hosted MinerU contract: one POST {mineru_endpoint}/file_parse answering inline with md_content. Both newer deployments break today:

  • the current self-hosted API (mineru master) is task-based (POST /tasks → 202 → status_url/result_url → zip result) — the inline endpoint is gone;
  • the online batch API (mineru.net/api/v4) has no /file_parse (404), no /health (so the startup preflight hard-fails the service), and requires a Bearer token that had nowhere to be configured.

Change

  • pdf.mineru_api_mode ("auto" | "sync" | "async", default "auto"): auto detects the protocol from the first response — inline completed+results stays v1; a task-shaped payload (or a 404 on /file_parse) switches to the task flow — so existing v1 deployments need zero config changes.
  • pdf.mineru_token sent as Authorization: Bearer ... on every call.
  • v2-tasks flow: POST /tasks → poll status_url until status=completed → download result_url zip.
  • online-batch flow: POST /extract/task/batch → poll extract-results/batch/{batch_id} until state=done → download full_zip_url.
  • zip handling reuses safe_extract_zip (zip-slip guarded), persists images/ into the media store and rewrites markdown references — mirroring the existing v1 image path.
  • startup preflight (/health) is skipped for the async flavors (the online service has no such endpoint; this was part of the reported startup errors).
  • docs (en/zh): protocol table + online example.

Wire-format evidence: the v2 contract was taken from the official mineru/cli/api_client.py (submit → 202 → status_url/result_url, pending/processing/completed); the online endpoints were verified live (/api/v4/extract/task/batch and /api/v4/extract-results/batch/{id} respond 401 login required without a token, /api/v4/file_parse responds 404 — which is exactly the reported failure).

Testing

tests/parse/test_pdf_mineru.py (new, 14 tests): v1 unchanged under sync and auto; v2 flow end-to-end (202 → poll → zip → markdown + image rewrite); online batch flow (with 404 fallthrough from /tasks); auto fallback on 404; task failure detail propagation; polling timeout; missing-task-endpoint actionable error; zip without markdown; config validation.

All green locally (48 passed incl. test_mineru_preflight, test_pdf_bookmark_extraction, test_parser_config_wiring, test_markdown_local_image_refs).

The PDF parser assumed the legacy self-hosted contract: a single
POST /file_parse answering inline with md_content. The current
self-hosted /tasks API and the online batch API are task-based
(create -> poll -> zip download), so both made OpenViking fail:
the online service has no /file_parse endpoint at all (404) and no
/health endpoint for the startup preflight.

- add pdf.mineru_api_mode ("auto" | "sync" | "async", default "auto"):
  auto detects the flavor from the first response, so existing v1
  deployments keep working without config changes
- add pdf.mineru_token, sent as Authorization: Bearer on every call
- v2-tasks flow: POST /tasks -> poll status_url (pending/processing/
  completed) -> download the result zip
- online-batch flow: POST /extract/task/batch -> poll
  extract-results/batch/{batch_id} until state=done -> download
  full_zip_url
- unpack the zip via safe_extract_zip, persist images/ into the media
  store and rewrite markdown references (mirrors the v1 image path)
- skip the /health startup preflight for the async flavors (the
  online service has no such endpoint)
- docs (en/zh): protocol table and online example

Fixes volcengine#4708
@q2000s

q2000s commented Sep 8, 2026

Copy link
Copy Markdown

Enhancement: Add mineru-first strategy

Motivation

当前 �uto 策略是 本地 pdfplumber 优先 → MinerU 回退。但对于需要高精度解析的场景(复杂版面、表格、公式),用户可能希望 MinerU 优先 → 本地 pdfplumber 回退。

在本地测试中,MinerU 的解析质量明显优于 pdfplumber,尤其是表格结构识别、数学公式保留、多栏版面处理。因此建议增加 mineru-first 策略。

Changes

1. parser_config.py - 验证逻辑

`python

Before:

if self.strategy not in ("local", "mineru", "auto"):

After:

if self.strategy not in ("local", "mineru", "auto", "mineru-first"):
`

2. pdf.py - 新增策略分支

python elif self.config.strategy == "mineru-first": # Try MinerU API first; fall back to local pdfplumber on failure try: return await self._convert_mineru(pdf_path, resource_name=resource_name) except Exception as e: logger.warning(f"MinerU API failed: {e}") logger.info("Falling back to local pdfplumber") return await self._convert_local(pdf_path, resource_name=resource_name)

3. core.py - 预检逻辑

`python

Before:

pdf_config.strategy == "mineru"

After:

pdf_config.strategy in ("mineru", "mineru-first")
`

Strategy Comparison

Strategy 优先级 回退行为 适用场景
local 本地 only 无 无 MinerU 部署
mineru MinerU only 无 强制使用 MinerU
�uto 本地 → MinerU 本地失败才用 MinerU 默认,兼容现有部署
mineru-first MinerU → 本地 MinerU 失败才用本地 高精度需求

Test Results

已验证 mineru-first 策略正常工作:

  1. MinerU API 可用时:使用 MinerU 解析
  2. MinerU API 不可用时:自动回退到本地 pdfplumber

@q2000s

q2000s commented Sep 8, 2026

Copy link
Copy Markdown

@now-ing I have made an improvement based on your PR, see #4818 . Functions are illustrated as above. Please take a look and decide if you want to combine my PR into yours...

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

[Feature]: 增加对最新在线mineru精确解析API服务格式的支持

2 participants