Skip to content

feat(pdf): add mineru-first strategy for MinerU-priority fallback - #4818

Open
q2000s wants to merge 7 commits into
volcengine:mainfrom
q2000s:feat/iss-4708-mineru-async-api
Open

q2000s wants to merge 7 commits into
volcengine:mainfrom
q2000s:feat/iss-4708-mineru-async-api

Conversation

@q2000s

@q2000s q2000s commented Sep 8, 2026 •

Copy link
Copy Markdown

Enhancement: Add mineru-first strategy for MinerU-priority fallback

Motivation

PR #4725 引入了 mineru_api_mode (auto/sync/async) 来支持 MinerU 的三种协议形态,非常棒的功能。但在实际部署中发现一个需求缺口:

当前 auto 策略的行为是 本地 pdfplumber 优先 → MinerU 回退。但对于需要高精度解析的场景(复杂版面、表格、公式),用户可能希望 MinerU 优先 → 本地 pdfplumber 回退。

在本地测试中,MinerU 的解析质量明显优于 pdfplumber,尤其是:

  • 表格结构识别
  • 数学公式保留
  • 多栏版面处理
  • 图片位置关联

因此建议增加 mineru-first 策略。

Changes

1. openviking_cli/utils/config/parser_config.py

# Before:
strategy: str = "auto"  # "local" | "mineru" | "auto"

# After:
strategy: str = "auto"  # "local" | "mineru" | "auto" | "mineru-first"
# Before (validate method):
if self.strategy not in ("local", "mineru", "auto"):
    raise ValueError(
        f"Invalid strategy '{self.strategy}'. Must be 'local', 'mineru', or 'auto'"
    )

# After:
if self.strategy not in ("local", "mineru", "auto", "mineru-first"):
    raise ValueError(
        f"Invalid strategy '{self.strategy}'. Must be 'local', 'mineru', 'auto', or 'mineru-first'"
    )

2. openviking/parse/parsers/pdf.py

# Before:
        if self.config.strategy == "local":
            return await self._convert_local(pdf_path, resource_name=resource_name)

        elif self.config.strategy == "mineru":
            return await self._convert_mineru(pdf_path, resource_name=resource_name)

        elif self.config.strategy == "auto":
            # Try local first
            try:
                return await self._convert_local(pdf_path, resource_name=resource_name)
            except Exception as e:
                logger.warning(f"Local conversion failed: {e}")

                # Fallback to MinerU if configured
                if self.config.mineru_endpoint:
                    logger.info("Falling back to MinerU API")
                    return await self._convert_mineru(pdf_path, resource_name=resource_name)
                else:
                    raise ValueError(
                        f"Local conversion failed and no MinerU endpoint configured: {e}"
                    )

# After:
        if self.config.strategy == "local":
            return await self._convert_local(pdf_path, resource_name=resource_name)

        elif self.config.strategy == "mineru":
            return await self._convert_mineru(pdf_path, resource_name=resource_name)

        elif self.config.strategy == "mineru-first":
            # Try MinerU API first; fall back to local pdfplumber on failure
            try:
                return await self._convert_mineru(pdf_path, resource_name=resource_name)
            except Exception as e:
                logger.warning(f"MinerU API failed: {e}")
                logger.info("Falling back to local pdfplumber")
                return await self._convert_local(pdf_path, resource_name=resource_name)

        elif self.config.strategy == "auto":
            # Try local first
            try:
                return await self._convert_local(pdf_path, resource_name=resource_name)
            except Exception as e:
                logger.warning(f"Local conversion failed: {e}")

                # Fallback to MinerU if configured
                if self.config.mineru_endpoint:
                    logger.info("Falling back to MinerU API")
                    return await self._convert_mineru(pdf_path, resource_name=resource_name)
                else:
                    raise ValueError(
                        f"Local conversion failed and no MinerU endpoint configured: {e}"
                    )

3. openviking/service/core.py

# Before:
        pdf_config = self._config.pdf
        should_preflight_mineru = (
            pdf_config.strategy == "mineru"
            or (pdf_config.strategy == "auto" and pdf_config.mineru_endpoint is not None)
        ) and pdf_config.mineru_api_mode != "async"

        if should_preflight_mineru and pdf_config.mineru_endpoint:
            try:
                await wait_for_mineru_ready(pdf_config.mineru_endpoint)
                logger.info("MinerU preflight passed: %s", pdf_config.mineru_endpoint)
            except RuntimeError as exc:
                if pdf_config.strategy == "mineru":
                    raise

# After:
        pdf_config = self._config.pdf
        should_preflight_mineru = (
            pdf_config.strategy in ("mineru", "mineru-first")
            or (pdf_config.strategy == "auto" and pdf_config.mineru_endpoint is not None)
        ) and pdf_config.mineru_api_mode != "async"

        if should_preflight_mineru and pdf_config.mineru_endpoint:
            try:
                await wait_for_mineru_ready(pdf_config.mineru_endpoint)
                logger.info("MinerU preflight passed: %s", pdf_config.mineru_endpoint)
            except RuntimeError as exc:
                if pdf_config.strategy in ("mineru", "mineru-first"):
                    raise

Strategy Comparison

Strategy 优先级 回退行为 适用场景
local 本地 only 无 无 MinerU 部署
mineru MinerU only 无 强制使用 MinerU
auto 本地 → MinerU 本地失败才用 MinerU 默认,兼容现有部署
mineru-first MinerU → 本地 MinerU 失败才用本地 高精度需求, MinerU 优先

Configuration Example

{
  "pdf": {
    "strategy": "mineru-first",
    "mineru_endpoint": "http://127.0.0.1:8899",
    "mineru_api_mode": "auto",
    "mineru_token": "your-online-api-token"
  }
}

Test Results (Local)

已验证 mineru-first 策略在以下场景正常工作:

  1. MinerU API 可用时:使用 MinerU 解析,返回高质量结果
  2. MinerU API 不可用时:自动回退到本地 pdfplumber,不中断流程
  3. 与 mineru_api_mode: auto 配合:自动检测协议形态(v1-sync/v2-tasks/online-batch)

Summary

This enhancement adds mineru-first as a new strategy option that mirrors the existing auto strategy but with reversed priority. It's a small, self-contained change that gives users more flexibility in balancing quality vs. availability.

…nd mineru-first

Port the MinerU work onto current main.

PDF parsing
- Add async MinerU support: the task-based contracts (v2-tasks and
  online-batch) are now driven through _mineru_run_task_flow /
  _mineru_continue_task, with polling, zip download, safe extraction
  and markdown picking. The legacy single-shot /file_parse contract is
  kept and reached through _convert_mineru_v1.
- Select the protocol via the new pdf.mineru_api_mode
  ("auto" | "sync" | "async"). "auto" probes /file_parse first and
  falls back to the task protocol when the endpoint no longer offers it.
- Add pdf.mineru_token, sent as "Authorization: Bearer <token>" on
  every MinerU request, so the cloud API can be used.
- Add the "mineru-first" strategy: try MinerU, fall back to local
  pdfplumber.

Preflight
- Skip the synchronous readiness probe for "mineru-first" and for the
  async api modes: those paths use the task endpoints, so the probe does
  not describe the request that is actually made. Failures surface (and
  fall back) on the first real parse instead.
@q2000s
q2000s force-pushed the feat/iss-4708-mineru-async-api branch from 73c4f40 to 9d9f672 Compare October 2, 2026 03:50
@ZaynJarvis ZaynJarvis added the kernel Core server, runtime, storage, retrieval, SDK, and CLI. label Oct 5, 2026
MinerU v4 has no multipart upload and never exposed /file_parse, so the
existing MinerU support could not talk to mineru.net at all. Against
https://mineru.net/api/v4 every parse died with 405 Method Not Allowed,
because:

  - auto mode only fell through to the task protocol on 404, and the cloud
    API answers an unknown POST path with 405;
  - an unrecognized response body was coerced to "v2-tasks" and then raised
    KeyError: 'task_id' instead of trying the next endpoint;
  - HTTP 200 with a non-zero application "code" was treated as success;
  - _mineru_online_busy did not know the documented "waiting-file" state, so
    the upload handshake was reported as a failure.

Local files now use the v4 three-step flow that the docs describe:
POST /file-urls/batch (JSON) to reserve an OSS url, PUT the file, then poll
/extract-results/batch/{batch_id}. The image field also differs per protocol
(files for self-hosted, file for the hosted batch API), so it is a parameter.

v0.5.0 dropped the whole task-based MinerU block (881 lines, no
_mineru_run_task_flow), so this grafts the branch's MinerU section onto the
0.5.0 source rather than copying files wholesale. The two files that also
needed edits are shipped as install-time patches under tools/patches/0.5.0
instead of being committed over the release, because they sit outside the
MinerU block: core.py skips the sync preflight for async and mineru-first,
and parser_config.py adds mineru_api_mode, mineru_token and the
mineru-first strategy. Keeping them as patches avoids grafting unrelated
post-0.4.22 code (such as the reindex_processor import) onto older wheels.

tools/apply-openviking-patch.ps1 applies a version-matched patch set, backs
up every file it replaces, and can roll back. It refuses to run against a
version it has no patch set for, so an upgrade cannot silently land a build
grafted onto the wrong release.

Verified on 0.5.0: three PDFs import through mineru.net/api/v4 in ~288s,
producing markdown, extracted images and per-directory summaries, with no
failed files; existing indexed resources remain readable.
-Rollback restored only files recorded in the newest backup, so a file that was
already patched when it was last written came back still patched: rolling back
once did not undo the patch. That defeats the point of having a rollback, since
it can leave a half-applied tree behind with no obvious sign.

Walk every backup directory and keep the earliest copy of each file instead,
which is the state from before this script touched anything.
The v4 precision API needs a bearer token and answers 401 without one, which
aborted the whole parse before any fallback could run: HTTPStatusError is not
ValueError or MinerUAPIError, so the existing except clauses let it through.

Treat 401/403 from the v4 submit as an endpoint rejection, then fall back to the
Agent API, which needs no credentials: POST /api/v1/agent/parse/file for a
signed url, PUT, then poll /api/v1/agent/parse/{task_id} for markdown_url. It
returns markdown at a CDN link rather than a zip and extracts no images, so the
result is used as-is with images_saved=0.

Also fixes a missing import that only the zip path hit: _mineru_unpack_zip calls
normalize_zip_filenames, which v0.5.0's pdf.py does not import because the
release has no MinerU code at all. Without it every v4 parse that got as far as
unpacking raised NameError.

Verified on 0.5.0 with the same three PDFs:
  - with a token, v4 online-batch, 7-14s, 2-9 images each;
  - without a token, agent-file, 7-9s, markdown only;
  - through ov add-resource, 3/3 with no failed files, 21 artifacts and 6 images.
@q2000s

q2000s commented Oct 10, 2026

Copy link
Copy Markdown
Author

Now, this PR has been compatible with v0.5 of OV. Please take a look.

ov add-resource returns as soon as the task is queued, so waiting on its output
never ends and tells you nothing. Every wait here has a ceiling, and giving up
never cancels the task: the server owns it and may still finish.

Two bugs in the sync script came out of adding the waits:

  - it recorded a project's fingerprint straight after submitting the import, so
    a task that later failed still left the project marked synced and it was
    never retried. Failed and timed-out projects now keep their old fingerprint;
  - -Verify was declared but never used, so a dry run started importing
    everything, which is how a "check" turned into a 15-minute import.

Wait-OvTask reports rather than kills on timeout, and Invoke-OvAddResource reads
the task result (root_uri, failed_files) instead of parsing submit output.

Verified: -Verify now returns immediately reporting 0 projects to reimport.
…archive

Directory imports pack the tree into data/temp/upload/upload_<id>.zip. While one
import is still running, the next one fails with

  [PERMISSION_DENIED] [WinError 32] another process is using this file

which reads like a permissions fault and is not one: retrying just leaves more
orphaned archives behind. Ten of them had piled up to 10 GB before this was
noticed. Assert-OvIdle now checks for a running add_resource up front and exits
with status 2 having submitted nothing.

Also normalises -Project so it accepts comma-separated values. Start-Process
-ArgumentList has no array syntax, so "A,B" arrived as one literal path named "A,B"
and was reported as missing rather than as two projects.
OpenViking's memory extraction does not ask for a JSON array; the orchestrator in
session/compressor_v3.py drives it through tool calling (memory.extract.candidates.
tool_skill), and after four turns it raises "Final response could not be parsed as
operations". Measured against the extraction prompt:

  model                        plain JSON   tool calling
  tencent/Hunyuan-MT-7B        3/3, 3.3s    HTTP 400
  Qwen/Qwen3-8B                1/3, ~140s   yes
  PaddlePaddle/PaddleOCR-VL-1.5 0/3         HTTP 400
  deepseek-ai/DeepSeek-OCR     0/3          HTTP 400

Hunyuan wins the plain-JSON benchmark by a wide margin and is 40x faster, but it is
a translation model and rejects tools outright, so it cannot drive the orchestrator.
Qwen3-8B is the only candidate that satisfies both, and with it a real commit
completes in 5.5s instead of failing after 100s.

Window sizes, measured: DeepSeek-OCR 8192, PaddleOCR-VL-1.5 16384,
Qwen/Qwen3.5-4B 262144. The 8192 window was the original blocker, since a session
reached 9992 input tokens; max_tokens is now 4096.

Embedding and rerank were benchmarked on retrieval quality rather than latency
(recall@1 over six query/document pairs, three repeats). BAAI/bge-m3 and
BAAI/bge-large-zh-v1.5 both reach recall@1 0.83 / MRR 0.92 at 1024 dims, so the
current embedding model is kept. BAAI/bge-reranker-v2-m3 scores 0.83 and stays in
place. One real finding: SiliconFlow rejects the `dimensions` field on
/embeddings with 400 for bge-m3, so the vector length is whatever the model
returns and `dimensions` must not be sent.

ov.conf.production is the working configuration with credentials redacted.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kernel Core server, runtime, storage, retrieval, SDK, and CLI.

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

3 participants