Repository navigation
Conversation
…nd mineru-first
Port the MinerU work onto current main.
PDF parsing
- Add async MinerU support: the task-based contracts (v2-tasks and
online-batch) are now driven through _mineru_run_task_flow /
_mineru_continue_task, with polling, zip download, safe extraction
and markdown picking. The legacy single-shot /file_parse contract is
kept and reached through _convert_mineru_v1.
- Select the protocol via the new pdf.mineru_api_mode
("auto" | "sync" | "async"). "auto" probes /file_parse first and
falls back to the task protocol when the endpoint no longer offers it.
- Add pdf.mineru_token, sent as "Authorization: Bearer <token>" on
every MinerU request, so the cloud API can be used.
- Add the "mineru-first" strategy: try MinerU, fall back to local
pdfplumber.
Preflight
- Skip the synchronous readiness probe for "mineru-first" and for the
async api modes: those paths use the task endpoints, so the probe does
not describe the request that is actually made. Failures surface (and
fall back) on the first real parse instead.
q2000s
force-pushed
the
feat/iss-4708-mineru-async-api
branch
from
October 2, 2026 03:50
73c4f40 to
9d9f672
Compare
MinerU v4 has no multipart upload and never exposed /file_parse, so the existing MinerU support could not talk to mineru.net at all. Against https://mineru.net/api/v4 every parse died with 405 Method Not Allowed, because: - auto mode only fell through to the task protocol on 404, and the cloud API answers an unknown POST path with 405; - an unrecognized response body was coerced to "v2-tasks" and then raised KeyError: 'task_id' instead of trying the next endpoint; - HTTP 200 with a non-zero application "code" was treated as success; - _mineru_online_busy did not know the documented "waiting-file" state, so the upload handshake was reported as a failure. Local files now use the v4 three-step flow that the docs describe: POST /file-urls/batch (JSON) to reserve an OSS url, PUT the file, then poll /extract-results/batch/{batch_id}. The image field also differs per protocol (files for self-hosted, file for the hosted batch API), so it is a parameter. v0.5.0 dropped the whole task-based MinerU block (881 lines, no _mineru_run_task_flow), so this grafts the branch's MinerU section onto the 0.5.0 source rather than copying files wholesale. The two files that also needed edits are shipped as install-time patches under tools/patches/0.5.0 instead of being committed over the release, because they sit outside the MinerU block: core.py skips the sync preflight for async and mineru-first, and parser_config.py adds mineru_api_mode, mineru_token and the mineru-first strategy. Keeping them as patches avoids grafting unrelated post-0.4.22 code (such as the reindex_processor import) onto older wheels. tools/apply-openviking-patch.ps1 applies a version-matched patch set, backs up every file it replaces, and can roll back. It refuses to run against a version it has no patch set for, so an upgrade cannot silently land a build grafted onto the wrong release. Verified on 0.5.0: three PDFs import through mineru.net/api/v4 in ~288s, producing markdown, extracted images and per-directory summaries, with no failed files; existing indexed resources remain readable.
-Rollback restored only files recorded in the newest backup, so a file that was already patched when it was last written came back still patched: rolling back once did not undo the patch. That defeats the point of having a rollback, since it can leave a half-applied tree behind with no obvious sign. Walk every backup directory and keep the earliest copy of each file instead, which is the state from before this script touched anything.
The v4 precision API needs a bearer token and answers 401 without one, which
aborted the whole parse before any fallback could run: HTTPStatusError is not
ValueError or MinerUAPIError, so the existing except clauses let it through.
Treat 401/403 from the v4 submit as an endpoint rejection, then fall back to the
Agent API, which needs no credentials: POST /api/v1/agent/parse/file for a
signed url, PUT, then poll /api/v1/agent/parse/{task_id} for markdown_url. It
returns markdown at a CDN link rather than a zip and extracts no images, so the
result is used as-is with images_saved=0.
Also fixes a missing import that only the zip path hit: _mineru_unpack_zip calls
normalize_zip_filenames, which v0.5.0's pdf.py does not import because the
release has no MinerU code at all. Without it every v4 parse that got as far as
unpacking raised NameError.
Verified on 0.5.0 with the same three PDFs:
- with a token, v4 online-batch, 7-14s, 2-9 images each;
- without a token, agent-file, 7-9s, markdown only;
- through ov add-resource, 3/3 with no failed files, 21 artifacts and 6 images.
Author
|
Now, this PR has been compatible with v0.5 of OV. Please take a look. |
ov add-resource returns as soon as the task is queued, so waiting on its output
never ends and tells you nothing. Every wait here has a ceiling, and giving up
never cancels the task: the server owns it and may still finish.
Two bugs in the sync script came out of adding the waits:
- it recorded a project's fingerprint straight after submitting the import, so
a task that later failed still left the project marked synced and it was
never retried. Failed and timed-out projects now keep their old fingerprint;
- -Verify was declared but never used, so a dry run started importing
everything, which is how a "check" turned into a 15-minute import.
Wait-OvTask reports rather than kills on timeout, and Invoke-OvAddResource reads
the task result (root_uri, failed_files) instead of parsing submit output.
Verified: -Verify now returns immediately reporting 0 projects to reimport.
…archive Directory imports pack the tree into data/temp/upload/upload_<id>.zip. While one import is still running, the next one fails with [PERMISSION_DENIED] [WinError 32] another process is using this file which reads like a permissions fault and is not one: retrying just leaves more orphaned archives behind. Ten of them had piled up to 10 GB before this was noticed. Assert-OvIdle now checks for a running add_resource up front and exits with status 2 having submitted nothing. Also normalises -Project so it accepts comma-separated values. Start-Process -ArgumentList has no array syntax, so "A,B" arrived as one literal path named "A,B" and was reported as missing rather than as two projects.
OpenViking's memory extraction does not ask for a JSON array; the orchestrator in session/compressor_v3.py drives it through tool calling (memory.extract.candidates. tool_skill), and after four turns it raises "Final response could not be parsed as operations". Measured against the extraction prompt: model plain JSON tool calling tencent/Hunyuan-MT-7B 3/3, 3.3s HTTP 400 Qwen/Qwen3-8B 1/3, ~140s yes PaddlePaddle/PaddleOCR-VL-1.5 0/3 HTTP 400 deepseek-ai/DeepSeek-OCR 0/3 HTTP 400 Hunyuan wins the plain-JSON benchmark by a wide margin and is 40x faster, but it is a translation model and rejects tools outright, so it cannot drive the orchestrator. Qwen3-8B is the only candidate that satisfies both, and with it a real commit completes in 5.5s instead of failing after 100s. Window sizes, measured: DeepSeek-OCR 8192, PaddleOCR-VL-1.5 16384, Qwen/Qwen3.5-4B 262144. The 8192 window was the original blocker, since a session reached 9992 input tokens; max_tokens is now 4096. Embedding and rerank were benchmarked on retrieval quality rather than latency (recall@1 over six query/document pairs, three repeats). BAAI/bge-m3 and BAAI/bge-large-zh-v1.5 both reach recall@1 0.83 / MRR 0.92 at 1024 dims, so the current embedding model is kept. BAAI/bge-reranker-v2-m3 scores 0.83 and stays in place. One real finding: SiliconFlow rejects the `dimensions` field on /embeddings with 400 for bge-m3, so the vector length is whatever the model returns and `dimensions` must not be sent. ov.conf.production is the working configuration with credentials redacted.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Enhancement: Add
mineru-firststrategy for MinerU-priority fallbackMotivation
PR #4725 引入了
mineru_api_mode(auto/sync/async) 来支持 MinerU 的三种协议形态,非常棒的功能。但在实际部署中发现一个需求缺口:当前
auto策略的行为是 本地 pdfplumber 优先 → MinerU 回退。但对于需要高精度解析的场景(复杂版面、表格、公式),用户可能希望 MinerU 优先 → 本地 pdfplumber 回退。在本地测试中,MinerU 的解析质量明显优于 pdfplumber,尤其是:
因此建议增加
mineru-first策略。Changes
1.
openviking_cli/utils/config/parser_config.py2.
openviking/parse/parsers/pdf.py3.
openviking/service/core.pyStrategy Comparison
localmineruautomineru-firstConfiguration Example
{ "pdf": { "strategy": "mineru-first", "mineru_endpoint": "http://127.0.0.1:8899", "mineru_api_mode": "auto", "mineru_token": "your-online-api-token" } }Test Results (Local)
已验证
mineru-first策略在以下场景正常工作:mineru_api_mode: auto配合:自动检测协议形态(v1-sync/v2-tasks/online-batch)Summary
This enhancement adds
mineru-firstas a new strategy option that mirrors the existingautostrategy but with reversed priority. It's a small, self-contained change that gives users more flexibility in balancing quality vs. availability.