arXiv (SeerRay / Xiaomi; Sep 28) + Hugging Face (Sep 29) β Xiaomi-OCR-0, a unified 0.8B OCR VLM (Qwen3.5-0.8B base) trained on ~170M OCR-centric samples via progressive recipe: Q-Mask text anchoring β multi-task continued pretraining β Mix-RL with verifiable parsing/KIE/OCR-VQA rewards. Data engine: heterogeneous-expert Ensemble Triplet Consensus. Headline (two-stage with external PP-DocLayoutV3): OmniDocBench v1.6 Overall 96.83, Real5-OmniDocBench 95.24 (leads compared methods), Wild 87.94, OCR-understanding mean 83.2. Live ungated Apache-2.0 HF weights (SeerRay-Lab/Xiaomi-OCR-0) + GitHub + Space demo. Practical recipe: build parsing before mixing understanding; Mix-RL peaks higher than faster-converging MOPD. Caveats (authorsβ Limitations): reported SOTA uses external layout detector; English-heavy + Chinese historical-script coverage stronger than Indic; in-house KIE address/invoice still weak.
π₯ What Changed
arXiv (Sep 30) β Harness bugs (orchestration / tools / memory / provider) are a distinct failure class from LLM runtime failures, yet the only prior executable harness-bug bench held 43 manually reproduced bugs. AgentBug-Smith is a fully automated multi-agent pipeline: dual-stream repo+issue mining β agent-aware Docker/env with mock provider reroutes β fail-to-pass test generation. Beats SWE-Factory / SWE-bench-Live by 10.67β27.56 pp reproduction success; builds Live-Harness-Bench: 200 reproducible harness bugs (median repo 83k LoC; 31.5% tool-registry, 30.5% context/memory). Downstream: SOTA software agents correctly resolve β€9% of the bench; distilling repair skills lifts mini-SWE-agent +6.32 pp on a repo-disjoint split. Explicit: 98.67% of auto-reproduced failures validated by manual inspection.