Daily Engineering Intelligence

Technology Newspaper

← Back to Home

Engineering Intelligence β€” Wednesday

September 30, 2026

πŸ”₯ What Changed

Xiaomi-OCR-0 β€” compact 0.8B unified OCR VLM + live Apache-2.0 weights

arXiv (SeerRay / Xiaomi; Sep 28) + Hugging Face (Sep 29) β€” Xiaomi-OCR-0, a unified 0.8B OCR VLM (Qwen3.5-0.8B base) trained on ~170M OCR-centric samples via progressive recipe: Q-Mask text anchoring β†’ multi-task continued pretraining β†’ Mix-RL with verifiable parsing/KIE/OCR-VQA rewards. Data engine: heterogeneous-expert Ensemble Triplet Consensus. Headline (two-stage with external PP-DocLayoutV3): OmniDocBench v1.6 Overall 96.83, Real5-OmniDocBench 95.24 (leads compared methods), Wild 87.94, OCR-understanding mean 83.2. Live ungated Apache-2.0 HF weights (SeerRay-Lab/Xiaomi-OCR-0) + GitHub + Space demo. Practical recipe: build parsing before mixing understanding; Mix-RL peaks higher than faster-converging MOPD. Caveats (authors’ Limitations): reported SOTA uses external layout detector; English-heavy + Chinese historical-script coverage stronger than Indic; in-house KIE address/invoice still weak.

Why you care: Sole keepable shippable Document AI primary on a Tier-1-empty Wednesday (Go/DistSys/HQ EMPTY; Authority HARD SKIP). Live Apache-2.0 weights beat pure arXiv agent papers as a What Changed signal. Frame as open compact OCR VLM + data/Mix-RL recipe β€” not an Indic/production SLA SOTA claim.
Read on arXiv β†’ Β· Hugging Face β†’ Β· GitHub β†’

AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems

arXiv (Sep 30) β€” Harness bugs (orchestration / tools / memory / provider) are a distinct failure class from LLM runtime failures, yet the only prior executable harness-bug bench held 43 manually reproduced bugs. AgentBug-Smith is a fully automated multi-agent pipeline: dual-stream repo+issue mining β†’ agent-aware Docker/env with mock provider reroutes β†’ fail-to-pass test generation. Beats SWE-Factory / SWE-bench-Live by 10.67–27.56 pp reproduction success; builds Live-Harness-Bench: 200 reproducible harness bugs (median repo 83k LoC; 31.5% tool-registry, 30.5% context/memory). Downstream: SOTA software agents correctly resolve ≀9% of the bench; distilling repair skills lifts mini-SWE-agent +6.32 pp on a repo-disjoint split. Explicit: 98.67% of auto-reproduced failures validated by manual inspection.

Why you care: Best new production-harness reliability primary β€” a continuously refreshed bug surface rather than another LLM-failure taxonomy. Complements Anatomy-of-Harness HARD SKIP. Frame as automated harness-bug repro + live bench + agent repair gap β€” not a claim that any specific agent product is unsafe.
Read on arXiv β†’