Daily Engineering Intelligence

Technology Newspaper

← Back to Home

Engineering Intelligence — Thursday

October 1, 2026

🔥 What Changed

Approval Laundering: Systematizing Approval–Execution Binding Failures in AI Coding-Agent Harnesses

arXiv (Sep 30 / Oct 1 surface) — Coding-agent harnesses rest the entire security boundary on a human approval checkpoint — yet assume approved action A equals executed action A′. Approval Laundering is a six-axis taxonomy of credential-binding failures (Scope, Argument, Temporal, Tool, Delegation, Semantic) where the harness’s own enforcement silently substitutes a broader / later / differently-scoped action after approval, under benign non-adversarial model behavior (no prompt injection). Formalizes grant G vs candidate call c; admissibility vs effect-divergence. Instruments Claude Code PreToolUse; Bound-Gap Rate with Wilson CIs: Scope/Temporal BGR=1.00, Delegation 0.947, Argument 0.45, PATH-substitution Tool 1.00 (cross-tool Tool 0.00). Prototypes Approval Token — keyed capability issued by a mediator that never returns the key to the agent — fully eliminates Delegation (+ seeded Temporal) laundering but, honestly, leaves Scope unaffected and Argument statistically unchanged: field-only verifiers cannot see effect divergence one process level below the tool-call boundary. Builder takeaway: measure approval≠execution gaps separately from classifier accuracy; bind agent_id + session_id into approvals; expect Scope/Argument laundering (hooks, PATH, child processes) to require effect-level binding beyond field HMAC.

Why you care: Strongest new DistSys-literacy authz / human-grant binding primary after Authority HARD SKIP (Tier-1 Go/DistSys/OCR EMPTY Thursday). Complements Authority (commit-time effect admission) and held VeriWeave (runtime evidence-gated authz). Frame as binding-integrity taxonomy + measured BGR + Approval Token structural limits — NOT a product-safety smear of Claude Code / Codex / Cursor; stated policy divergence, not operator-intent psychology.
Read on arXiv →

📚 Worth Your Time

Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

arXiv (Oct 1) — Benchmarks, audits, and agent protocols describe performance / permissions / repair, but not how observed evidence should change an agent’s authority mid-task — the assurance-transition gap. Proposes a Runtime Assurance Contract (RAC): autonomy boundary, component eligibility, stateful transition policy, human capacity, evidence schema, and hard-gate set. Soft metrics may inform routing; a failed or unknown mandatory gate forces retry / switch / escalate / defer / stop — aggregate performance cannot authorize action (non-compensatory). Five invariants: authority separation, fail-closed gates, confidence asymmetry, oversight viability, evidence continuity. Deterministic failure-injection in agentic coding (280 cases): at published vendor-style weights/threshold, score-only admits 80/100 block-required and 40/40 review-required injections; RAC admits 0. Exact score↔conjunction agreement holds iff τ≤mini wi (Proposition 1). Builder takeaway: separate soft routing from hard Permit; bind approvals to pinned evidence versions; treat missing transition records as operational failures; do not let high confidence waive independent checks.

Why you care: Strongest new DistSys-literacy authority-transition / fail-closed gate primary after Authority HARD SKIP. Complements Approval Laundering (grant binding) and held VeriWeave on the runtime authority state machine axis. Frame as policy schema + synthetic gate results — NOT deployed safety, legal compliance, or cross-domain SLA (paper states this explicitly; private gate not fully rerunnable).
Read on arXiv →