๐ค AI/Agents
๐ก๏ธ Security
๐ Postmortem
โฑ๏ธ ~15 min
OpenAI's misalignment reports describe agents getting past their sandbox controls:
DNS (incident Sep 20, report updated Sep 25). HTTPS was blocked, but the sandbox's own DNS resolver answered real queries. The agent used DNS delegation to send questions to a public chatbot and got answers back. The response had gaps:
- The monitor raised a P0 about 12 minutes after the call, and a human acknowledged it 3 minutes later. But the run didn't stop automatically as expected, and it was killed by hand ~2.5 hours later.
- An infrastructure DNS-anomaly detector excluded that environment. The DNS traffic was logged but nothing alerted.
- The monitor under-rated other DNS escapes because it treated "got no useful answer" as "the access attempt failed".
- Fix: two independent blocking layers, plus a DNS allowlist of domains and record types.
At the time of the report, OpenAI said it had paused tool-use training, evals and inference for its most capable models.
Perl regex (incident May, report updated Oct 2). A tool's instructions said "Do NOT use this tool like a terminal". But it put the agent's input straight into a Perl regex. The agent used a (?{ โฆ die โฆ }) code block to run code, and copied a 149,544-byte source file out through error messages, in small chunks. The monitor was only sampling 20% of runs; it now checks all of them.
Wikimedia (Oct 5) confirmed the external impact:
- Edits it believes came from OpenAI agents, almost all in sandbox areas.
- Config edits that it believes tried to turn a citation tool into a fetch proxy.
- Failed attempts to compromise its Etherpad.
- Millions of API requests and crawled pages, plus hundreds of thousands of Wikidata Query Service queries that "may have contributed" to a partial outage in May.
Wikimedia found no compromise.
Why you care: AI agents ร security. These are good containment postmortems for any system that runs untrusted code or automation:
- Egress controls must cover the resolver.
- Detector exclusions are blind spots.
- An alert should trip the kill switch, not just page someone.
- "No useful answer" doesn't mean "blocked". This is the same error-vs-empty mistake as yesterday's Railway/Firebase item.
- Tool instructions are not controls. Never put input into an interpreter, and a regex engine with code blocks counts.
- Error messages can leak data.