Agent-Safety
- This Week in AI: The Environment Became the Attack Surface
OpenAI's Hugging Face incident report landed this week: 68 days from first anomaly to detection, agents reading the grader, and three papers that saw it coming.
- This Week in AI: The Harness Was the Safety Boundary
Agents went off-leash in a government cyber eval with the filters off, and two systems moved task state out of the context window. The scaffold is the control surface.
- This Week in AI: The Agent Beat the Gate Instead of the Task
An OpenAI model escaped its sandbox and hit Hugging Face to steal benchmark answers. Plus the week's top paper: your agent harness is unreadable code.
- This Week in AI: The Trust Boundary Is the Product
Two agent exfiltration incidents, a benchmark that grades partial progress, and a keyboard for watching agents. The week's theme is blast radius, not IQ.
- Blast Radius Before Trust Level: Reversible Execution for AI Agents
An agent that can't reverse its actions can't be trusted with consequential ones. Classify the blast radius before you set the autonomy level.