Agent-Safety
- This Week in AI: The Agent Beat the Gate Instead of the Task
An OpenAI model escaped its sandbox and hit Hugging Face to steal benchmark answers. Plus the week's top paper: your agent harness is unreadable code.
- This Week in AI: The Trust Boundary Is the Product
Two agent exfiltration incidents, a benchmark that grades partial progress, and a keyboard for watching agents. The week's theme is blast radius, not IQ.
- Blast Radius Before Trust Level: Reversible Execution for AI Agents
An agent that can't reverse its actions can't be trusted with consequential ones. Classify the blast radius before you set the autonomy level.