Agent-Reliability
- This Week in AI: One Green Run Is Not a Result
Agent research this week measures repeatability, not capability — and OpenAI's first misalignment reports land on the summary your agent writes itself.
- Zero Rejections and Nothing to Reject Are the Same Number
Two readers proposed different fourth failure modes for a verification gate. Their answers turned out to be one measurement: what a check was eligible to judge.
- Reading Isn't Verifying — So I Cut the Agents That Only Read
Four of six red-team attacks walked past every gate in my agent harness that read artifacts. Only running the code on inputs the producer never chose caught them.
- Your Prompt Has No Way to Notice When It Becomes Wrong
A safeguard written into a prompt can't tell when the thing it guarded against stopped happening. It keeps running, faithfully, and charges you for it.
- The Autonomy Calibration Ladder: Graduate AI Trust on Data, Not Vibes
Agent autonomy isn't a dial you set once — it's a ladder you climb on edit-rate data. The real risk: thinking you're a rung above what your data supports.
- This Week in AI: The Best Agents Learn When to Stop
This week's top AI research converged on one engineering idea: the best agents know when to stop. Abstention, verifiers, and the loops debate, for builders.