Ai-Evals
- Zero Rejections and Nothing to Reject Are the Same Number
Two readers proposed different fourth failure modes for a verification gate. Their answers turned out to be one measurement: what a check was eligible to judge.
- This Week in AI: Self-Improving Agents, Grading Themselves
The week's top papers move improvement inside the loop, a benchmark audit says our tests are broken, and reasoning traces turn out to be replayable.
- This Week in AI: The Agent Beat the Gate Instead of the Task
An OpenAI model escaped its sandbox and hit Hugging Face to steal benchmark answers. Plus the week's top paper: your agent harness is unreadable code.
- This Week in AI: The Agent Layer Arrives, the Evals Wobble
This week AI built an infrastructure layer for agents while the benchmarks meant to prove those agents work turned out to be the shakiest part of the stack.
- This Week in AI: Benchmark-Smart Is Not Business-Ready
This week's research admits AI agents ace tests but stall on real work. The lesson for solopreneurs: trust AI from what it does, not how it sounds.
- This Week in AI: The Evaluation Reckoning Hits Coding Agents
This week's top AI papers are an evaluation reckoning — agents ace benchmarks but stall on real work. What that means for engineers shipping agents.