Ai-Agents
- This Week in AI: The Harness Is the Upgrade
The week's top papers moved the gains outside the model weights: a runtime that made a $575 run cost $15, and a skill library that quietly competes with itself.
- This Week in AI: Self-Improving Agents, Grading Themselves
The week's top papers move improvement inside the loop, a benchmark audit says our tests are broken, and reasoning traces turn out to be replayable.
- This Week in AI: The Harness Was the Safety Boundary
Agents went off-leash in a government cyber eval with the filters off, and two systems moved task state out of the context window. The scaffold is the control surface.
- This Week in AI: The Trust Boundary Is the Product
Two agent exfiltration incidents, a benchmark that grades partial progress, and a keyboard for watching agents. The week's theme is blast radius, not IQ.
- This Week in AI: The Agent Layer Arrives, the Evals Wobble
This week AI built an infrastructure layer for agents while the benchmarks meant to prove those agents work turned out to be the shakiest part of the stack.
- This Week in AI: The Best Agents Learn When to Stop
This week's top AI research converged on one engineering idea: the best agents know when to stop. Abstention, verifiers, and the loops debate, for builders.
- This Week in AI: Agent Memory Becomes a Real Engineering Subsystem
This week's top agent papers stop treating memory and runtime state as afterthoughts and start treating them as systems — with tiers, costs, and audit trails.
- This Week in AI: Cheap Code Raises the Discipline Bill
This week's AI signal points one way: as code gets cheap and disposable, the binding constraint becomes engineering discipline — memory, stress, isolation.
- This Week in AI: The Evaluation Reckoning Hits Coding Agents
This week's top AI papers are an evaluation reckoning — agents ace benchmarks but stall on real work. What that means for engineers shipping agents.
- This Week in AI: The Agent Infrastructure Layer Is Arriving
This week's top AI research is infrastructure: cheaper inference, agent safety frameworks, grounded search. What it means for engineers shipping agents.