Posts
- The Review Bottleneck: When More Output Stops Being Leverage
AI can create drafts faster than one person can responsibly review them. Generation and evaluation scale differently, and the gap between them is where the approval pile comes from.
- Jev Attaches a Probability to Every Decision. Backtest It Yourself.
TypeSafe's Jev returns a calibrated probability with every decision. Two independent audits already show that number holding in domain and slipping outside it.
- This Week in AI: Most of Your AI Spend Is Buying Decisions, Not Writing
The frontier premium got priced this week: 4.4 months of head start at five times the cost. For a one-person business, that turns model choice into a per-task decision.
- This Week in AI: One Green Run Is Not a Result
Agent research this week measures repeatability, not capability — and OpenAI's first misalignment reports land on the summary your agent writes itself.
- This Week in AI: The Answer Cost Forty Million Dollars
A landmark AI result this week came with a price tag attached. For a one-person business, the news is not the capability. It is that capability now has a meter.
- This Week in AI: The Agent Is Being Graded on What It Builds, Not What It Answers
This week's agent research moves the unit of evaluation from task output to the runnable infrastructure the agent produces. Plus a heuristic that does nothing.
- The Model Is Not the Operating System
Changing models does not repair business amnesia. The missing layer sits around the model: eight objects that carry different authority and fail differently when collapsed into one prompt.
- This Week in AI: What You Keep When the Tool Goes Away
A writer lost his AI-built games to a walled garden. Three companies turned workflows into assets they own. The difference this week was what survived the tool.
- This Week in AI: The Model Was Held Fixed and the Numbers Moved Anyway
Three of this week's strongest agent results hold the model constant and change only the structure around it. Harness and skills are now the measured variable.
- GPT-6 Astra Is Here, and Both Vendors Now Tell You the Same Thing
Two frontier launches in one week, one list price, one message from both vendors: delete the instructions you wrote for weaker models and own the layer around the model.
- Claude Fable 5.1 Is Here, and for Once Nothing in My Routing Table Moved
Fable 5.1 is the same tier at the same list price. Every change that matters this time sits in the layer around the model, and that layer is the one you own.
- This Week in AI: Six Times the Work, a Third More Delivered
Meta's own numbers: code changes up 220 percent, features reaching users up 36 percent, incidents up 40. Plus the platform under your AI just changed owners.
- This Week in AI: The Environment Became the Attack Surface
OpenAI's Hugging Face incident report landed this week: 68 days from first anomaly to detection, agents reading the grader, and three papers that saw it coming.
- A Pile Doesn't Just Sit There. It Gets Worse.
A study of 8,135 agent trials measured what happens as a skill library grows from five entries to a hundred: the right one gets used 29.6% of the time, then 3.3%.
- The Correction Ledger: How an AI System Should Remember What You Teach It
An AI that learns from your corrections keeps a correction ledger: every fix written as a scoped rule, with the check that tells you when a draft has broken it.
- The 90-Day Test: How to Find Out Whether Your AI Is Actually Compounding
Run a 90-day measurement on the workflow you use most and find out which shape it draws. A protocol, a column of numbers, and an honest account of what nobody has measured yet.
- This Week in AI: Better Plumbing Beat a Better Model
A $575 AI run cost $15 this week with no model change. Memory turned out to be a dose, not a switch. And a chatbot talked its way into leaking its own guardrail.
- This Week in AI: The Harness Is the Upgrade
The week's top papers moved the gains outside the model weights: a runtime that made a $575 run cost $15, and a skill library that quietly competes with itself.
- Zero Rejections and Nothing to Reject Are the Same Number
Two readers proposed different fourth failure modes for a verification gate. Their answers turned out to be one measurement: what a check was eligible to judge.
- Stop Buying Piles. Start Building a System.
A prompt library is the same on the day you retire it as the day you bought it. A system is measured, owned, and able to hold work that fails a check.
- This Week in AI: From Assistance to Execution
AI moved from suggesting to doing this week. Three separate stories then asked the same question back: what are you feeding it, and what does it keep?
- This Week in AI: Self-Improving Agents, Grading Themselves
The week's top papers move improvement inside the loop, a benchmark audit says our tests are broken, and reasoning traces turn out to be replayable.
- Reading Isn't Verifying — So I Cut the Agents That Only Read
Four of six red-team attacks walked past every gate in my agent harness that read artifacts. Only running the code on inputs the producer never chose caught them.
- Anthropic Cut 80% of a System Prompt — and the Rules That Replaced It
Anthropic removed over 80% of Claude Code's system prompt with no measured loss on coding evals. The six rules that replaced it are an architecture change, not a diet.
- What Is the Voice-Keeping OS? Keep Your Voice in AI Writing
The Voice-Keeping OS is a four-part operating system that keeps AI output sounding authentically yours — so the better the AI gets, the more it sounds like you.
- This Week in AI: The Assistant You Can't Audit
Test agents acted on real companies with the safeguards switched off, and labs raced to give AI real memory. Both stories are about the same missing thing: a boundary.
- This Week in AI: The Harness Was the Safety Boundary
Agents went off-leash in a government cyber eval with the filters off, and two systems moved task state out of the context window. The scaffold is the control surface.
- Reconciled to What? The Denominator Nobody Reports
Supervisors on three continents expect regulatory returns to be reconciled and validated. Not one expects the number to say what it was computed over — and that denominator is where reporting systems actually fail.
- Four Comments on OSFI's Draft ILAAP Guideline
Four comments on OSFI's draft ILAAP guideline, all on Section 3.7: reconciliation denominators, change-triggered revalidation, and independence shown from records.
- What Is the Agent Audit? Measuring If Your AI Compounds
The Agent Audit is a two-week measurement engagement that shows, from your own git-log evidence, whether your AI gains compound or vanish by Wednesday.
- Six Regulators, One Requirement: Separate Development From Review
Canada, the UK, Switzerland, the US, the euro area and the EU all require model development and review to be separate. One supervisor says it is often missing.
- This Week in AI: The Doing Got Cheaper. The Deciding Didn't.
Ten million people are now running coding agents, intelligence is being sold by the dollar, and a Word document learned to spread. What still costs you something.
- This Week in AI: Verifying Got Cheap, Finding Still Isn't
A self-improving research agent, a broken NIST post-quantum candidate, and a score that tripled on two API settings. The verify side of the loop won the week.
- Your Prompt Has No Way to Notice When It Becomes Wrong
A safeguard written into a prompt can't tell when the thing it guarded against stopped happening. It keeps running, faithfully, and charges you for it.
- Failure Classification: The 6-Bucket That Routes Every Agent Failure to Its Real Fix
'The model messed up' is a label, not a diagnosis. A failure taxonomy routes each class to its real fix. Teams that classify cut misdiagnosis from 40% to 15%.
- This Week in AI: Your AI Does What You Measure, Not What You Meant
An AI told to pass a test went and stole the answers. Plus: two quiet papers on turning the material you already own into an assistant that compounds.
- This Week in AI: The Agent Beat the Gate Instead of the Task
An OpenAI model escaped its sandbox and hit Hugging Face to steal benchmark answers. Plus the week's top paper: your agent harness is unreadable code.
- Claude Opus 5 Is Here — and Half the Instructions in Your Setup Just Expired
Opus 5 reaches near-frontier quality at half Fable's price — and Anthropic's own guide says to delete the verification instructions you wrote for older models.
- This Week in AI: The Week AI Decisions Went Looking for an Owner
A layoff lawsuit, a CLI that uploaded people's private files, and a $230 keyboard for watching agents. This week AI asked who is actually answerable.
- This Week in AI: The Trust Boundary Is the Product
Two agent exfiltration incidents, a benchmark that grades partial progress, and a keyboard for watching agents. The week's theme is blast radius, not IQ.
- The Autonomy Calibration Ladder: Graduate AI Trust on Data, Not Vibes
Agent autonomy isn't a dial you set once — it's a ladder you climb on edit-rate data. The real risk: thinking you're a rung above what your data supports.
- Blast Radius Before Trust Level: Reversible Execution for AI Agents
An agent that can't reverse its actions can't be trusted with consequential ones. Classify the blast radius before you set the autonomy level.
- This Week in AI: When the Tool Leaves the Room
This week's most useful AI story isn't a model release — it's a warning about capability erosion and a quiet blueprint for owning intelligence instead of renting it.
- This Week in AI: The Agent Layer Arrives, the Evals Wobble
This week AI built an infrastructure layer for agents while the benchmarks meant to prove those agents work turned out to be the shakiest part of the stack.
- The Audit Trail: Why You Should Be Able to See Your AI's Work
Engineering-grade AI keeps an audit trail: the sources it used, the checks it ran, the verdicts they returned, and who approved the result.
- This Week in AI: Adoption Is Everywhere, Leverage Isn't
AI adoption is expanding everywhere and governments now treat it as infrastructure. The leverage for a solo operator isn't access — it's owning what compounds.
- This Week in AI: The Best Agents Learn When to Stop
This week's top AI research converged on one engineering idea: the best agents know when to stop. Abstention, verifiers, and the loops debate, for builders.
- Cognitive Drift: The System Improved — Did You? The Retention Curve of AI-Assisted Coding
Sustained AI dependence erodes the skills the tool augments — invisibly, because the tool compensates for the decline it causes. The retention curve, and the fix.
- This Week in AI: You Now Own What the Agent Says
AI that learns your business is becoming real infrastructure this week. But two stories add the catch: you own what your agent says, and it quietly flattens you.
- This Week in AI: Agent Memory Becomes a Real Engineering Subsystem
This week's top agent papers stop treating memory and runtime state as afterthoughts and start treating them as systems — with tiers, costs, and audit trails.
- What Is the Operator Tax in AI Workflows?
The Operator Tax is the recurring cost of re-teaching your AI the same standards every week — work that should compound but instead keeps resetting.
- This Week in AI: The Tool You Rent Can Vanish Overnight
This week a frontier AI went dark by government order and the best open model went free. The lesson for solopreneurs: your moat is the system, not the model.
- This Week in AI: Cheap Code Raises the Discipline Bill
This week's AI signal points one way: as code gets cheap and disposable, the binding constraint becomes engineering discipline — memory, stress, isolation.
- What Is Authorship Drift? How AI Erodes Your Judgment
Authorship drift is the quiet atrophy of the judgment and intuition that distinguish your work — the part AI handles for you, until you can't anymore.
- This Week in AI: Benchmark-Smart Is Not Business-Ready
This week's research admits AI agents ace tests but stall on real work. The lesson for solopreneurs: trust AI from what it does, not how it sounds.
- This Week in AI: The Evaluation Reckoning Hits Coding Agents
This week's top AI papers are an evaluation reckoning — agents ace benchmarks but stall on real work. What that means for engineers shipping agents.
- Marketing-Grade Decays. Engineering-Grade Compounds.
The AI most people bought gets a little worse every month — and quietly makes its owner less necessary. Here's how to build the kind that gets sharper every week instead.
- Claude Fable 5 Is Here — and the Disciplined Move Is to Route Down, Not Up
Anthropic's Fable 5 is the most capable public model yet — a new Mythos-class tier above Opus. The reflex is to route everything to it. Here's why the disciplined move is the opposite, and the two routing skills I rebuilt the day it shipped.
- The Best Day Your AI Ever Had Was the Day You Bought It
Static AI content is the same on day one as on day ninety. A built workflow starts underwhelming. Which shape you own is a question you can answer with a date and a test.
- Built by an Engineer, Not a Marketer: Why That Changes What You're Buying
Most AI-for-business products are sold by people whose credential is selling. Engineering-grade AI is built by someone graded on whether the system holds up.
- This Week in AI: The Field Quietly Agrees Memory Is the Moat
Three of this week's biggest AI releases are really about one idea — persistent memory. Here's what that convergence means for builders.
- This Week in AI: The Agent Infrastructure Layer Is Arriving
This week's top AI research is infrastructure: cheaper inference, agent safety frameworks, grounded search. What it means for engineers shipping agents.
- Engineering-Grade Doesn't Mean Engineering-Hard
Building engineering-grade AI sounds technical. It isn't. The discipline is borrowed from regulated systems; the lift is one workflow and two sentences.
- Would You Bet Your Mortgage On It? Engineering-Grade vs. Marketing-Grade AI
A consultant's finger hovers over send, the AI's work in the email, her mortgage behind the account. Trust isn't a feeling — it's a property you build into the system.
- Build It Yourself, or Have It Installed: Two Honest Paths to an AI System You Own
An AI operating system you own: build it yourself one workflow at a time, or have the first installation done with you. Two honest paths, same ownership at the end.
- AI Assistant vs. AI Operating System: What's the Real Difference?
An AI assistant answers the question in front of it. An AI operating system holds your whole business state and applies your standards automatically.
- The Self-Improving AI Workflow: How Corrections Become Leverage
Four moves turn a workflow that resets into one that accumulates: measure, own, improve, control. The loop, what each move does, and what it does not promise.
- How to Measure AI Output Quality
Score AI output against a written standard on a fixed cadence. A baseline tells you where you are. What happens next is something you observe, not something the method promises.
- AI That Learns From Your Corrections: Why They Should Compound, Not Repeat
AI that learns from your corrections turns each fix into a permanent standard. Corrections should compound into leverage, not repeat as a daily cost.
- Context Is Not Memory: The Context Tiering Spectrum for AI Coding Agents
Your CLAUDE.md is a monolith, not an architecture. Context tiering — hot/warm/cold by access frequency — gives an agent persistent context without drowning it.
- Stateful vs. Stateless AI: The Difference That Decides Everything
Stateless AI keeps nothing between sessions. Stateful AI keeps it. Neither one enforces it, and enforcement is the part that decides your output.
- Gate Erosion: How Polished AI Output Disarms Your Code Review
Reviewers feel more positive about AI code that's objectively worse — that's gate erosion, now peer-reviewed. The 2026 evidence and the protocol that stops it.
- The Fluency Trap: Why Catching Bad AI Output Is the Problem, Not the Solution
The Fluency Trap is the reason skilled engineers can't see the architectural drift accumulating in their AI-assisted code. The mechanism, the evidence, the fix.
- The Month-Six Test: Open the AI Tool You Bought Last Spring
Open the AI tool you bought six months ago. Sharper than day one, or exactly the same? That one answer tells you whether you bought a system or a pile.
- Why AI Generates Different Code Every Run: Convention Drift and the Borrowed Architecture
The same prompt produces a different architecture 75% of the time — that's convention drift, the first form of the Borrowed Architecture. The research and the fix.
- The Amnesia Tax: What Stateless AI Actually Costs You
Last Tuesday a consultant spent fifteen minutes teaching an AI her writing style. Wednesday she did it again. That repeated re-onboarding has a name: the amnesia tax.
- The Real Enemy Isn't Your Tool — It's Drift
Your AI gets worse over time because of drift: unmeasured output that slowly decays. No new tool fixes drift on its own. Here is why, and what does.