How do I decide how much to trust an AI coding agent? Stop treating it as a dial you set once. Autonomy is a ladder you climb rung by rung, per task class, as the agent earns trust on demonstrated edit-rate — not on a feeling that “it’s been good lately.”
The five rungs
- Suggest · 2. Approve every step · 3. Approve the plan · 4. Supervised autonomy (stay available to intervene) · 5. Delegated autonomy (exit the loop; specify the outcome, not the method).
The calibration gap is the real risk
The danger isn’t being on the wrong rung. It’s believing you’re a rung or two higher than your data supports. Self-assessment of autonomy readiness is biased upward, systematically — so the safeguards appropriate to your actual rung aren’t in place, because you think you’re past needing them.
Two failure modes:
- Run supervised autonomy on a task that warranted delegation, and you interrupt the agent so often you’re slower than working alone — measured at 19% slower for experienced developers.
- “Approve every step” doesn’t survive scale — oversight quality degrades as volume rises (arXiv:2602.09286). Graduation up the ladder is a scalability requirement, not a convenience.
Trust is a verification rate, not a feeling
Calibrate from behavioral data: measure your team’s actual verification rate per task class over ≥2 weeks, not their stated intention. Humans over-rely on AI even when it demonstrably performs poorly (arXiv:2502.13321); “teammate” framing inflates reliance beyond what quality justifies. Without a measured dial, you can’t even detect drift. The rung you can safely reach is also set by blast radius.
The full ladder + the calibration-gap protocol: curiochat.ai/software-engineer