How do I decide how much to trust an AI coding agent? Stop treating it as a dial you set once. Autonomy is a ladder you climb rung by rung, per task class, as the agent earns trust on demonstrated edit-rate — not on a feeling that “it’s been good lately.”

The five rungs

  1. Suggest · 2. Approve every step · 3. Approve the plan · 4. Supervised autonomy (stay available to intervene) · 5. Delegated autonomy (exit the loop; specify the outcome, not the method).

The calibration gap is the real risk

The danger isn’t being on the wrong rung. It’s believing you’re a rung or two higher than your data supports. Self-assessment of autonomy readiness is biased upward, systematically — so the safeguards appropriate to your actual rung aren’t in place, because you think you’re past needing them.

Two failure modes:

  • Run supervised autonomy on a task that warranted delegation, and you interrupt the agent so often you’re slower than working alone — measured at 19% slower for experienced developers.
  • “Approve every step” doesn’t survive scale — oversight quality degrades as volume rises (arXiv:2602.09286). Graduation up the ladder is a scalability requirement, not a convenience.

Trust is a verification rate, not a feeling

Calibrate from behavioral data: measure your team’s actual verification rate per task class over ≥2 weeks, not their stated intention. Humans over-rely on AI even when it demonstrably performs poorly (arXiv:2502.13321); “teammate” framing inflates reliance beyond what quality justifies. Without a measured dial, you can’t even detect drift. The rung you can safely reach is also set by blast radius.

The full ladder + the calibration-gap protocol: curiochat.ai/software-engineer