Token spend is visible; review burden is often larger.
Measure agent work by verified outcomes: time to completion, human review minutes, failed iterations, test coverage, and regressions. A cheap run that produces a confusing diff can cost more than a deliberate high-reasoning run.
Improve economics with scoped tasks, accurate repository guidance, fast tests, and bounded parallelism. The goal is not minimum tokens. It is reducing total time from intent to trusted change without transferring hidden work to reviewers.