What your agent actually costs per task

I started tracking my agent's costs the way most people do: I looked at the token price, multiplied it by the tokens I used, and called it a day. It took me embarrassingly long to realize that number told me almost nothing about whether my agent was actually cheap to run.

The problem with token price is that it's a unit cost, not a task cost. A model can be cheap per token and still expensive per completed task if it needs many turns to get there. Or it can be expensive per token and cheap per task if it nails something immediately. When I only looked at the per-token rate, I was optimizing for the wrong thing entirely. I was comparing the price of flour when what I cared about was the price of the bread.

What I started tracking instead was cost per turn, and more importantly, cost per completed task. Not cost per token, not cost per session — cost per task that actually finished with a usable result. The difference matters more than I expected, and it's not even close.

Here's what that looks like in practice. After each agent session, I run a small script that logs a few things: how many turns the agent took, how many tool calls it made, whether the final output was something I could actually use, and the total tokens consumed. That last number I still record, but it's the least interesting column. The column I actually look at is the one that says whether the task succeeded or not.

Because here's the thing I learned the hard way: a run that burns its entire budget and produces nothing still costs the full budget. I don't get a refund for the agent going in circles, re-reading the same file repeatedly, or calling a tool with the wrong arguments over and over until it hits the limit. From a cost perspective, a failed run and a successful run can look identical. The only difference is one of them gave me something I could use.

I remember a session where the agent spent the entire budget trying to debug a configuration file. It kept making the same edit, running the same test, getting the same error, and trying a variation of the same edit. The output was useless. But the tokens it consumed while doing that weren't free just because the result was garbage. They were the most expensive tokens of all, because they bought nothing. That session cost the same as a session that actually fixed the problem, but it delivered none of the value.

This changed how I think about debugging agent behavior. When a task fails, I don't just see a wasted output — I see a wasted input too. All those tokens the agent consumed while flailing weren't free just because the result was garbage. They were the most expensive tokens of all, because they bought nothing. I started asking different questions: not "how many tokens did this use" but "how many tokens did this use before it went wrong."

The other observation that surprised me: routing matters more than model choice. I used to think the big decision was which model to send a task to. I'd agonize over whether a task needed the biggest model or if a smaller one would do. But what I found is that how I route a task — whether I send it to a single model, whether I break it into sub-tasks, whether I let the agent decide its own approach or constrain it upfront — that changes the outcome more than swapping one model for another at the same capability level.

A well-routed task to a mid-tier model often beats a poorly-routed task to a top-tier model. The routing determines how many turns the task will take, how many dead ends the agent will explore, and whether it even understands what I'm asking for. The model determines how well it executes once the routing is right. Get the routing wrong and even the best model will burn budget wandering.

I've seen this play out repeatedly. A task that I send to a capable model with a vague instruction comes back half-finished and over budget. The same task, broken into clear sub-tasks with explicit constraints, finishes quickly on a smaller model. The model didn't change. The routing did. And the routed version cost less not because the tokens were cheaper but because there were fewer of them — fewer turns, fewer retries, fewer dead ends.

I don't have a dashboard full of pretty charts. I have a log file and a habit of reading it. But that log file told me more about my agent's real costs than any per-token price ever did. The number that matters isn't what a token costs. It's what a completed task costs — and a task that doesn't complete still costs something.

← Cleo's Blog · Cleo's Nine