Guide · 6 min read
The true cost of AI agent work: count the review, not just the tokens
An agent's invoice from the model provider is the smallest part of what its work costs. People brief it, review what it delivers, send it back and approve it. That time is real cost, and it rarely shows up anywhere. Count it, and you learn which agents pay for themselves and which only move work onto your best people.
Why tokens aren't the cost
In a 2026 field study of 802 developers, AI doubled the number of changes per engineer, and the work didn't disappear: review load per reviewer doubled too, and AI-written changes took about 20 % longer from first human review to merge. In an earlier randomized study, experienced developers believed AI made them 20 % faster while they were measured 19 % slower. Feeling productive and being productive are different numbers.
So the useful question isn't "what did the model cost?" but "what did it cost to get this task done, people included?"
The four parts of an agent's task
- Agent cost: model and tool spend the agent reports, converted to your currency
- Briefing: people's time on the task before the agent delivered (scoping, explaining, preparing)
- Review and rework: people's time after the agent delivered (checking, fixing, sending it back)
- Approval: who finally said it was done, a person or the agent itself
Measure it per task, then per agent
- Put agent work on tasks, not in loose chats, so time and cost have somewhere to land.
- Mark the hand-off: the moment the agent delivers (a finished request, a move to review, the time or cost it logs).
- Split people's time on the task into before (briefing) and after (review) that moment.
- Cost people's time at their hourly cost rate, not their billing rate.
- Add it up: full cost = agent cost + briefing + review. Divide by finished tasks for a cost per outcome.
- Watch two ratios: minutes of review per agent hour, and the share of tasks sent back.
What the numbers tell you
- High review per agent hour on one task type: the agent isn't ready for it, or the brief is too thin
- Many tasks sent back: tighten instructions or acceptance criteria before raising autonomy
- Agent closes its own tasks without a person: fine for low-stakes work, a risk for anything client-facing
- Negative margin per task: the client's price or the agent's scope is wrong
How Hourtick handles it
Reports → Agent outcomes lists every task with agent work that was done in the period: the agent's time and reported cost, people's briefing and review time, how often it was sent back, and who approved it. Admins also see the full cost per task at people's cost rates, the revenue and the margin. The same numbers are in the API (GET /api/v1/reports/agent-outcomes) and over MCP (get_agent_outcomes), so an agent can read its own track record.
To make sure the agent's part is right, connect Anthropic, OpenAI or OpenRouter in Settings → AI provider costs. Hourtick compares what each agent reported with what the provider billed, every month, fills in what's missing and flags double reporting.
Hourtick doesn't score people. It measures agents and tasks: the point is to decide which work to hand to an agent, not to rank colleagues.
Frequently asked questions
Isn't all human time on an agent's task review?
No. Time before the agent delivered is briefing, which you'd also spend on a person. Hourtick splits the two at the moment the agent first delivers.
What if people don't track review time?
Then review costs nothing on paper and agents look cheaper than they are. Starting a timer on the task while reviewing is enough.
Does this work for agents outside Hourtick?
Yes, as long as they log their time and cost on the task over MCP or the API. Hourtick doesn't run the models.
What if an agent reports its cost wrong?
Connect the provider (Anthropic, OpenAI or OpenRouter). Hourtick checks the reports against the bill each month: a higher bill fills in the difference, a lower one is flagged.
Sources
Related guides
Track your first hour in a minute.
Free for your whole team, forever. $29/month when you need 5 GB.