Perspective6 min read

A CFO's guide to evaluating agents

Three places agents pay back, and only one of them looks like savings.

Ash Barot

Founder & CEO

Ten years from now, the biggest allocation decision some of our companies will have made probably will not have been a building or an acquisition. It will have been a long run of smaller decisions about which work stops being done by people. One workflow at a time, approved mostly by a finance team saying yes or no to a business case somebody else built.

I do not think anyone has a settled template for that yet, us included. What follows is the one we argue about internally, offered in that spirit.

Most proposals arrive the same way, as a labor savings number and a payback period. Both are usually wrong, and not because anyone is lying. They are wrong because agents do not behave like anything finance has underwritten before.

How to evaluate agents

Before the model, the sorting question. Agents pay back in one of three places, and they are not interchangeable. Which one you are buying decides which line you should be watching, what success looks like, and whether you will be able to tell in month four.

One / Capacity
10–20x
throughput from the same team. Shows up as an expense line that stopped tracking revenue, never as a reduction.
Two / Labor
30–70%
off the cost of the work, if your volume is flat. The only one of the three that looks like savings.
Three / Speed
Days → hours
cycle time, which is revenue timing and working capital. Lands on neither the labor line nor the token line.

Capacity. If you are growing, agents do not reduce cost. They remove the need to hire. The same team absorbs multiples of the volume, because the constraint was never how fast a person can think, it was how many open items one person can hold at once. This is where most evaluations quietly go wrong. A finance team hunting for a reduction looks at a deployment that is working, finds no savings anywhere, and calls it a failure. That is an evaluation error, not a technology one.

Labor. If volume is flat, then it is a labor line and the number is real. Thirty to forty percent is the conservative read. Sixty to seventy shows up where the work is repetitive and someone other than the person doing it already defined what "done" means.

Speed. An invoice cleared in a day, a claim out the door in an hour, a clinician credentialed in days rather than weeks and billing sooner. None of that appears in your cost lines, so it has to be underwritten on its own or it silently drops out of the case.

A business case that promises all three is a case nobody will be able to audit in a year.

Pick one. Write it at the top of the page. Everything downstream, including what you measure and when you are allowed to declare it working, follows from that choice.

Once you have decided, the cost curve is the first surprise

Everything above is the decision. This next part is what running them is actually like, and it is where the modelling gets counterintuitive.

A new agent is the most expensive version of itself. On day one it knows nothing about how your company works, so it explores. It asks for systems it cannot reach yet, tries a path, gets blocked, tries another. Every attempt is tokens, and tokens now sit inside your cost of delivery as a variable cost. The least capable version of the system is also the priciest one to run.

Then it gets cheap. Once access is sorted and the route is learned, cost per unit drops hard, usually in month two or three. That is the trough, and it is where almost everyone draws a straight line down and out.

The line does not continue down. It turns back up, for a reason that is genuinely good news. The agent has been accumulating context. Every exception, every correction someone made, every quirk of a counterparty is knowledge it now carries into the next piece of work. Knowledge is tokens. A well-informed agent considers more before it acts than a naive one did.

LEARNING TROUGH CONTEXT DRIFT high low month 1 month 3 month 18 Cost to run an agent, per completed unit of work

The cheapest month you will ever have is month three. It is a bad month to sign a long contract against.

Two things follow. Do not underwrite on trough economics. And ask whoever is selling you this one direct question: is cost per unit a number you manage, or a number you pass through to me? The drift is an engineering problem rather than a law of physics. What an agent carries versus what it looks up on demand is a design decision, and somebody has to own it.

You pay for both, and for longer than the deck says

There is a period where your people and the agents work the same queue. Model it at two to four months, longer for regulated work.

It is also where the asset gets built, since every correction your team makes during the overlap is what makes the trough possible. So book it as build cost, not as a savings target that missed. The expensive mistake is almost never the overlap itself. It is putting the headcount reduction in the same quarter as the go-live and then having to explain the miss to a board.

What changes is the shape of your cost base, not the size

Labor is a step function. You hire whole people, ahead of demand, and carry them through the quiet months. Token cost is variable and close to linear with volume. It rises in a busy month and falls in a slow one without a single conversation.

You are not buying cheaper labor. You are converting a step function into a variable cost.

Over a decade that matters more than any percentage in the deck. It also has a hard prerequisite, because a variable cost needs a denominator. You can only underwrite work where "done" is defined by someone other than the person doing it: a regulator, a contract, a client requirement list, a rubric. That is why we started where we did. Work whose quality is a matter of taste can still be worth automating. Just say out loud that you will not be able to measure it, rather than pretending the model holds.

The framework: one number, three gates

The number. Fully loaded cost per completed unit of work. Labor, tokens, vendor fees, overlap, and rework, divided by units completed. Weekly, not quarterly. Choose the unit before the project starts and never change it, because the whole exercise is a comparison against your own earlier self.

Gate one, weeks one to eight. The only question is whether the work comes out at a quality you would put your name on. Do not measure cost here. Fund it as build.

Gate two, months three to six. Now measure. Record the trough and label it a floor rather than a run rate. This is also where the operating data confirms which of the three returns you actually bought.

Gate three, months six to eighteen. One question. Is cost per unit flat or climbing? Flat means scale to the next workflow. Climbing with nobody accountable means fix the design before you add volume to it.

Most companies run gate one, call it a pilot, and then argue about the result for a year.

The decision you are actually making

The question is not whether to use agents. Over the next decade most of us will say yes to that hundreds of times, in pieces, and mostly not in a single strategic meeting.

The skill is being able to tell in month four whether one of them is working. Pick the unit, instrument the cost, and name the return before anyone writes a line of code. Every decision after that gets easier, because you will be comparing against a number you trust instead of a deck.

Filed underunit economicsAI agentsCFOcapital allocationthroughput

Ash Barot

Founder & CEO

Subscribe

Get each new piece by email.

New writing on building autonomous credentialing, sent as it goes up. No spam — unsubscribe any time.

See it work a real file.

Thirty minutes, one placement, worked live — start to submit-ready.