Engineering practice

How we cut an AI agent's running cost by splitting the work between two models

We helped a tax accountant in India put an AI agent into their filing work. It worked on day one, but at that cost it would not have lasted a filing season. Here is what we changed, what each phase cost, and the open-source toolkit that came out of it.

By Ashish Tonse 9 min read

Build phase
~$400
First two returns
~$50
Projected per client
~$1

We recently helped a tax accountant in India put an AI agent to work in their return-filing process. The agent reads government tax statements, reconciles credits for tax deducted at source, and builds the computation sheets behind each filing. Two machines run Claude Code. Ours runs a large frontier model and builds the tools. The accountant's machine runs a small model, billed per token, and does the day-to-day filing. The accountant is not an engineer.

The agent worked on the first day. On that same first day, it was on course to cost too much to last a filing season. The same work is now on track to cost about a dollar per client, on the same small model. A better prompt did not fix it. We changed which model does the thinking, and when. I call the pattern Papa Claude and Baby Claude.

First make it correct, then make it cheap

As with any process, you get it right before you make it fast. The first phase costs money on purpose. A frontier model read the tax code, the schema the government publishes for electronic returns, and last year's accepted return. Then it wrote code that checks its output against all three. Getting one real return correct from end to end is worth the tokens. Measured afterwards, the build cost about $400 at API prices. Our flat-rate subscription hid that number at the time.

The bill that did arrive was for daily use: nearly $50 of per-token sessions on the accountant's machine for the first two returns. The system was correct, but too expensive to run. So we did the usual engineering thing and optimized what already worked.

What each phase cost A bar chart at API prices, on a scale from zero to 400 dollars. Building the tools with the large model cost about 400 dollars, once. The first two returns on the small model cost nearly 50 dollars, and four long sessions caused about 95 percent of that. With the changes, each client is projected to cost about 1 dollar on the same small model. That last figure is a projection from the design, not yet a measurement. WHAT EACH PHASE COST, AT API PRICES $0 $100 $200 $300 $400 Building the tools LARGE MODEL, ONCE ~$400 First two returns SMALL MODEL, PER TOKEN FOUR LONG SESSIONS: ABOUT 95 PERCENT ~$50 Each client, projected SAME SMALL MODEL ~$1 A PROJECTION FROM THE DESIGN, NOT YET MEASURED
Building the tools cost about $400, once. The first two returns cost nearly $50 on the small model. With the changes, each client is projected to cost about $1 on the same small model. The first two figures are measured, and the third is a projection from the design.

Where the money went

When the bill jumped, I pulled the session transcripts, removed duplicates and added up the cost. The first few days came to nearly $50, and four long sessions caused about 95 percent of it. Each of those sessions had built up 110,000 to 145,000 tokens of context. They ran from 150 to 750 requests each, and every request re-read that context from the start. The most expensive session cost well over ten dollars, almost all of it spent re-reading history it had already paid for.

Every request re-read the whole session Where the first few days of per-token sessions went, nearly 50 dollars in total. Four long sessions caused about 95 percent of it, and everything else caused the rest. Inside one of those sessions, each request re-reads everything that came before it, plus the little that is new. The first request reads a little, the second a little more, and so on, until the last request, after 150 to 750 of them, re-reads 110,000 to 145,000 tokens of context from the start. Most of what each request pays for is history it has already paid for. THE FIRST FEW DAYS: NEARLY $50 OF PER-TOKEN SESSIONS Four long sessions: about 95 percent EVERYTHING ELSE INSIDE ONE OF THOSE SESSIONS: WHAT EACH REQUEST READS Request 1 Request 2 Request 3 Request 4 The last request OF 150 TO 750 110,000 TO 145,000 TOKENS, READ AGAIN FROM THE START HISTORY IT ALREADY PAID FOR, READ AGAIN NEW IN THIS REQUEST
Four long sessions caused about 95 percent of the first bill. In each one, every request re-read the whole session from the start, so by the end each request was paying again for 110,000 to 145,000 tokens of history.

What was the model doing in those sessions? Reading 30-page government PDFs straight into its context. Patching a document parser by hand, three separate times. And once, retrying an unexpected capital-gains case again and again, paying for the full context each time. In other words, it was thinking. Thinking is what a small model does worst, and it gets expensive because the model does not know it is failing.

The model was not bad at tax. The setup was expensive to run.

Split the work between two models

Papa Claude is the maintainer: the large model, on our machine, with an engineer watching. It writes tested tools that give the same result every time and check their own output. It wrote parsers for every document the government produces. It wrote an extraction step that turns a folder of portal downloads into a draft data file and a short list of the decisions a person needs to make. It wrote one command that checks a finished return in several independent ways. Use the most expensive model to build the tools, and keep it out of the daily filing.

Baby Claude is the operator: the small model, on the accountant's machine. It never opens a document. It reads the short summaries, resolves the listed questions and runs commands. The handbook in the repository states the rule: the logic lives in scripts, and the model does not reason its way through a filing. If the small model is reading a document, a tool is missing.

Papa Claude builds the tools, Baby Claude runs them On our machine, a large model with an engineer watching writes tested tools: document parsers, an extraction step and one command that checks a finished return. Those tools run on the accountant's machine, where a small model billed per token reads short summaries, resolves the listed questions and runs commands, and never opens a document. When the tools cannot do something, the small model adds an open entry to a file of missing capabilities in git. The large model builds the capability and marks the entry done. OUR MACHINE, AN ENGINEER WATCHING Papa Claude, the maintainer LARGE FRONTIER MODEL Writes tested tools PARSERS, AN EXTRACTION STEP, ONE CHECK COMMAND EACH ONE CHECKS ITS OWN OUTPUT TESTED TOOLS THE ACCOUNTANT'S MACHINE Baby Claude, the operator SMALL MODEL, BILLED PER TOKEN Reads summaries, runs commands RESOLVES THE LISTED QUESTIONS NEVER OPENS A DOCUMENT ADDS AN OPEN ENTRY BUILDS IT, MARKS IT DONE File of missing capabilities A MARKDOWN FILE IN GIT NO QUEUE SERVICE, NO FRAMEWORK
The large model writes tested tools on our machine, and the small model on the accountant's machine only runs them. When a tool is missing, the small model writes it down in a file in git, and the large model builds it.

Why tax is a good place to start

This pattern does not suit every process. Tax has two features that make it a strong fit, and both make the output easy to check.

First, the rules are public and precise. The tax code, the schema for electronic returns and the validation rules the processing center applies are all documented. A capable model is very good at reading those documents, turning them into code, and checking its code and output against them. We do not rely on the model to know tax. We rely on it to check its work against the documents that define the rules.

Second, there is a known correct answer to test against: last year's return, which the tax authority already accepted. The maintainer uses the prior filing as a template and the published rules as the test, and produces nothing that fails either one. A field with good documentation and a history you can check is where this kind of AI works well, because you never depend on the model's confidence.

Two things every return is checked against Why tax suits this approach. This year's return, built by the tools, is checked against two things. The first is the test: the published rules, meaning the tax code, the schema for electronic returns and the validation rules the processing center applies. The second is the template: last year's return, which the tax authority already accepted. The tools produce nothing that fails either one, so the result never depends on the model's confidence. WHY TAX IS EASY TO CHECK THE TEST The published rules THE TAX CODE, THE E-RETURN SCHEMA, THE PROCESSING CENTER'S VALIDATION RULES THE TEMPLATE Last year's return ALREADY ACCEPTED BY THE TAX AUTHORITY: A KNOWN CORRECT ANSWER BUILT BY THE TOOLS This year's return CHECKED AGAINST BOTH THE RULE Nothing that fails either one NEVER THE MODEL'S CONFIDENCE WE DO NOT RELY ON THE MODEL TO KNOW TAX. WE RELY ON IT TO CHECK ITS WORK AGAINST THE DOCUMENTS THAT DEFINE THE RULES.
Every return is checked against the published rules and against last year's accepted return. Nothing that fails either one is produced.

Seven techniques that keep the cost down

Use git to pass requests between the models

When the small model meets something the tools cannot do, it does not improvise tax law. It adds an OPEN entry to a markdown file of missing capabilities, commits, pushes and sets that filing aside. On our machine, the large model pulls the change, builds the missing capability with tests and marks the entry done. The small model's job is to report the gap. The large model fills it. There is no queue service and no orchestration framework, only a text file in git.

Stop after two failures

Every error message ends with exactly one next step. A rule of two failures, then set the filing aside, stops retry loops. When a check fails, the message says, word for word, do not retry or work around this. Small models follow clear instructions about failure very well. Given a vague error, they retry, and every retry re-reads the whole context. So a vague error message costs money.

Block the next step until the questions are answered

The extraction step writes its open questions into the data file it produces, and every later step refuses to run while any remain. You cannot file a return over a placeholder. The tools make the careless path impossible, so we do not have to rely on the model being careful.

Two rules that close off the expensive path Two rules the tools enforce on the small model. First, block the next step until the questions are answered: the extraction step writes a draft data file with its open questions inside it, every later step refuses to run while any question remains, so a return can never be filed over a placeholder. Second, stop after two failures: a failed check ends with exactly one next step and says not to retry or work around it, and after a second failure the filing is set aside. Without that rule, a vague error leads to retries, and every retry re-reads the whole context. BLOCK THE NEXT STEP UNTIL THE QUESTIONS ARE ANSWERED Extraction step READS THE PORTAL DOWNLOADS Draft data file OPEN QUESTIONS WRITTEN INSIDE IT Every later step REFUSES TO RUN WHILE ANY QUESTION REMAINS Filing NEVER OVER A PLACEHOLDER STOP AFTER TWO FAILURES A check fails FIRST FAILURE Exactly one next step "DO NOT RETRY OR WORK AROUND THIS" It fails again SECOND FAILURE Set the filing aside NO THIRD ATTEMPT WITHOUT THIS RULE: A VAGUE ERROR, A RETRY, AND THE WHOLE CONTEXT READ AGAIN, EVERY TIME.
Open questions lock every later step, and a second failure sets the filing aside. Both rules close off the paths where a small model would otherwise think, retry and re-read.

Keep what next year needs in a small file

Every verified return writes a small carry-forward file: the computed totals, losses carried forward, and the few numbers next year's return needs from this year's. That is a few hundred bytes of JSON. Without it, the model would re-read 100,000 tokens of context.

Ask for data in a machine-readable format

We spent three rounds patching a parser for PDF layouts before anyone asked whether the government portal offered the same data as JSON. It did, in the same download dialog, one extra click away. The best way to save tokens can be to change the input before the model ever sees it.

Remove warnings that do not matter

The return generator used to print 74 standard warnings on every run. When we collapsed them into one summary line, the 5 real warnings hidden among them became visible, and the small model stopped spending tokens on problems that did not exist.

Seventy-four standard warnings, collapsed to one line The return generator's output on every run, before and after. Before, it printed 74 standard warnings, and 5 real warnings were hidden among them, drawn here as 5 highlighted marks among 74 plain ones. After, the 74 standard warnings are collapsed into one summary line, and the 5 real warnings stand on their own lines where they are easy to see. The small model stopped spending tokens on problems that did not exist. EVERY RUN, BEFORE 74 STANDARD WARNINGS, 5 REAL ONES AMONG THEM EVERY RUN, AFTER 74 STANDARD WARNINGS: ONE SUMMARY LINE REAL WARNING REAL WARNING REAL WARNING REAL WARNING REAL WARNING THE 5 REAL WARNINGS, NOW VISIBLE THE SMALL MODEL STOPPED SPENDING TOKENS ON PROBLEMS THAT DID NOT EXIST.
Before, 5 real warnings sat among 74 standard ones on every run. After, the 74 are one summary line and the 5 are easy to see.

Measure before you optimize

One fix looked obvious: set a cheaper model in the environment config. It quietly broke prompt caching. The same request then cost 14 times as much, about 25 cents against under 2 cents. We caught it only because we ran both settings live before rolling out the change.

One config change, fourteen times the cost The same request under two settings, drawn in blocks, where each block is what the request cost with prompt caching working. Before the change, with caching working, the request cost under 2 cents: one block. After a cheaper model was set in the environment config, caching quietly broke and the same request cost about 25 cents: fourteen blocks, 14 times as much. It was caught only because both settings were run live before the change rolled out. ONE REQUEST, TWO SETTINGS EACH BLOCK: WHAT IT COST WITH CACHING WORKING Caching working BEFORE THE CHANGE Under 2 cents Caching broken CHEAPER MODEL SET IN CONFIG 14 times as much: about 25 cents CAUGHT ONLY BECAUSE BOTH SETTINGS RAN LIVE BEFORE THE CHANGE ROLLED OUT.
Setting a cheaper model in the environment config broke prompt caching, and the same request went from under 2 cents to about 25 cents.

The same approach works outside tax

The split between a maintainer and an operator works for any team putting agents in front of people who are not engineers: operations, finance, support, compliance. One expensive model, supervised by an engineer, writes tested tools that check their own output. Many cheap sessions read the summaries and run commands. A shared file sits between them, so every request for a new capability is written down, and the small model never has to judge it in the middle of a task.

Do the expensive thinking once, when you build the tools. Daily use should be routine.

All of it is public at github.com/kznconsulting/india-itr-toolkit: the parsers, the checks, the file of missing capabilities, and the handbook the small model works from. Client folders never go into git, so the repository holds the tools and nobody's tax data. The results of the build phase are in there too, so if you clone it, you start at the cheap phase.

About the numbers

We measured the build and first-returns figures from session transcripts, priced at API rates. The one dollar per client is a projection from the design, and we will label it that way until a full client has gone through the pipeline.

Putting AI in front of people who are not engineers?

We are happy to talk it through. Bring the workflow and the questions your team has about it.

How we implement AI

Book a 30-minute call