We recently helped a tax accountant in India put an AI agent to work in their return-filing process. The agent reads government tax statements, reconciles credits for tax deducted at source, and builds the computation sheets behind each filing. Two machines run Claude Code. Ours runs a large frontier model and builds the tools. The accountant's machine runs a small model, billed per token, and does the day-to-day filing. The accountant is not an engineer.
The agent worked on the first day. On that same first day, it was on course to cost too much to last a filing season. The same work is now on track to cost about a dollar per client, on the same small model. A better prompt did not fix it. We changed which model does the thinking, and when. I call the pattern Papa Claude and Baby Claude.
First make it correct, then make it cheap
As with any process, you get it right before you make it fast. The first phase costs money on purpose. A frontier model read the tax code, the schema the government publishes for electronic returns, and last year's accepted return. Then it wrote code that checks its output against all three. Getting one real return correct from end to end is worth the tokens. Measured afterwards, the build cost about $400 at API prices. Our flat-rate subscription hid that number at the time.
The bill that did arrive was for daily use: nearly $50 of per-token sessions on the accountant's machine for the first two returns. The system was correct, but too expensive to run. So we did the usual engineering thing and optimized what already worked.
Where the money went
When the bill jumped, I pulled the session transcripts, removed duplicates and added up the cost. The first few days came to nearly $50, and four long sessions caused about 95 percent of it. Each of those sessions had built up 110,000 to 145,000 tokens of context. They ran from 150 to 750 requests each, and every request re-read that context from the start. The most expensive session cost well over ten dollars, almost all of it spent re-reading history it had already paid for.
What was the model doing in those sessions? Reading 30-page government PDFs straight into its context. Patching a document parser by hand, three separate times. And once, retrying an unexpected capital-gains case again and again, paying for the full context each time. In other words, it was thinking. Thinking is what a small model does worst, and it gets expensive because the model does not know it is failing.
The model was not bad at tax. The setup was expensive to run.
Split the work between two models
Papa Claude is the maintainer: the large model, on our machine, with an engineer watching. It writes tested tools that give the same result every time and check their own output. It wrote parsers for every document the government produces. It wrote an extraction step that turns a folder of portal downloads into a draft data file and a short list of the decisions a person needs to make. It wrote one command that checks a finished return in several independent ways. Use the most expensive model to build the tools, and keep it out of the daily filing.
Baby Claude is the operator: the small model, on the accountant's machine. It never opens a document. It reads the short summaries, resolves the listed questions and runs commands. The handbook in the repository states the rule: the logic lives in scripts, and the model does not reason its way through a filing. If the small model is reading a document, a tool is missing.
Why tax is a good place to start
This pattern does not suit every process. Tax has two features that make it a strong fit, and both make the output easy to check.
First, the rules are public and precise. The tax code, the schema for electronic returns and the validation rules the processing center applies are all documented. A capable model is very good at reading those documents, turning them into code, and checking its code and output against them. We do not rely on the model to know tax. We rely on it to check its work against the documents that define the rules.
Second, there is a known correct answer to test against: last year's return, which the tax authority already accepted. The maintainer uses the prior filing as a template and the published rules as the test, and produces nothing that fails either one. A field with good documentation and a history you can check is where this kind of AI works well, because you never depend on the model's confidence.
Seven techniques that keep the cost down
Use git to pass requests between the models
When the small model meets something the tools cannot do, it does not improvise tax law. It adds an OPEN entry to a markdown file of missing capabilities, commits, pushes and sets that filing aside. On our machine, the large model pulls the change, builds the missing capability with tests and marks the entry done. The small model's job is to report the gap. The large model fills it. There is no queue service and no orchestration framework, only a text file in git.
Stop after two failures
Every error message ends with exactly one next step. A rule of two failures, then set the filing aside, stops retry loops. When a check fails, the message says, word for word, do not retry or work around this. Small models follow clear instructions about failure very well. Given a vague error, they retry, and every retry re-reads the whole context. So a vague error message costs money.
Block the next step until the questions are answered
The extraction step writes its open questions into the data file it produces, and every later step refuses to run while any remain. You cannot file a return over a placeholder. The tools make the careless path impossible, so we do not have to rely on the model being careful.
Keep what next year needs in a small file
Every verified return writes a small carry-forward file: the computed totals, losses carried forward, and the few numbers next year's return needs from this year's. That is a few hundred bytes of JSON. Without it, the model would re-read 100,000 tokens of context.
Ask for data in a machine-readable format
We spent three rounds patching a parser for PDF layouts before anyone asked whether the government portal offered the same data as JSON. It did, in the same download dialog, one extra click away. The best way to save tokens can be to change the input before the model ever sees it.
Remove warnings that do not matter
The return generator used to print 74 standard warnings on every run. When we collapsed them into one summary line, the 5 real warnings hidden among them became visible, and the small model stopped spending tokens on problems that did not exist.
Measure before you optimize
One fix looked obvious: set a cheaper model in the environment config. It quietly broke prompt caching. The same request then cost 14 times as much, about 25 cents against under 2 cents. We caught it only because we ran both settings live before rolling out the change.
The same approach works outside tax
The split between a maintainer and an operator works for any team putting agents in front of people who are not engineers: operations, finance, support, compliance. One expensive model, supervised by an engineer, writes tested tools that check their own output. Many cheap sessions read the summaries and run commands. A shared file sits between them, so every request for a new capability is written down, and the small model never has to judge it in the middle of a task.
Do the expensive thinking once, when you build the tools. Daily use should be routine.
All of it is public at github.com/kznconsulting/india-itr-toolkit: the parsers, the checks, the file of missing capabilities, and the handbook the small model works from. Client folders never go into git, so the repository holds the tools and nobody's tax data. The results of the build phase are in there too, so if you clone it, you start at the cheap phase.
About the numbers
We measured the build and first-returns figures from session transcripts, priced at API rates. The one dollar per client is a projection from the design, and we will label it that way until a full client has gone through the pipeline.