It is four in the afternoon. A proposal is due, and somebody on the team needs to know what a particular section is supposed to contain, whether the budget justification has to break out a category, or how the program office reads a specific review criterion.
They will ask whoever in the research office has answered it forty times before. That person is a bottleneck made entirely of institutional memory, and they are also in a meeting.
We built the thing that answers at four in the afternoon. The interesting decision was working out which questions it should be allowed to think about.
Two kinds of question
Sit with a research administrator for a day and the questions sort themselves into two piles almost immediately.
The first pile has right answers. What goes in this section. Whether this cost category is allowable. What the page limit is. How the participant list is supposed to be structured. These are not judgment calls. Somebody in that office has already worked out the correct answer, checked it against the program guidance, and given it forty times in the same words.
The second pile is genuinely open. Is this framing of the broader impacts likely to land with this panel. Does this collaboration read as real or as bolted on. Is this section doing the work the criterion is asking for. These need reading, comparison, and something closer to taste.
Almost every system built in this space sends both piles to the same place. That is the mistake.
The expert answer, served as written
For the first pile, the system holds the answer the office already wrote, and returns it.
Not paraphrased. Not regenerated with the same meaning in fresh words. The text somebody senior sat down and got right, delivered exactly as they wrote it, every time.
Three things follow from that, and each of them is worth the trouble on its own.
It is correct on the thousandth answer. A generated answer is correct in distribution. A written one is correct. When the question is whether a cost category is allowable, "usually right" is not a category anyone in research administration is willing to occupy.
It is fixable in one place. When the guidance changes, someone edits one paragraph and every future answer is right. Compare that with discovering the model has been producing a subtly outdated answer for a month and trying to prompt your way out of it.
It is attributable. The office can point at the text and say yes, that is our position. That sentence is doing a lot of quiet work. It is the difference between a tool the office endorses and a tool the office tolerates.
What is left for the model
The open pile, grounded in two corpora that matter.
The first is the program guidance and solicitation, which is the obvious one. The second is the office's own funded proposals, which is the one people skip, and it is where most of the value is.
Program guidance tells you what is being asked for. Funded proposals from that same office show you what an answer that worked looked like, in this institution's voice, for this kind of team. The gap between those two is exactly the gap a first-time principal investigator falls into. They can read the solicitation perfectly well. What they cannot see is the shape of a successful response.
Answers here stream, because they are long enough that waiting in silence feels broken. Answers here also carry their sources, because advice about a proposal that you cannot trace back to the guidance is advice you have to independently verify, which means it saved you nothing.
This is not laziness, it is where determinism belongs
There is a reflex in this work that says the more the model does, the more advanced the system is. It is backwards.
The skill is knowing which parts of a problem have a correct answer and refusing to let a probabilistic system anywhere near them. A calculator does not estimate. A system that returns a page limit should not either.
Practically, the split also fixes the two problems that sink these projects. It is cheaper, because the highest-frequency questions never reach an API. And it is faster on exactly the questions people ask most, which is what makes somebody come back to it a second time.
The same instinct shows up in a different form in Papa Claude and Baby Claude: put the expensive thinking where it is needed once, not where it repeats.
What it costs to keep
An honest note, because this shape is not free.
Somebody has to own the written answers. When the guidance changes, when a program adds a requirement, when the office's position on something shifts, the text has to be updated by a person who knows. That is a real, recurring obligation and it is the thing that decides whether the system is still good in a year.
We think it is the right cost. The alternative is not "no maintenance", it is maintenance you cannot see: a system slowly drifting away from what the office actually believes, with nobody able to point at the moment it happened.
The general version
Before you build anything in this shape, sort the questions.
Take the fifty most common, and put each one in a pile: does this have a right answer that somebody already knows, or does it need judgment. Then build two systems that share a front door. The first pile is a well-maintained set of expert answers with a good matcher in front of it. The second is where the retrieval and the model earn their place.
Most teams build only the second and let it swallow the first. It demos beautifully, and it is wrong just often enough, on exactly the questions where being wrong is expensive.