A proposal is due this week. Someone on the team needs to know what a particular section should contain, whether the budget justification has to list a cost category separately, or how the program office reads a specific review criterion.
They will ask the person in the research office who has answered it forty times before. That person holds years of knowledge nobody wrote down, so every question waits for them. And they are in a meeting.
We built a tool that answers while that person is busy: the EPIIC Assistant, for the National Science Foundation's Enabling Partnerships to Increase Innovation Capacity (EPIIC) program. The key decision was which questions the AI should be allowed to answer on its own.
Two kinds of question
Spend a day with a research administrator and the questions fall into two groups almost at once.
The first group has right answers. What goes in this section. Whether this cost category is allowed. What the page limit is. How the participant list should be structured. These are not judgment calls. Someone in the office has already worked out the correct answer, checked it against the program guidance, and given it forty times in the same words.
The second group is open. Will this description of broader impacts work for this panel? Does this collaboration look real, or added at the last minute? Does this section do what the review criterion asks? These need reading, comparison and judgment.
Most systems built for this work send both groups to the AI. That is the mistake.
Common questions get the expert's own answer
For the first group, the system stores the answer the office already wrote, and returns it word for word.
It does not paraphrase it or have the model restate it in new words. The reader gets the text a senior person wrote and got right, exactly as they wrote it, every time.
That gives three benefits, and each one is worth the effort on its own.
It is correct every time. A generated answer is correct most of the time. A stored answer is correct every time. When the question is whether a cost category is allowed, nobody in research administration will accept an answer that is right most of the time.
You fix it in one place. When the guidance changes, someone edits one paragraph, and every later answer is right. The alternative is finding out that the model has given a slightly outdated answer for a month, then trying to fix it with prompts.
The office can stand behind it. The office can point at the text and say, yes, that is our position. That is what lets the office endorse the tool.
What the AI answers
The AI handles the open questions, and it answers them from the program guidance and the solicitation.
That guidance tells a team what the program asks for. It does not show what a strong answer looks like. That missing piece is where a first-time principal investigator struggles. They can read the solicitation perfectly well. What they cannot see is what a successful response looks like. If an office can share its own funded proposals, adding them to what the model can search gives writers examples in the institution's own voice, for the same kind of team.
These answers appear on screen as they are written, because they are long, and a blank screen for that long looks broken. They also show their sources. Advice about a proposal that you cannot trace back to the guidance is advice you have to check yourself, so it saves you no time.
Use fixed answers where a correct answer exists
A common belief in this work is that the more the model does, the more advanced the system is. We think the opposite is true.
The skill is knowing which parts of a problem have a correct answer, and keeping the model away from them. A calculator does not estimate. A system that tells you a page limit should not estimate either.
The split also solves the two problems that sink these projects: cost and speed. It costs less, because the most frequent questions never reach the model. And it is fastest on the questions people ask most, which is what brings them back a second time.
The same idea appears in a different form in our post on cutting AI agent costs: do the expensive thinking once, where it is needed, and not every time a question repeats.
What it costs to maintain
Someone has to own the stored answers. When the guidance changes, when a program adds a requirement, or when the office changes its position, a person who knows the subject has to update the text. That work never ends, and it decides whether the system is still good a year from now.
We think that cost is worth paying. A system without stored answers still needs maintenance, but nobody can see where. It drifts slowly away from what the office believes, and nobody can say when it happened.
Before you build one
Sort the questions first.
Take the fifty most common questions and put each one in one of two groups: it has a right answer that someone already knows, or it needs judgment. Then build two parts behind one search box. The first group goes to a well-maintained set of expert answers, with a good matching step in front of it. The second group goes to search and the model.
Most teams build only the second part and send every question to it. It looks impressive in a demo. Then it is wrong on exactly the questions where a wrong answer costs the most.