Research centers

Why our proposal advisor answers common questions without AI

We built a proposal advisor for a National Science Foundation program. Common questions get an answer that an expert already wrote. AI answers the rest from the program guidance. The part worth copying is how many questions it answers without AI.

By Ashish Tonse 6 min read

A proposal is due this week. Someone on the team needs to know what a particular section should contain, whether the budget justification has to list a cost category separately, or how the program office reads a specific review criterion.

They will ask the person in the research office who has answered it forty times before. That person holds years of knowledge nobody wrote down, so every question waits for them. And they are in a meeting.

We built a tool that answers while that person is busy: the EPIIC Assistant, for the National Science Foundation's Enabling Partnerships to Increase Innovation Capacity (EPIIC) program. The key decision was which questions the AI should be allowed to answer on its own.

Two kinds of question

Spend a day with a research administrator and the questions fall into two groups almost at once.

The first group has right answers. What goes in this section. Whether this cost category is allowed. What the page limit is. How the participant list should be structured. These are not judgment calls. Someone in the office has already worked out the correct answer, checked it against the program guidance, and given it forty times in the same words.

The second group is open. Will this description of broader impacts work for this panel? Does this collaboration look real, or added at the last minute? Does this section do what the review criterion asks? These need reading, comparison and judgment.

Most systems built for this work send both groups to the AI. That is the mistake.

One search box, two ways to answer A question typed into one search box goes to a matching step. If it is a common question with a right answer, the tool returns the answer an expert in the office already wrote, word for word. That answer is correct every time, the office owns and updates it, and the question never reaches the model. If the question needs judgment, AI answers it from the program guidance and the solicitation, and the answer appears as it is written, with its sources. THE QUESTION MATCHING STEP WHO ANSWERS WHAT THE READER GETS A question TYPED IN ONE SEARCH BOX A common question? YES, A RIGHT ANSWER Stored expert answer RETURNED WORD FOR WORD OWNED AND UPDATED BY THE OFFICE The office's own words CORRECT EVERY TIME NEVER REACHES THE MODEL NO, IT NEEDS JUDGMENT Answered by AI FROM THE PROGRAM GUIDANCE AND THE SOLICITATION An answer with sources SHOWN AS IT IS WRITTEN TRACED TO THE GUIDANCE
Both paths sit behind one search box. A question with a right answer gets the office's own stored answer, word for word. A question that needs judgment goes to the AI, which answers from the program guidance and shows its sources.

Common questions get the expert's own answer

For the first group, the system stores the answer the office already wrote, and returns it word for word.

It does not paraphrase it or have the model restate it in new words. The reader gets the text a senior person wrote and got right, exactly as they wrote it, every time.

That gives three benefits, and each one is worth the effort on its own.

It is correct every time. A generated answer is correct most of the time. A stored answer is correct every time. When the question is whether a cost category is allowed, nobody in research administration will accept an answer that is right most of the time.

You fix it in one place. When the guidance changes, someone edits one paragraph, and every later answer is right. The alternative is finding out that the model has given a slightly outdated answer for a month, then trying to fix it with prompts.

The office can stand behind it. The office can point at the text and say, yes, that is our position. That is what lets the office endorse the tool.

A guidance change, fixed in one place Two rows of answers over time, before and after the program guidance changes. With a stored answer, someone edits one paragraph when the guidance changes, and every later answer is right. With a generated answer, the model keeps giving a slightly outdated answer, for a month in this example, until someone finds out and then tries to fix it with prompts. EACH DOT IS ONE ANSWER GIVEN Stored answer THE OFFICE'S OWN TEXT Generated answer THE MODEL RESTATES IT THE GUIDANCE CHANGES One paragraph edited EVERY LATER ANSWER IS RIGHT SLIGHTLY OUTDATED, FOR A MONTH Someone finds out, then tries to fix it with prompts. TIME Current answer Outdated answer
When the guidance changes, someone edits one stored paragraph and every later answer is right. A generated answer stays slightly outdated until someone notices, and then the fix is a round of prompt changes.

What the AI answers

The AI handles the open questions, and it answers them from the program guidance and the solicitation.

That guidance tells a team what the program asks for. It does not show what a strong answer looks like. That missing piece is where a first-time principal investigator struggles. They can read the solicitation perfectly well. What they cannot see is what a successful response looks like. If an office can share its own funded proposals, adding them to what the model can search gives writers examples in the institution's own voice, for the same kind of team.

These answers appear on screen as they are written, because they are long, and a blank screen for that long looks broken. They also show their sources. Advice about a proposal that you cannot trace back to the guidance is advice you have to check yourself, so it saves you no time.

An open question, answered by AI with sources A mock of the EPIIC Assistant answering an open question: does this section do what the review criterion asks? The AI's answer appears line by line as it is written, because the answers are long and a blank screen for that long looks broken. Source links to the program guidance and the solicitation sit inside the answer, because advice you cannot trace back to the guidance saves you no time. The tool searches the program guidance and the solicitation, and, if the office shares them, its own funded proposals, which give writers examples in the institution's own voice. EPIIC ASSISTANT Does this section do what the review criterion asks? ANSWERED BY AI PROGRAM GUIDANCE SOLICITATION WHAT IT SEARCHES Program guidance Solicitation Funded proposals, if shared Appears as it is written The answers are long. A blank screen for that long looks broken. Shows its sources Advice you cannot trace to the guidance saves you no time. The office's funded proposals If shared: examples in its own voice.
An open question gets an answer from the AI. The answer is written on screen as it is generated, and it shows its sources in the program guidance and the solicitation.

Use fixed answers where a correct answer exists

A common belief in this work is that the more the model does, the more advanced the system is. We think the opposite is true.

The skill is knowing which parts of a problem have a correct answer, and keeping the model away from them. A calculator does not estimate. A system that tells you a page limit should not estimate either.

The split also solves the two problems that sink these projects: cost and speed. It costs less, because the most frequent questions never reach the model. And it is fastest on the questions people ask most, which is what brings them back a second time.

The most asked questions never reach the model An illustrative chart, not measured data. Questions are ranked from most asked to least asked, and each bar shows how often a question is asked. The most asked questions have right answers, so they get stored expert answers: no model call, the fastest response and the lowest cost. The rest need judgment and go to search and the model. QUESTIONS, MOST ASKED FIRST ILLUSTRATIVE, NOT MEASURED Stored expert answers NO MODEL CALL FASTEST, AND COSTS LEAST Search and the model QUESTIONS THAT NEED JUDGMENT HOW OFTEN EACH QUESTION IS ASKED
The questions people ask most often have right answers, so they never reach the model. The tool is cheapest and fastest exactly where it is used most. The chart shows the shape, not measured data.

The same idea appears in a different form in our post on cutting AI agent costs: do the expensive thinking once, where it is needed, and not every time a question repeats.

What it costs to maintain

Someone has to own the stored answers. When the guidance changes, when a program adds a requirement, or when the office changes its position, a person who knows the subject has to update the text. That work never ends, and it decides whether the system is still good a year from now.

We think that cost is worth paying. A system without stored answers still needs maintenance, but nobody can see where. It drifts slowly away from what the office believes, and nobody can say when it happened.

Before you build one

Sort the questions first.

Take the fifty most common questions and put each one in one of two groups: it has a right answer that someone already knows, or it needs judgment. Then build two parts behind one search box. The first group goes to a well-maintained set of expert answers, with a good matching step in front of it. The second group goes to search and the model.

Sort the questions before you build The fifty most common questions are sorted into two groups. The first group has a right answer that someone already knows, such as what goes in a section, whether a cost category is allowed, what the page limit is, and how the participant list should be structured. The second group needs judgment, such as whether a broader impacts section will work for this panel, whether a collaboration looks real, and whether a section does what the review criterion asks. The first group is built as a maintained set of expert answers with a matching step in front of it. The second group is built as search and the model. Both parts sit behind one search box. SORT THE FIFTY MOST COMMON QUESTIONS It has a right answer someone knows SAME ANSWER, SAME WORDS, EVERY TIME It needs judgment READING, COMPARISON AND JUDGMENT What goes in this section? Will this broader impacts section work for this panel? Is this cost category allowed? Does this collaboration look real? What is the page limit? Does this section do what the review criterion asks? How should the participant list be structured? Build: a set of expert answers WELL MAINTAINED, A MATCHING STEP IN FRONT Build: search and the model ANSWERS FROM THE PROGRAM GUIDANCE BOTH PARTS SIT BEHIND ONE SEARCH BOX
Sort the fifty most common questions first. The ones with a right answer become stored expert answers. The ones that need judgment go to search and the model. Both parts sit behind one search box.

Most teams build only the second part and send every question to it. It looks impressive in a demo. Then it is wrong on exactly the questions where a wrong answer costs the most.

Building something similar for your center?

We build websites, researcher directories and AI assistants for research centers and programs. Tell us about your center in a 30-minute call.

Our work for research centers

Book a 30-minute call