A researcher wants to know which site in a multi-site program has an instrument that can measure in-plane strain in a particular two-dimensional material.
That question has a real answer. It is on an instrument page, in a publication and in a program record. The person asking will not find it, because they would first need to know which of a dozen sites to look at. So they do what everyone does: they email the one person they know and hope that person passes it on.
We built a tool to shorten that search. The search itself uses standard methods. This post is about a different decision: what the system is allowed to do with an answer once it has found one.
Why not use ChatGPT?
A reviewer will ask this, and it is a fair question. A public assistant that can browse the web answers many of these questions for free. We estimated about two thirds of them.
The other third is where this tool earns its place. Here is what it offers.
One trusted source. A public assistant answers from whatever a search engine returned that day. The result changes with the wording of the question, and it can be a page from four years ago. Our tool answers from a maintained index of every instrument page, publication and program record, crawled again on a schedule.
A page the program owns. A chat session disappears. It has no web address, it is not a deliverable, and nobody can cite it in a report. A branded page that the program owns can be linked, cited and reported on.
Numbers the program can report. A public assistant tells the program nothing about who is asking what. Our tool logs every query, anonymized and reported only in total. A director who says "there is outside interest in our instruments" can then show the number.
Fast corrections. When a public model is wrong about your program, you wait for the next model release. When our tool is wrong, the error is in a document in the index, and fixing it takes an afternoon.
A person to contact. This matters most, so it has its own section.
Every answer ends with a person to contact
Every answer ends with the name and contact details of the facility manager who runs that instrument at that site.
This is the core of the design.
A researcher who asks about an instrument wants to start a conversation with the person who can say yes. The paragraph about the technique only helps them start that conversation well. A system that gives the right name has done the whole job. A system that writes three fluent paragraphs about the technique has done the easy half. It leaves the researcher where they started, only more confident.
It also makes the answer easy to check. Grading free text is hard, as everyone in this field knows. Grading a referral is simple: either the instrument is at that site and that person runs it, or it is not and they do not. When you can check whether an answer is correct, you can improve it.
What it refuses to do
The limits we set matter more than the text the model writes.
It answers only from pages crawled in the last thirty days. Out-of-date facility information is worse than none: a lead time that changed six months ago produces a confident answer that costs a researcher a month.
It declines any question the index does not cover. When it is not confident, it says plainly that it does not have that information, and gives the contact for the site most likely to have it.
Every factual claim links to the page it came from. The link sits on the claim itself, where a reader will see it, and not in a source list at the bottom.
We set tools like this to decline more often than a consumer product would. That makes the tool a little less satisfying to use, and we think it is the right trade. A researcher who gets a refusal emails somebody. A researcher who gets a confident wrong answer about instrument availability books travel.
Build only on data that already exists
One decision shaped the whole design: every feature runs on public data we can crawl, plus the program's own existing records. No site has to take on a new task.
This decides whether the tool lasts.
If a system needs eleven sites to each keep something up to date, it gets worse whenever the busiest site falls behind, and every site is the busiest one at some point in the year. Six months in, three sites have stopped, the data is visibly wrong, and nobody trusts the tool. When the system is built on what sites already publish, it improves when they do their normal work, and it stays accurate when they are overloaded.
It also answers the funding question that any program office will ask. The work builds on crawlers the program has already paid for, so it does not start from nothing.
Two things to know before you build one
The first version finds instruments and names the contact. Nothing else. No booking, no scheduling, no calendar integration. Each of those would turn a read-only system into one that writes into someone's calendar. That means a different security review, new ways to fail, and a new conversation with every site.
Write the data-residency paragraph before anyone asks for it. University IT departments push back on AI tools out of concern that their data will be used to train models. That concern does not apply to systems that call a commercial model through its application programming interface (API). Answer it in the proposal: name the provider, commit to US-only regions and isolated containers, and state that commercial API providers do not train on customer data by default. That paragraph takes an afternoon to write, and it does more for a procurement conversation than an accuracy benchmark.
The same idea applies elsewhere
Most search-and-answer systems are built to produce an answer. First ask what the answer is for.
When the real goal is a decision that a specific person has to make, the most useful output is the shortest correct path to that person, with enough context that the conversation starts in the right place. Models write fluent text easily. The harder design decision is when to stop writing and name the person to contact.