An institution scores itself three out of five on one dimension: its capacity to turn research into public benefit.
Six months later, someone asks why. It might be a reviewer, a provost or the next person in the job. In almost every assessment process ever run, the answer is that a group of people discussed it in a room and three felt right.
Nobody was hiding anything. That is what happens when the tools are a workshop and a spreadsheet. The reasoning was real and thorough, and then it was lost, because nothing in the process was built to keep it.
We built a tool that keeps it.
The evidence matters more than the score
With any assessment framework, the instinct is to treat the scores as the result: nine dimensions, a number on each, and a shape you can compare with last time or with a peer.
The scores are the part of the exercise that lasts least. They summarize a conversation, and the conversation held the information: which programs someone named as working, which gap three different people raised on their own, which strength turned out to depend on one person who is about to retire.
If the tool captures only the summary, it is an expensive way to produce nine numbers.
So entering evidence is a main feature of the tool, with its own place in every step. Each score is linked to the specific things that justify it, entered during the assessment by the people making the claim. The record builds up as people do the work, so nobody has to write it up afterwards. With busy people, that is the only way a record like this survives.
How a rubric differs from a survey
A survey asks people what they think. A rubric asks them to place themselves against a written description of each level, and then show why.
That difference changes the results. On a survey, a generous respondent and a strict one give different numbers for identical institutions, and you cannot tell which you are looking at. When people score against a written description with evidence attached, disagreements come out in the open, where they are useful. Two people reading the same description and reaching different conclusions is the most productive thing that happens in one of these sessions.
Nine dimensions is a deliberate choice, and more than most institutions expect. That is enough that a real weakness cannot hide behind a nearby strength, and few enough that the exercise finishes.
People score, and the tool keeps the record
The tool does not score anyone.
A facilitator runs the process. People argue. The tool holds the structure, keeps the evidence with each claim, and makes sure the same questions are asked of every dimension.
We could have built something that reads institutional data and produces a readiness score. It would look better in a demo than what we built, and it would be worth less, for a reason that has nothing to do with model quality. Much of the value of an assessment like this comes from the argument. An institution that receives a score learns a number. An institution that had to defend its own three out of five to its own colleagues learns where it stands, and ends up with a room full of people who agree on it.
What the evidence gives you later
Three things, and the third is the one people do not expect.
Reviews take less time. When someone asks how you got a number, you send a link and skip the meeting.
The second assessment is worth more than the first. A score with no evidence behind it cannot be compared with anything, because you cannot tell whether a change is real or whether this year's group was harder on itself. With the evidence kept, a change in score means something.
It survives staff turnover. The person who ran the assessment leaves. With a spreadsheet, everything they knew leaves with them, and the next cycle starts from nothing. Nobody plans for this, and everybody meets it eventually.
The same idea in our other tools
The same idea sits behind the instrument finder and the plain-English warehouse queries.
If a person cannot check an output, they have to take it on trust, and nobody takes a number on trust when their own credibility depends on it. So they either check it themselves, which means the system saved them nothing, or they quietly stop using it.
The fix is the same every time, and it is simple. Attach the reasoning to the result when the result is produced, in a form the reader can follow. Everything else gets easier after that.