An institution scores itself a three out of five on a dimension of its capacity to translate research into public benefit.
Six months later somebody asks why. Maybe a reviewer, maybe a provost, maybe the next person in the job. And the honest answer, in almost every assessment process that has ever run, is that a group of people discussed it in a room and three felt right.
That is not dishonesty. It is what happens when the tool you used was a workshop and a spreadsheet. The reasoning was real and it was thorough, and then it evaporated, because nothing in the process was built to hold it.
We built the thing that holds it.
The number is not the deliverable
The instinct with any assessment framework is that the output is the scores. Nine dimensions, a number on each, a shape you can compare against last time or against a peer.
The scores are the least durable part of the exercise. They are a summary of a conversation, and the conversation is where all the information actually was. Which programs somebody named as working. Which gap three different people raised independently. Which strength turned out on examination to rest on one person who is retiring.
If the platform captures only the summary, you have built a very expensive way to produce nine integers.
So evidence entry is a first-class thing in the product rather than a notes field at the bottom. A score is attached to the specific things that justified it, entered as the assessment happens, by the people making the claim. The trail is a product of doing the work rather than an act of documentation afterwards, which is the only version of this that ever survives contact with busy people.
Why a rubric is not a survey
A survey asks people what they think. A rubric asks them to place themselves against described states and then show why.
The difference sounds procedural and it is not. On a survey, a generous respondent and a hard marker produce different numbers from identical institutions, and you cannot tell which you are looking at. Against a described state with evidence attached, the disagreement surfaces in the open, where it is useful. Two people both looking at the same described level and reaching different conclusions is the single most productive thing that happens in one of these sessions.
Nine dimensions is also a deliberate number, and it is more than most institutions expect. Enough that a genuine weakness cannot hide behind an adjacent strength, few enough that the exercise finishes.
Facilitated, not automated
This one is worth stating plainly, because it runs against the current: the platform does not score anybody.
A facilitator runs the process. People argue. The platform holds the structure, keeps the evidence with the claim, and makes sure the same question gets asked of every dimension.
We could have built something that ingests institutional data and emits a readiness score. It would demo better than what we built and it would be worth less, for a reason that has nothing to do with model quality. The value of an assessment like this is substantially in the arguing. An institution that receives a score has learned a number. An institution that had to defend its own three out of five to its own colleagues has learned where it actually stands, and, more to the point, has a room full of people who now agree on it.
Automate that and you have optimised away the product.
What the trail buys you later
Three things, and the third is the one people do not anticipate.
The review conversation gets shorter. When somebody asks how you got that number, the answer is a link rather than a meeting.
The second assessment is worth more than the first. A score with no evidence behind it cannot be compared to anything, because you have no idea whether the change is real or whether this year's group was simply harder on themselves. With the trail intact, movement means something.
It survives turnover. The person who ran the assessment leaves. In a spreadsheet world, everything they knew leaves with them and the next cycle starts from nothing. This is the failure mode nobody plans for and everybody eventually meets.
The general version
This is the same principle as the retrieval systems in the capability finder and the warehouse chat layer, arriving from a different direction.
An output that a person cannot check is an output that person has to take on faith, and nobody takes a number on faith when their own credibility is attached to it. So they either verify it independently, which means your system saved them nothing, or they quietly stop using it.
The fix is the same every time and it is not clever. Attach the reasoning to the result, at the moment the result is produced, in a form the reader can actually follow. Everything else is easier once that is true.