Research centers

Why every score in our assessment tool comes with its evidence

We built a facilitated self-assessment tool for a national research-impact network. It has nine dimensions and a five-step process. One rule shaped all of it: every score keeps the evidence that produced it.

By Ashish Tonse 5 min read

Dimensions
9
Steps in the process
5

An institution scores itself three out of five on one dimension: its capacity to turn research into public benefit.

Six months later, someone asks why. It might be a reviewer, a provost or the next person in the job. In almost every assessment process ever run, the answer is that a group of people discussed it in a room and three felt right.

Nobody was hiding anything. That is what happens when the tools are a workshop and a spreadsheet. The reasoning was real and thorough, and then it was lost, because nothing in the process was built to keep it.

We built a tool that keeps it.

The evidence matters more than the score

With any assessment framework, the instinct is to treat the scores as the result: nine dimensions, a number on each, and a shape you can compare with last time or with a peer.

The scores are the part of the exercise that lasts least. They summarize a conversation, and the conversation held the information: which programs someone named as working, which gap three different people raised on their own, which strength turned out to depend on one person who is about to retire.

If the tool captures only the summary, it is an expensive way to produce nine numbers.

So entering evidence is a main feature of the tool, with its own place in every step. Each score is linked to the specific things that justify it, entered during the assessment by the people making the claim. The record builds up as people do the work, so nobody has to write it up afterwards. With busy people, that is the only way a record like this survives.

Each score keeps the evidence behind it An institution scores itself 3 out of 5 on one of nine dimensions: its capacity to turn research into public benefit. People set the score, and the tool keeps the record. During the assessment, the people making the claim link the score to its evidence, such as the programs someone named as working, a gap three people raised on their own, and a strength that rests on one person. Six months later, a reviewer, a provost or the next person in the job can ask why and follow the score back to that evidence. ONE OF NINE DIMENSIONS Capacity to turn research into public benefit 3 OUT OF 5 SCORED BY THE PEOPLE IN THE ROOM THE EVIDENCE, KEPT WITH THE SCORE Programs named as working A gap three people raised A strength resting on one person ENTERED DURING THE ASSESSMENT ASKS WHY Six months later A REVIEWER, A PROVOST, OR THE NEXT PERSON IN THE JOB
The people making each claim attach the evidence during the assessment. When someone asks about a score later, the reasoning is already linked to it.

How a rubric differs from a survey

A survey asks people what they think. A rubric asks them to place themselves against a written description of each level, and then show why.

That difference changes the results. On a survey, a generous respondent and a strict one give different numbers for identical institutions, and you cannot tell which you are looking at. When people score against a written description with evidence attached, disagreements come out in the open, where they are useful. Two people reading the same description and reaching different conclusions is the most productive thing that happens in one of these sessions.

A survey compared with a rubric On the left, a survey asks people what they think. Two respondents rate the same institution: a generous one gives a high number and a strict one gives a low number. The result is two different numbers, and no way to tell which one you are looking at. On the right, a rubric gives each of five levels a written description. Two people place themselves against the same descriptions and attach evidence, and they reach different levels. Their disagreement is out in the open, which is the most productive part of a session. SURVEY: WHAT DO YOU THINK? The same institution A generous respondent A strict respondent NO DESCRIPTION OF WHAT EACH NUMBER MEANS Two different numbers NO WAY TO TELL WHICH YOU ARE LOOKING AT RUBRIC: WHERE DO YOU STAND, AND WHY? LEVEL 5 LEVEL 4 LEVEL 3 LEVEL 2 LEVEL 1 B EVIDENCE A EVIDENCE SAME TEXT, DIFFERENT READINGS EACH LEVEL HAS A WRITTEN DESCRIPTION The disagreement is in the open THE MOST PRODUCTIVE PART OF A SESSION
A survey gives you numbers that depend on who answered. A rubric ties each score to a written description and its evidence, so a disagreement shows up where people can discuss it.

Nine dimensions is a deliberate choice, and more than most institutions expect. That is enough that a real weakness cannot hide behind a nearby strength, and few enough that the exercise finishes.

People score, and the tool keeps the record

The tool does not score anyone.

A facilitator runs the process. People argue. The tool holds the structure, keeps the evidence with each claim, and makes sure the same questions are asked of every dimension.

We could have built something that reads institutional data and produces a readiness score. It would look better in a demo than what we built, and it would be worth less, for a reason that has nothing to do with model quality. Much of the value of an assessment like this comes from the argument. An institution that receives a score learns a number. An institution that had to defend its own three out of five to its own colleagues learns where it stands, and ends up with a room full of people who agree on it.

A tool that scores, compared with people who score What we could have built: a tool that reads institutional data and produces a readiness score. It looks better in a demo, and the institution learns a number. What we built: a facilitator runs the process, and colleagues argue it out and defend their own three out of five. The institution learns where it stands, and the room ends up agreeing on it. Underneath, the tool keeps the record: it asks the same questions of every dimension and keeps the evidence with each claim. The tool does not score anyone. WHAT WE COULD HAVE BUILT Reads institutional data Produces a readiness score LOOKS BETTER IN A DEMO The institution learns a number WHAT WE BUILT A facilitator runs it Colleagues argue it out AND DEFEND THEIR OWN 3 OUT OF 5 It learns where it stands AND THE ROOM AGREES ON IT The tool keeps the record SAME QUESTIONS FOR EVERY DIMENSION EVIDENCE KEPT WITH EACH CLAIM IT DOES NOT SCORE ANYONE
A tool that scores the institution hands it a number. When colleagues score themselves and the tool keeps the record, the institution learns where it stands.

What the evidence gives you later

Three things, and the third is the one people do not expect.

Reviews take less time. When someone asks how you got a number, you send a link and skip the meeting.

The second assessment is worth more than the first. A score with no evidence behind it cannot be compared with anything, because you cannot tell whether a change is real or whether this year's group was harder on itself. With the evidence kept, a change in score means something.

It survives staff turnover. The person who ran the assessment leaves. With a spreadsheet, everything they knew leaves with them, and the next cycle starts from nothing. Nobody plans for this, and everybody meets it eventually.

What the kept evidence does over time Two ways of running assessments, followed through four moments: the first assessment, someone asking why a score is what it is, the person who ran the assessment leaving, and the next assessment. With a workshop and a spreadsheet, the first assessment produces nine scores and the reasoning is not kept. A question needs a meeting to rebuild the answer. When the organizer leaves, what they knew leaves with them, and the next assessment starts from nothing, with no way to tell whether a change is real or the group was stricter. With scores kept with their evidence, each of the nine scores has its evidence. A question is answered by sending a link, with no meeting. The record stays for the next person in the job. And in the next assessment, a change in score means something. So reviews take less time, the second assessment is worth more than the first, and the record survives staff turnover. FIRST ASSESSMENT SOMEONE ASKS WHY ITS ORGANIZER LEAVES NEXT ASSESSMENT A workshop and a spreadsheet Scores kept with their evidence Nine scores THE REASONING IS NOT KEPT Nine scores EACH WITH ITS EVIDENCE A meeting TO REBUILD THE ANSWER Send a link SKIP THE MEETING Knowledge leaves WITH THE PERSON The record stays FOR THE NEXT PERSON IN THE JOB Starts from nothing IS A CHANGE REAL, OR A STRICTER GROUP? A change in score MEANS SOMETHING FASTER REVIEWS. A SECOND ASSESSMENT WORTH MORE. A RECORD THAT SURVIVES TURNOVER.
The evidence pays off after the assessment ends: when someone asks why, when the organizer leaves, and when the next assessment compares against this one.

The same idea in our other tools

The same idea sits behind the instrument finder and the plain-English warehouse queries.

If a person cannot check an output, they have to take it on trust, and nobody takes a number on trust when their own credibility depends on it. So they either check it themselves, which means the system saved them nothing, or they quietly stop using it.

What a reader does with a result they cannot check Top: a score or an answer arrives on its own, with no reasoning attached. The reader cannot check it and would have to take it on trust. So they either check it by hand, which means the system saved them nothing, or they quietly stop using it. Bottom: the same result arrives with its reasoning attached when it was produced. The reader follows the reasoning and uses the result, checked rather than taken on trust. The fix is to attach the reasoning to the result when the result is produced. A RESULT ON ITS OWN A score or an answer NO REASONING ATTACHED The reader CANNOT CHECK IT, MUST TAKE IT ON TRUST Checks it by hand THE SYSTEM SAVED NOTHING Quietly stops using it THE RESULT WITH ITS REASONING A score or an answer REASONING ATTACHED WHEN IT IS PRODUCED The reader FOLLOWS THE REASONING Uses the result CHECKED, NOT TAKEN ON TRUST THE FIX: ATTACH THE REASONING TO THE RESULT WHEN THE RESULT IS PRODUCED.
A result the reader cannot check ends in extra work or in disuse. The same result with its reasoning attached gets used.

The fix is the same every time, and it is simple. Attach the reasoning to the result when the result is produced, in a form the reader can follow. Everything else gets easier after that.

Building something similar for your center?

We build websites, researcher directories and AI assistants for research centers and programs. Tell us about your center in a 30-minute call.

Our work for research centers

Book a 30-minute call