Skip to content
Predictive Systems

Case study · Oral assessment

The tie that wasn’t

Three AI models finished within 0.2 points of each other. Predictive Systems shipped the one that lost a category outright, and says the loss is the reason the test was worth running.

Allan Tan · Founder and Chief AI Scientist


The scoreboard read 97.7, 97.5, 97.7.

By any ordinary reading, the three models Predictive Systems put through its July evaluation were the same model. Two tied. The third trailed by two tenths of a point, a gap narrower than the noise in most benchmarks. A procurement committee handed that table would have picked on price and moved on.

The Manila company picked differently, and it took fifteen numbers instead of one.

The models were competing for a job inside better-ed, PSI’s oral assessment product, which runs today at Barcelona Academy, at Kulosaari Secondary School in Helsinki, at Karelia University of Applied Sciences, at Jubilee Christian Academy, and at the Philippine Science High School’s main campus. The task sounds narrow and is not: write the questions a teacher asks a student out loud.

Get it wrong and the failure is not a bad score on a dashboard. It is a fourteen-year-old being asked something that gives away its own answer, or a question that punishes a student for not already knowing the thing being taught.

The average is the thing that hides it

PSI’s evaluation suite scores a generated assessment on fifteen separate dimensions. Does the question read like a person speaking, or like an exam paper? Can the answer be read straight off the question? When a student stumbles, does the follow-up break the problem into a smaller step, or hand over the solution? Is the question about the objective it was tagged with, or something adjacent?

Each is judged separately. The overall score is the average of all fifteen, and averaging is exactly where three distinct models became one indistinguishable number.

15 numbers, not 1

What it took to tell three near-identical models apart

One model was perfect at breaking hard questions into smaller steps and lost fifteen points on matching its follow-up to what the student had actually done. Another was flawless at staying on topic and weakest at sounding like a human being. The third, the one PSI shipped, was best or tied on thirteen of the fifteen dimensions and ran the whole suite in 35 seconds, roughly 1.5 times faster than the incumbent it replaced.

EdTech exampleWhich model writes the best question?
DimensionModel 1Model 2Model 3shipped
Overall scoreThe average across all fifteen dimensions below. The one number that hides everything under it.97.797.597.7
Follow-up fitDoes the follow-up match what the student actually did: affirm a correct answer, clarify a partial one, re-engage a wrong one?100.085.0100.0
ConversationalDoes it read like a tutor speaking aloud, rather than a written exam paper?94.695.0100.0
No answer leakCan the answer be read straight off the question? One leak fails the whole case.97.597.9100.0
Objective alignmentIs the question about the objective it was tagged with, rather than something next to it?96.598.3100.0
Probe qualityDoes a deeper follow-up stay on the same idea, or wander to a different one?93.192.893.1
ScaffoldingWhen a student stumbles, does the follow-up break the question into a smaller step, without handing over the answer?95.0100.080.0
Cultural alignmentDo the everyday references, money, names, places, units, fit the locale the assessment was written for?100.099.295.0
Procedural, not computationalIn maths and science, does it ask the student to talk through the method rather than compute a number? Recitation is spoken.92.395.697.9
Answer plausibilityDoes the model answer sound like a student saying it out loud, rather than a textbook paragraph?98.899.299.3
Standalone contextDoes each question make sense at the point it is asked? The first one has no earlier question to lean on.98.8100.0100.0
Age appropriateAre the vocabulary and sentence length calibrated to the grade it was written for?100.0100.0100.0
Difficulty accuracyDoes the real difficulty match the easy, medium or hard label the question carries?100.0100.0100.0
Language consistencyIs every part of the question in one language, and the language the source material is in?99.0100.0100.0
Subject styleDoes the question fit how the subject is actually asked about? History invites cause, literature invites interpretation.100.0100.0100.0
Bloom accuracyDoes the question demand the level of thinking its objective asks for, rather than settling two levels below?100.0100.0100.0
Time per caseWall clock to generate one full assessment, averaged over the run.51s91s35s

We shipped Model 3. Best or tied on thirteen of fifteen, and 1.5× faster. All three land within 0.2 on the overall score, and up to 20 points apart underneath it.

All fifteen dimensions · Recitation eval, three-way comparison, 8 Jul 2026

It also scored 80.0 on scaffolding, against a rival’s 100.

Shipping the model that lost a row

That number is on the table PSI publishes. It was not rounded off, moved into an appendix, or quietly dropped from the comparison.

One-shotting the entire question cannot generate good scaffolds.

The model was being asked to produce a question, its expected answer, and the smaller supporting steps a teacher would fall back on, all in a single pass. The supporting steps were what suffered.

The fix was not a better prompt. It was a second agent, whose only job is writing the scaffolds.

That is a structural answer rather than a tuning one, and it is the kind of answer an average would never have surfaced. Had PSI been reading a single composite score, the shipped model and the incumbent would have looked identical, the weakness would have gone to production unnoticed, and it would have been discovered the way these things usually are: by a teacher, in a classroom, with a student waiting.

The judges get graded too

There is a harder problem sitting behind all of this, and PSI’s reports name it.

The fifteen dimensions are scored by AI judges, which raises the obvious question of who judges the judges. PSI’s answer is a human-labelled gold set of 164 cases. Every judge is measured against it on two axes: raw agreement, and Cohen’s kappa, a statistic that discounts the agreement you would get by chance alone.

A judge is only trusted as a gate, a pass-or-fail check, when it reaches 0.90 agreement and a kappa of 0.6. Seven now clear that bar. The rest are kept as diagnostics, which explain rather than block.

Scaffolding is one of the diagnostics. Its kappa sits at 0.47, and it moves between runs.

Which means the dimension where the shipped model looks worst is also the dimension its judge is least sure about. PSI publishes both facts, in the same report, and treats the second as a reason to keep watching the first rather than a reason to discount it.

What the test found that wasn’t about models

The most interesting result was not about any of the three models.

Not all teachers think alike, or teach alike. Some metrics go against another teacher’s preference.

That is a genuinely awkward finding for anyone building assessment software. A scaffold one teacher considers supportive, another considers a giveaway. A follow-up one reads as probing, another reads as badgering. The disagreement is not noise in the measurement; it is the actual state of the profession, and any metric that scores pedagogy has to decide whose pedagogy it is scoring.

An evaluation suite cannot resolve that. What it can do is make the choice visible, which is a considerably more honest position than a single number that quietly encodes one philosophy of teaching and presents it as accuracy.

The number nobody should trust

The lesson PSI takes from July is not that its model won.

Two of the three models tied on the headline figure. The one it shipped concedes a category on the published table. The judge for that category is the least reliable of the fifteen. And the most useful thing the exercise produced was a question about teaching that no amount of further measurement will settle.

None of that is visible from 97.7.


Want the same measured before it ships?

Let’s Talk AI