Instrument state: never proof-tested UX4Tech working note, part one
UX4Tech Blog Contact

Can Jev tell you the question was wrong?

A new class of cheap classifier promises judgement everywhere. We spent thirty hours leaning on our own instruments instead of judgement, and kept the record. Try the one below before you read on.

leak_scan0.1

outgoing: "Draft summary attached for review."

CLEAN

It printed CLEAN. That looks like assurance. Plant a term it is supposed to catch and run it again.

UX4Tech, 21 September 2026. Working note, part one. About twelve minutes.

The short answer

Our own checks looked fine and were not. A scan ran against an empty list and printed CLEAN. A scoring script read the words "No BLOCKING findings" and recorded a block. Both printed results that looked completely normal, and neither was capable of printing a bad one. We had green lights that could never have turned red, and we were an inch from publishing numbers based on them.

So when a new class of cheap classifier arrived promising that we can finally afford to put judgement everywhere, we had a specific question rather than an opinion. Can Jev tell you the question itself was wrong?

What we found, stated plainly before the evidence: You can give Jev an option such as “question invalid”, and that lets it flag a possible problem — but we have not tested how reliably it does so. And its typed response does not itself explain anything: someone, or another step, still has to investigate and confirm the cause. That is a flag. It is not a diagnosis, and the difference is the whole of this piece. Our own evaluation was flawed enough that it establishes neither that Jev is reliable enough for this task nor that it is not.

A check that has never been shown to fail is not a check. It is a decoration that happens to be green.

Safety engineering has a name for the step we had skipped. A safety function can sit idle for years waiting for a demand that may never come, so it is exercised on a schedule to show it can still trip — a proof test. Until an instrument has been shown to fail, a clean result from it tells you very little. A smoke alarm that stayed quiet all year may be working, or may have a dead battery; the silence is identical either way. That is why you press the test button.

We run a consulting practice with a team of AI agents. Over one weekend we leaned on four instruments of our own to tell us whether our work was fit to publish. Not one had ever been asked whether it could fail.

A working note, not a finished argument. The numbers are ours and some of them are embarrassing. It will be revised, and the open questions at the end are real.

What Jev actually changes

Classifiers are not new. What held them back was never the capability — it was that each one needed its own labelled data, its own training, its own upkeep. That setup cost is why millions of small judgements scattered through ordinary software never got made. Not because they were hard. Because no single one of them was worth the build.

Jev removes the training. You describe the categories in words instead of first building and labelling a classifier of your own. You still have to define the options, the context, the escalation rules and the privacy controls, and then validate the whole workflow — our weekend is evidence of exactly that work. We agree with most of what is being claimed for that, and we measured the economics ourselves rather than taking them from a vendor page.

run, 20 September 2026our one task
Sixtyscreening decisions over our own documents
A tenth of a pennyfor all sixty together, on Jev
A third of a secondeach, roughly

That price genuinely changes what you can afford to build. A second opinion behind every alert was never worth costing out before; now it is. Put Jev where a miss is cheap and a flag is useful — triage, routing, ranking, deciding what a person looks at first — and it earns its place immediately.

Our one disagreement is narrow, and it is about where you put it, not what it is. What a wrong answer costs depends entirely on where you put it. Routing a support ticket, picking the best hundred questions out of ten thousand, choosing which button to press — each of those can be cheap or expensive depending on the setting. Ours was a publishing gate, where a wrong answer means a retracted claim goes out silently, and that is the consequence our experiment did not evaluate reliably.

The rest of this is how we found that out, by breaking our own tools first.

What we actually did

The evidence, if you want it. Four instruments of ours, what each one got wrong, and what is wired now. If you only wanted the answer, you already have it.

Our agents write a lot. A few times this year they produced something we later had to retract, so they built a checker: a list of retracted claims and a pattern-match that a person runs before anything is published.

Last week we finally checked the checker.

Instrument one: the claim checker1.0
30 sentences flagged in our own documents, then read by hand, one at a time
"do not publish this number"a sentence forbidding the claimFLAGGED
a sentence criticising the claimFLAGGED
a note about work not yet startedread as a claim it was doneFLAGGED
most of the thirty were fine
What failed
It fired constantly and had never once been tested against a miss. Nobody knew its hit rate, in either direction.
What caught it
Reading all thirty flags by hand.
What is wired now
Every rule now carries its own canary: a sentence it must catch. The screen runs all eleven before it is allowed to report anything, and refuses to report at all if any rule fails to catch its own. This checks the rules that are configured; it does not check whether the rule list covers everything it should. It also cannot yet catch a rule removed together with its canary — the proof test's own open edge.
Instrument two: the measuring script0.1
rulethe checker blocked something if its output contains "BLOCK"
output, when nothing is wrongNo BLOCKING findings.
scored asblocked
What failed
The script written to score instrument one mistook a message saying nothing was blocked for a blocked result, because both contained the same word. That is the exact bug it was measuring, and it produced a result we were an inch from publishing.
What caught it
Its author reading the raw output instead of the score.
What is wired now
Scoring asserts on structure, never on words in a log. Throwaway scripts get the same rule as committed ones, because that is where every one of these lived.
Instrument three: the scan with nothing to scan for0.1

CLEAN

What failed
It searched outgoing text against a list of terms we never publish. It ran with an empty list and printed CLEAN. A scan with nothing to scan for cannot fail, and it did not.
What caught it
A reviewer who asked for evidence that the scan could fire, not evidence that it had passed.
What is wired now
The guard plants a known term first and refuses to clear anything unless the scan catches it. It sits on the execution path, so it cannot be skipped by forgetting. That is the instrument at the top of this page.

Two of these printed passes that could never have been fails. The first was the mirror image: it fired constantly and had never once been tested against a miss.

And one that was worse than useless

We fixed the checker's false alarms. The fix was measured, improved things on every axis we tested, and was approved.

Instrument four: the fix1.1
false alarmsdown
every axis we testedbetter
reviewapproved
then a reviewer tried to break it
real violations, cases built to break it5 of 7 passed
What failed
A checker that cries wolf is annoying. A checker that fails silently while being cited as a gate is worse than no checker, because a pass looks like assurance. We had spent the day telling people to run it.
What caught it
A reviewer on a different model family from the builders, whose only job was to break the fix.
What is wired now
Current policy: we treat the checker as advisory. It decides what a person reads first, not what nobody reads. A clean result is not publication clearance, and we do not cite it as a gate.

We also discovered it had never been pushed to the other machines. It existed on one node for four days while two agents were being told to use it.

What actually caught all of this

Four reviewers, seven versions of one piece of writing, and every version was one we were ready to publish. The catches:

A difference that was noise. We reported one model outperforming another. Two items apart out of twenty-six — close enough that the gap could easily have been chance. The comparison should not have existed.

A number that was not a number. The replacement metric subtracted the share of flags that were genuine from the share of false alarms a model cleared. They were measuring different things: one percentage was out of 30 flags, the other out of 24 false alarms. It looked like a measurement.

Instructions that contradicted the answer key. Our own prompt told both models a particular phrasing was acceptable; our answer key then marked them wrong for accepting it. We instructed them to comply and graded them down for complying.

A voice that was not ours to use. The draft claimed work in the first person that a different party had done.

And a claim about automation that was false. The draft said the checker "scans everything before it goes out." It does not. It runs when someone remembers.

None of these came from a check. Reviewers raised them by reading the work and asking whether it was true; arithmetic and reproductions then tested what they raised. Review and testing worked together — and two of the objections overturned conclusions three reviewers had already approved.

Then we tried to define what judgement is, and our own reviewers split

This is the part we did not expect and the reason this is a working note rather than a conclusion.

We proposed a distinction: a classifier answers the question you asked; judgement notices the question is wrong. We put it to our own review layer.

Reviewer one: the distinction is too absolute

A fixed menu of answers does not stop it flagging a bad question. "Question invalid", "inconsistent instructions", "other, escalate" are all representable options. And a generative model can confidently miss a flawed premise too. So this is a question of how you build and test the thing, not a hard line between people and machines.

Reviewer two: it is about how the thing is built, and the first wording went in a circle

Defining judgement as "what a reviewer noticed" is a claim you could never prove wrong: every catch counts as judgement, every miss counts as classification. It explains everything, which means it predicts nothing. Anchor it in how the thing is built instead. A classifier can only choose from the list of answers you gave it. Hand it a question where none of them is right and it must still pick one. You can test that before you rely on it, instead of arguing about it afterwards.

We think both are right, and what settles it between them is the useful part.

A cheap flag is not an explanation. Finding the cause is a separate step.
the flagtyped
options you supplied:
  acceptable
  violation
  none_of_these   only there if you thought to add it

returns:
  choice: none_of_these
the diagnosisa reader

∵ Your instructions tell the model this phrasing is acceptable, and your answer key marks it wrong for accepting it.

The first panel shows the shape of a typed answer, not a measured model output. The second paraphrases what a reviewer told us.

You can put "none of these" in the option set, but you have to include an escape route, and test what happens when it is used, and what you get back is a signal, not an explanation. Something is off here is worth a great deal at these prices — though the cost we measured was for our screening task, and we have not tested an invalid-question detector at all. Here is what is off, and here is why the question was the wrong question is a different job, and this weekend it was done by a reviewer every single time.

What we would actually tell you to do

The placement rule

Use the cheap layer to decide what a person looks at first. Do not use it to decide what nobody looks at. Those sound similar and are opposites: one narrows attention, the other removes it.

And a specific one from our own record: if your fast check clears something and the second layer only ever sees flagged items, that second layer cannot catch the miss.

The proof-test rule

Before relying on a clean result, check that the instrument rejects a known violation. Then test the clean cases, and the failure modes that would actually cost you something. One planted example proves a single path works; it does not prove coverage. That applies to a regex, a scan, a scoring script, and a classifier you rent by the token.

This borrows the idea of exercising a safeguard on purpose. It is not a functional-safety validation of our software.

What we are not claiming

Our run does not establish how reliable any of these models are. The evaluation was invalidated by our own defects: the prompt contradicted the answer key, and we decided what counted as a right answer after we had already seen what the models said. Failure to demonstrate reliability is not the same as demonstrating unreliability, and we will not smuggle one in as the other. A proper holdout, labelled before anyone sees a model output, is what this question deserves and we have not run it yet.

We are also not claiming this is a human-versus-machine story. Three of the four reviewers who caught our mistakes were AI agents. One runs on a different model family from the others, deliberately, to get a different perspective — which is not a guarantee of independent errors, and it broke two conclusions the rest of us had signed off. Machines can review, and ours reviewed well.

But every one of those reviewers worked inside the frame we handed them. They checked whether our sentences were true, and killed nine that were not. They also pushed on the frame where they could see it — one reviewer challenged the headline directly, on the grounds that it promised a capability test we had not finished. But an earlier draft of this article named a product in its headline and then never mentioned it again, left its own central question unanswered until two-thirds of the way through, and was written in language its intended reader could not follow. None of those three surfaced in review. A person found all of them in one evening, by asking something none of us had been handed: is this the right thing to be checking at all?

Which is the same distinction this piece has been making about Jev, arriving from the other side. A closed menu can only return what you put on it. An open-ended reviewer can tell you the menu is wrong. Judgement is the part that steps outside the frame — and in our shop it still took a human to do it.

Where we have landed on Jev

We see enough promise to keep evaluating Jev for advisory triage — deciding what a person looks at first. The economics are real, we measured them ourselves, and they change what is worth building: a second opinion behind every alert, a sanity check on every draft, a triage pass over a whole queue. All of that was priced out of existence a month ago and is not any more. That is a genuine shift, and most of what is being said about the economics is right. Cost and latency on one task is what we measured; nothing else.

What we would not do is let it decide what nobody looks at. It gives you a flag, fast and cheap. The step after the flag — what is actually wrong here, and was this even the right question — is still ours. On one weekend of evidence we are content with that division of labour, and we would rather say so plainly than pretend we tested more than we did.

Open questions we are carrying

  • Does calibration hold where being wrong is expensive? We have not evaluated calibration on our own task, or established that published results transfer to it. We have not seen a reliability curve on a task with real consequences, and we have not run one.
  • Who sets the escalation threshold? Vendor-suggested bands are a guess about someone else's risk appetite.
  • Does a classifier in the outer loop degrade safely? Letting one choose between tool, cheap model, frontier model and human is appealing. It is also a control, and in a loop a wrong route compounds into a costly action rather than a mislabel.
  • What does its drift look like? Teams learned to recognise fine-tuned classifier decay over twenty years. We do not yet know its failure patterns as our inputs and policies change.
  • And the one we cannot answer from inside: how would we know if our next instrument were lying to us the way the last four were?

If you have run one of these against a task where being wrong is expensive and silent, we would like to hear what happened, especially if it contradicts the above.

Start with a Governed-Design Review