Engineering evidenceSeptember 28, 2026

Can a decision-only model judge news relevance?

TypeSafe's Jev returns typed decisions with probabilities instead of text, in about a tenth of a second, for 4.9 cents per thousand judgements in this run. We compared excerpt-only calls with 3,169 retained judgements, using questions adapted from our relevance judge's prompt and agreement gates we wrote down first. This setup was fast and cheap, but did not meet the gates for either proposed role.

Not adopted · gates predeclared3,169 judgements · $0.155 · 277 s

Outcome

Two roles, two failed gates

As a gate ahead of the judge, it keeps 99% of related items only at a threshold that still lets 73.1% of unrelated ones through; the rule allows 50%. As a replacement judge for bulk re-judging, it agrees with the judge of record on 66.7% of relationship calls; the rule asks for 90%.

Recall at the operating threshold
99.65%
threshold 0.10; the floor is 99%
Unrelated items still passing
73.1%
F3 allows at most 50%
Relationship agreement
66.7%
B1 asks for at least 90%
Direction reversals
3.3%
B3 allows at most 2%

The question

Where a model that cannot write could still help

Every news article TickerTrac retains is judged once per saved thesis that holds the company. The judge returns typed fields, whether the item bears on the thesis, which way it moves it, how material and how direct it is, and a written reason the investor reads beside the finding. Because of that reason, a model that returns only decisions can never be the judge of record.

Two smaller jobs remained. A gate could skip the judge when an article plainly cannot bear on a thesis; a week earlier we had declined one because a gate call cost a meaningful fraction of a judgement, and wrote down the trigger for revisiting: a gate model that costs a small fraction of a judgement. And a backfill judge could re-judge thousands of retained stories for Historical Replay, where no reason is shown. Both turn on one measurable question: does it agree with what the judge concluded?

Method

Retained sample, adapted questions, gates written first

Sample

The balanced sample our earlier gate calibration prepares from the retained judgement corpus: 600 articles at least one thesis found related, 600 that every thesis found unrelated, 3,169 article-thesis judgements in all (1,131 related, 2,038 unrelated), positives from 72 tickers. Thesis text was recovered from an exact saved version for 2,564 judgements and reconstructed from the version effective at judgement time for 605.

The call

One request per judgement: ticker, headline, the stored 600-character excerpt and the thesis as state; four typed questions worded from the judge's own prompt. Related as a probability; relationship, materiality and directness as choices with the judge's definitions as criteria. Model typesafe/jev-1.13-20260917, through OpenRouter systemone.

Gates

For the gate role, the pre-filter gates we already used: recall at least 99% at the operating threshold, no more than 50% of unrelated items passing there, positives from at least 40 tickers, and a named human review of every dropped related item. For the backfill role, new ones: at least 90% relationship agreement, 80% materiality agreement, and no more than 2% direction reversals.

Comparison limits: Jev received the stored 600-character article excerpt, thesis text capped at 4,000 characters and no prior-evidence context. The production judge can read up to 6,000 article characters, the full supplied thesis and prior evidence. We compared with historical verdicts, without rerunning that judge on Jev's inputs. These results measure agreement for this setup; they do not isolate model capability or establish which model was correct.

Results · gate role

Real signal, not enough separation

Jev's related probability orders items well: at 0.5 it keeps two thirds of related items and passes under a tenth of unrelated ones. A gate, though, must keep 99% of related items, and the highest threshold that does so passes 73.1% of the unrelated. Our embedding pre-filter failed the same gate at 98.1%; this is much better and still a filter that removes only about a quarter of unrelated judgements while dropping some related evidence.

ThresholdRecall on relatedUnrelated still passing
0.05100.00%95.2%
0.10operating point99.65%73.1%
0.1597.97%56.4%
0.2095.58%42.4%
0.3088.59%24.0%
0.5067.37%9.5%
0.7042.35%2.4%

Brier score 0.1237 on this deliberately balanced sample. Between 0.3 and 0.9 the observed related rate runs five to twelve points above what Jev states, and below 0.1 it runs below (0.7% observed against 6.5% stated). 4 related judgements sit under the operating threshold; they are held for the named review, which no longer changes the decision.

Results · backfill judge

Where the retained verdicts differ

Field (702 related judgements)AgreementCohen's kappaWhere Jev's confidence ≥ 0.9Gate
relationship66.7%0.4689.8% (127 rows)fails (floor 90%)
materiality72.1%0.2591.4% (35 rows)fails (floor 80%)
directness45.3%0.0871.1% (128 rows)reported only

Jev read 104 of the judge's 382 supporting verdicts as neutral and 47 of 167 contradicting ones as neutral; it reversed direction 23 times (3.3%), 22 of them a contradiction read as support. It called 340 of 409 judgements labelled direct by the original judge indirect. The different inputs mean these counts alone cannot explain why the models disagreed.

Agreement is higher among calls with confidence of 0.9 or more: about nine in ten on relationship and materiality. That covers a fifth of the 702 related rows with stored labels for relationship and a twentieth for materiality. Agreement with those labels is useful for this decision; it is not an independent measure of correctness or calibrated confidence.

Cost and speed

Cheap, fast, and the whole run cost fifteen cents

Judgements
3,169
Provider charge
$0.155
Per judgement
$0.000049 · 7.3× below the incumbent
Input tokens
3,694,416
Latency p50 / p95
129.3 ms / 290.2 ms
Errors / retries
0 / 1

Decision

Not adopted, and what would reopen it

Nothing in production changes. This setup failed the predeclared gates, so we did not adopt it or build a third provider transport. The protocol, runner and content-free report preserve the result and its limits. The result supports declining this setup; it does not establish how the models compare on identical inputs.

Higher volume or a later model could make another comparison useful. A confident-only hybrid, taking Jev's verdict at 0.9 and sending the rest to the judge, remains a possible design. Any future claim about model capability would need comparable inputs and a control; this run establishes neither the hybrid's coverage across all judgements nor its correctness.

Publication boundary

Publish the numbers. Protect the text.

The report carries counts, rates, thresholds and identifiers only. The headlines, excerpts and thesis texts the calls read, and the list of dropped items held for review, stay on the evaluation host; excerpts are licensed and theses belong to their writers. The hashes bind this page to the committed artifacts and to the runner exactly as it ran.

Protocol
9257d4e7d1bc09962e1f555e8308cdaf2204a676a156c28af90cb0fccd24ced2
Freeze
98770f4478bb188059f4f4dfd79017aef3c7492cbfbf0110c0576f7d532bb2a2
Report
5078134b7ca95cd5339cd70d1e8066873352412db27eb01b5173f85c36cd5cc2
Runner
c1f295430f5fa436e0ffb1ae92712421ea27b3950e6caad0709fcce7885868b6