Can a decision-only model judge news relevance?
TypeSafe's Jev returns typed decisions with probabilities instead of text, in about a tenth of a second, for 4.9 cents per thousand judgements in this run. We compared excerpt-only calls with 3,169 retained judgements, using questions adapted from our relevance judge's prompt and agreement gates we wrote down first. This setup was fast and cheap, but did not meet the gates for either proposed role.
Outcome
Two roles, two failed gates
As a gate ahead of the judge, it keeps 99% of related items only at a threshold that still lets 73.1% of unrelated ones through; the rule allows 50%. As a replacement judge for bulk re-judging, it agrees with the judge of record on 66.7% of relationship calls; the rule asks for 90%.
- Recall at the operating threshold
- 99.65%
- threshold 0.10; the floor is 99%
- Unrelated items still passing
- 73.1%
- F3 allows at most 50%
- Relationship agreement
- 66.7%
- B1 asks for at least 90%
- Direction reversals
- 3.3%
- B3 allows at most 2%
The question
Where a model that cannot write could still help
Every news article TickerTrac retains is judged once per saved thesis that holds the company. The judge returns typed fields, whether the item bears on the thesis, which way it moves it, how material and how direct it is, and a written reason the investor reads beside the finding. Because of that reason, a model that returns only decisions can never be the judge of record.
Two smaller jobs remained. A gate could skip the judge when an article plainly cannot bear on a thesis; a week earlier we had declined one because a gate call cost a meaningful fraction of a judgement, and wrote down the trigger for revisiting: a gate model that costs a small fraction of a judgement. And a backfill judge could re-judge thousands of retained stories for Historical Replay, where no reason is shown. Both turn on one measurable question: does it agree with what the judge concluded?
Method
Retained sample, adapted questions, gates written first
Sample
The balanced sample our earlier gate calibration prepares from the retained judgement corpus: 600 articles at least one thesis found related, 600 that every thesis found unrelated, 3,169 article-thesis judgements in all (1,131 related, 2,038 unrelated), positives from 72 tickers. Thesis text was recovered from an exact saved version for 2,564 judgements and reconstructed from the version effective at judgement time for 605.
The call
One request per judgement: ticker, headline, the stored 600-character excerpt and the thesis as state; four typed questions worded from the judge's own prompt. Related as a probability; relationship, materiality and directness as choices with the judge's definitions as criteria. Model typesafe/jev-1.13-20260917, through OpenRouter systemone.
Gates
For the gate role, the pre-filter gates we already used: recall at least 99% at the operating threshold, no more than 50% of unrelated items passing there, positives from at least 40 tickers, and a named human review of every dropped related item. For the backfill role, new ones: at least 90% relationship agreement, 80% materiality agreement, and no more than 2% direction reversals.
Comparison limits: Jev received the stored 600-character article excerpt, thesis text capped at 4,000 characters and no prior-evidence context. The production judge can read up to 6,000 article characters, the full supplied thesis and prior evidence. We compared with historical verdicts, without rerunning that judge on Jev's inputs. These results measure agreement for this setup; they do not isolate model capability or establish which model was correct.
Results · gate role
Real signal, not enough separation
Jev's related probability orders items well: at 0.5 it keeps two thirds of related items and passes under a tenth of unrelated ones. A gate, though, must keep 99% of related items, and the highest threshold that does so passes 73.1% of the unrelated. Our embedding pre-filter failed the same gate at 98.1%; this is much better and still a filter that removes only about a quarter of unrelated judgements while dropping some related evidence.
| Threshold | Recall on related | Unrelated still passing |
|---|---|---|
| 0.05 | 100.00% | 95.2% |
| 0.10operating point | 99.65% | 73.1% |
| 0.15 | 97.97% | 56.4% |
| 0.20 | 95.58% | 42.4% |
| 0.30 | 88.59% | 24.0% |
| 0.50 | 67.37% | 9.5% |
| 0.70 | 42.35% | 2.4% |
Brier score 0.1237 on this deliberately balanced sample. Between 0.3 and 0.9 the observed related rate runs five to twelve points above what Jev states, and below 0.1 it runs below (0.7% observed against 6.5% stated). 4 related judgements sit under the operating threshold; they are held for the named review, which no longer changes the decision.
Results · backfill judge
Where the retained verdicts differ
| Field (702 related judgements) | Agreement | Cohen's kappa | Where Jev's confidence ≥ 0.9 | Gate |
|---|---|---|---|---|
| relationship | 66.7% | 0.46 | 89.8% (127 rows) | fails (floor 90%) |
| materiality | 72.1% | 0.25 | 91.4% (35 rows) | fails (floor 80%) |
| directness | 45.3% | 0.08 | 71.1% (128 rows) | reported only |
Jev read 104 of the judge's 382 supporting verdicts as neutral and 47 of 167 contradicting ones as neutral; it reversed direction 23 times (3.3%), 22 of them a contradiction read as support. It called 340 of 409 judgements labelled direct by the original judge indirect. The different inputs mean these counts alone cannot explain why the models disagreed.
Agreement is higher among calls with confidence of 0.9 or more: about nine in ten on relationship and materiality. That covers a fifth of the 702 related rows with stored labels for relationship and a twentieth for materiality. Agreement with those labels is useful for this decision; it is not an independent measure of correctness or calibrated confidence.
Cost and speed
Cheap, fast, and the whole run cost fifteen cents
- Judgements
- 3,169
- Provider charge
- $0.155
- Per judgement
- $0.000049 · 7.3× below the incumbent
- Input tokens
- 3,694,416
- Latency p50 / p95
- 129.3 ms / 290.2 ms
- Errors / retries
- 0 / 1
Decision
Not adopted, and what would reopen it
Nothing in production changes. This setup failed the predeclared gates, so we did not adopt it or build a third provider transport. The protocol, runner and content-free report preserve the result and its limits. The result supports declining this setup; it does not establish how the models compare on identical inputs.
Higher volume or a later model could make another comparison useful. A confident-only hybrid, taking Jev's verdict at 0.9 and sending the rest to the judge, remains a possible design. Any future claim about model capability would need comparable inputs and a control; this run establishes neither the hybrid's coverage across all judgements nor its correctness.
Publication boundary
Publish the numbers. Protect the text.
The report carries counts, rates, thresholds and identifiers only. The headlines, excerpts and thesis texts the calls read, and the list of dropped items held for review, stay on the evaluation host; excerpts are licensed and theses belong to their writers. The hashes bind this page to the committed artifacts and to the runner exactly as it ran.
- Protocol
- 9257d4e7d1bc09962e1f555e8308cdaf2204a676a156c28af90cb0fccd24ced2
- Freeze
- 98770f4478bb188059f4f4dfd79017aef3c7492cbfbf0110c0576f7d532bb2a2
- Report
- 5078134b7ca95cd5339cd70d1e8066873352412db27eb01b5173f85c36cd5cc2
- Runner
- c1f295430f5fa436e0ffb1ae92712421ea27b3950e6caad0709fcce7885868b6