Engineering evidence ·

When better evidence changes the answer

An investment alert needs more than a plausible judgment. It needs the right source, a passage that supports the judgment, and a delivery path that respects it. Our quality investigation found failures in all three.

A 1,000-case study and targeted follow-ups led to source-specific model routing, safer citations and durable supporting evidence. They also led us to reject a prompt that looked cleaner but missed more useful alerts.

The error was bigger than the model

TickerTrac checks new evidence against an investor's saved Company thesis or Investment theme. In an initial audit, some judgments invented expectations, attributed a customer's economics to a supplier, or treated a conditional risk as a contradiction. Of 38 reviewable theme cases, eight had unsupported directions; another 22 selected cases had too little retained evidence to assess reliably. This was a sample chosen to find risks, not an estimate of the overall error rate.

A separate delivery defect had allowed 192 findings into accepted email receipts despite their saved decision not to interrupt the reader. The email loader was reconstructing permission to alert from relevance alone. Better model answers could not fix that path.

Earlier evaluations had checked judgments on supplied inputs. Their theme-specific follow-up covered only ten cases, including two clear positive alerts. They did not establish reliability across causal edge cases or test the complete path from saved judgment to email. Short retained excerpts also made later diagnosis harder.

A larger comparison changed the diagnosis

We compared 1,000 cases: 500 Company cases and 500 theme cases, using 800 retained news inputs and 200 constructed filing and transcript cases. Three model policies and multiple source lengths produced 3,600 comparisons. Source-first automated labels were reconciled before model results were revealed; they are not independent human ground truth.

Original study: useful alerts on scorable held-out cases. Each row has its own denominator.
SourceThen-current routeSonnet 5.5
News97 of 16352 of 163
Filings25 of 28; 3 invalid alerts17 of 28; 0 invalid alerts
Earnings transcripts28 of 9372 of 93

The transcript path was losing useful judgments at the citation boundary and producing weaker explanations. We moved retained transcripts to Sonnet 5.5. GPT-6 Luna stayed on news; replacing it with Sonnet would have cost more while missing more useful alerts in this comparison. Filings stayed on Haiku 4.5: the tested alternative traded fewer invalid alerts for lower useful recall.

A separate 52-pair filing analysis found that longer selected passages exposed additional useful facts but sometimes removed answerable material. It did not establish that wholesale passage replacement was better. Preserving existing facts while adding missing context remains the next candidate to test.

Preserve the evidence all the way to the reader

The delivery gate now honors the stored interruption decision. An append-only correction history preserves the original judgment and what was delivered. Sixteen stored records behind nine confirmed errors were corrected without sending new alerts. Future audits retain the bounded input and prior-evidence context actually shown to the model.

Review then found a subtler defect: cutting an exact quote to its first 350 characters could remove the decisive sentence while leaving its alert intact. The repaired selector requires complete sentences around a unique verbatim evidence anchor, including adjacent context. If that passage cannot fit, the judgment remains related but undirected and cannot alert. Exact wording is still not proof of a correct causal interpretation.

Replaying 840 saved replies made no provider calls. Prefix clipping admitted 476 alerts; both strict rejection and the new selector admitted 429. The selector recovered no additional real cases over rejection. Its recovery path is covered by controlled tests, not demonstrated as a recall gain here.

Article rereads now count attempts before dispatch, including failures and denied work, with separate limits of six fetch attempts and six judgment attempts per company per pass. An article-based direction or alert requires an exact supporting passage of at most 350 characters. That passage reaches the app and digests and survives pruning of the shorter-lived audit corpus, with its retrieval time, resolved source URL and judged-text hash. It is a retained excerpt, not an archive of the whole publisher page.

A cleaner prompt was not a better alert policy

We compared the original and an aligned transcript prompt on the same 140 sources, using the corrected parser and existing fallback policy. The original held-out slice has 126 cases, but these cases had already informed development. It is not fresh confirmation.

Paired transcript follow-up, including the incumbent fallback.
MeasureOriginalCandidate
Useful alerts76 of 9371 of 93
Correct retained directions82 of 10695 of 106
Invalid alerts on determinate labels00

The candidate improved directional judgments but lost five useful alerts, all in the Company slice, at essentially the same model cost. The clustered 95% interval for the useful-recall change was −12.5 to +1.1 percentage points. Only 13 cases called for quiet and 20 remained unscored, so zero invalid alerts does not establish a low production false-alert rate. We kept the original prompt.

More source text helped, with real qualifications

We selected 157 eligible cases whose short news inputs were related but had not produced an alert. Publisher retrieval returned 92 articles across 89 event clusters. Twenty-five sites refused access, 39 yielded too little text and one timed out. These failures stay in the denominator; inaccessible pages are not assumed to contain no useful facts.

All 32 proposed article alerts after source-label reconciliation. Unresolved alerts remain visible.
OutcomeCompanyThemeTotal
Supported useful alerts17421
Invalid alerts213
Unresolved alerts448

The article exposed 33 useful opportunities, versus five in the corresponding short inputs. The model recovered 21 of those 33, with a clustered 95% interval of 47.2%–79.3%. Across all 157 eligible cases, the observed useful yield was 21. End-to-end recall remains unknown because 65 pages were unavailable. The baseline produced no alerts by selection; this was not a random sample.

The evaluation itself needed correction. Initial labels scored 16 useful and 15 invalid alerts. Some required quantified financial impact for a concrete checkpoint, assumed a recap was already known despite empty prior context, or confused partial weakening with full invalidation. A second source-only pass covered every retrieved article and its short input, with candidate answers and first labels hidden. It produced the table above.

We specified that correction after seeing the initial outputs. These are reconciled results and sensitivity analysis, not a preregistered success. Twenty-eight sources remain unscored for interruption, including eight proposed alerts. Three false alerts among 31 quiet cases also leave meaningful noise. We retained the bounded article policy with stronger provenance; these results do not justify expanding it.

Quality and cost belong in the same decision

US dollars from returned usage and configured token prices. Research costs and recurring work are separate.
WorkObservedWithout cache savings
92 article second readings$0.0401$0.0478
Per 1,000 retrieved article assessments$0.436$0.520
Per 1,000 Sonnet transcript assessments, original study$9.734$13.635

The article judgment cost was about $0.00191 per confirmed additional useful alert, excluding the first reading, retrieval and research. Median fetch time was 0.715 seconds and median model time 9.43 seconds. Availability and eligibility in this selected sample do not support a monthly or per-user forecast.

The original 1,000-case study recorded $61.8319 in token-based cost plus $0.5750 reserved for uncertain outcomes. This follow-up admitted 898 requests: 897 completed and one remained uncertain. It recorded $17.5553 plus $0.072255 reserved. Most follow-up spending went to source labeling, reconciliation and output review, not the production-model article judgments. These are application-provider accounting figures, not settled invoices, and exclude the coding session.

What this establishes, and what remains open

The implemented repairs enforce stored alert decisions, preserve usable citations and bound failed work. They do not establish that every causal judgment is correct. Automated labels, reused cases, uncertain interruption thresholds and limited quiet controls constrain the quality claims. Publisher pages were fetched later; we did not independently verify that their contents were unchanged since publication.

The next useful comparison needs new quiet events, a calibrated interruption rubric and filing context that adds missing facts without discarding retained ones. We should measure useful recovery, unsupported claims, missed alerts and total cost together. Repeatedly relabeling this sample until it looks good would not answer those questions.

Download the aggregate follow-up results (JSON). The download includes all alert outcomes, both article label passes, cost accounting and a hash binding it to the internal aggregate receipt. It contains no private theses, account details, case identifiers or publisher text. A hash binds a version; it is not independent validation of the labels.