The numbers should match.
CrossCheck checks that they do.

An automated pipeline that extracts financial claims from earnings calls, investor decks, and SEC filings, then verifies they actually agree with each other.

Explore the claims → 96 real claims from 5 real companies, live from Postgres
Consistency snapshot live data
PPC · revenue.total $3,075,121,000 MATCH
DHR · revenue.total $4,343,100,000 DISCREPANT
GE · revenue.total $20,524,000,000 MATCH
ADTN · gross_profit $58,962,000 UNSCORED
41 consistent 8 unsupported
49 gold labels

UNSCORED = no hand-labeled reference exists yet for this claim, so neither checker has a verdict to compare.

5 companies analyzed, 5 real 2020 quarterly filings
PPC Pilgrim's Pride DHR Danaher GE General Electric TRV Travelers ADTN Adtran

Manual reconciliation doesn't scale.
So most of it never happens.

A public company's quarterly numbers appear in at least three independent places: the earnings call, the investor deck, and the SEC filing. They rarely get checked against each other, line by line.

RevenueFromContract... → $3,075,121,000
XBRL extraction
Structured facts pulled straight from SEC filings, matched by accession number, not by trusting the filing's own period labels.
page 15 word-position clustering
Deck PDF extraction
Financial tables parsed by real word position in the PDF, not simple text scraping, from actual investor-deck files.
Audio claim extraction
An LLM reads the real, imperfect speech-to-text transcript of the earnings call and pulls out the figures that were spoken aloud.
PPCDHRGETRVADTN
Multi-company coverage
The same pipeline, run across 5 companies and 4 fiscal quarters, Q1 through Q3 2020.

Two different ways to check.
They don't always agree either.

Every extracted claim is checked twice: once by a deterministic rule, once by an LLM judge. Running both is what actually surfaces disagreement.

Rule-based checker

Deterministic tolerance

Claims sharing a company, metric, and period are compared directly. No model call, no ambiguity in how the number was reached.

relative_diff = |a − b| / max(|a|, |b|)
≤ 0.5% → consistent
> 0.5% → discrepant
LLM judge

Claude Sonnet 5

Reads the same pair with full context, qualifiers included, and can recognize a spoken approximation a fixed tolerance can't.

Evidence traceability

Every claim is anchored

Each claim links back to exactly where it came from.

char_offset xbrl_fact page_bbox audio_ts
96 claims, 97.8% accuracy
on both checkers.

Real numbers from a live Postgres database at the time this page was generated, not projections.

5
real SEC filers, 4 fiscal quarters spanned (Q1-Q3 2020)
18
distinct financial metrics tracked, from revenue to combined ratio
$0.39
total real LLM spend across the whole project, cached after first run

GE's own SEC filing mislabels its fiscal quarter as 2019 Q3 instead of 2020 Q1. Found by reading the raw XBRL fact, not by trusting the filing's metadata.

Phase 8 · XBRL extraction

A CEO said sales grew to "$4.3 billion." The exact figure is $4,343.1M, 0.99% off. The rule-based checker flagged it discrepant. It shouldn't have.

Phase 8 · Consistency checker

Dense retrieval alone beat hybrid dense+BM25 search on this filing, the opposite of what the retrieval literature predicts.

Phase 1 · Text retrieval

Every claim, by company.
Start where the real disagreement is.

DHR's revenue.total is the DISCREPANT claim flagged above — this table is where it lives. Rows pairing a manual claim with an auto one are the same fact checked twice: a hand-labeled reference against an automated extraction, not a duplicate.

loading… ⇠ scroll for Source, Value, Method
Metric Source Value Method

Nine phases, each verified before the next.

00
Data foundation
5 companies, real filings, decks, XBRL, and audio downloaded and hand-labeled: 64 claims, 49 gold verdicts.
01
Text retrieval
189 filing chunks, dense vs. hybrid vs. reranked retrieval measured head to head.
02
Audio pipeline
VAD, ASR (13.0% WER), and speaker diarization (754 turns, 11 speakers) on the real earnings call.
03
XBRL extraction
Automated structured-fact extraction from SEC filings, matched by accession number.
04
Rule-based checker
A deterministic 0.5% tolerance rule compares claims across modalities.
05
LLM judge
Claude Sonnet 5 judges the same pairs independently, validated against the rule-based baseline.
06
Deck extraction
Investor-deck PDF tables parsed by real word position, not naive text scraping.
07
Audio extraction
An LLM reads the merged transcript and diarization output to extract spoken claims.
08
Multi-company scale-up
Every extractor and both checkers run across all 5 companies, surfacing the first real discrepancies.

96 claims. 2 real disagreements found. 1 live database.

Everything on this page traces back to a query against the same Postgres instance the pipeline writes to.

Explore the claims →