THE TONY B. FILES / CASE 02

Trust as code

0wav

Verify an answer once. Every re-ask after that is a read, not a re-roll.

THE BRIEF

An answer becomes an asset: stored with its provenance, checked independently, then recalled byte for byte. Disagreement goes to human review.

0wav is an asset-side determinism layer for multimodal AI: speech recognition, diarization, prosody, and encoder features, stored so that a verified answer is locked byte-exact with its provenance.

The verifier is never the model's own confidence. Five transcription engines from four vendor families run the same audio; where all five agree, the word locks. Where they argue, it goes to a human, and the ruling is hash-bound to the exact queue entry.

ON THE RECORD

PORTFOLIO SNAPSHOT / JUL 2026

Figures and source references supplied with this portfolio. Scope and limitations accompany each result.

Read speech, certified

≥99%

of accepted words correct, pre-registered

9 of the 12 misses were archaic spellings in the gold transcript.

Source & basis

Reported measurement · asr_gold_gate_alpha001_2026-07-12.json — 12/1,952, CP-95 upper bound 0.009941

Financial QA, certified

≥92.2%

of accepted answers correct at 90% confidence

Effective n ≈ the unique underlying questions. The certificate covers this suite's query mix.

Source & basis

Reported measurement · metaharness_descent2_pooled_2026-07-18.json — n=2,271 pooled, 21.6% coverage

Byte-exact replay

20 / 20

identical SHA across live runs

Source & basis

Reported measurement · T3a replay lane

Batch-shape drift

63.9%

of cells changed answer based on who shared the server

Stated as a loss on purpose. It is the industry default this work replaces.

Source & basis

Reported measurement · inference_batch_card.json — B ∈ {2,4,8}

Receipt chain

×36 intact

every iteration re-verifies from its own stored bytes

Source & basis

Reported measurement · loop-contract trail, ADR-043

Energy per re-ask

~100,000×

less than regenerating the answer

Declared constants, not a metered measurement. Shown with CO₂e beside cost.

Source & basis

Declared assumption · harness constants — 1.0 Wh fresh pass vs 0.00001 Wh recall read

Routing accuracy @1

FinanceBench · 144 docs · 150 questions
LLM alone9.6
Lexical52.7
RRF58.0
Descent61.3
Refiner65.3
Arbiter72.7

Handing the model the whole document tree scores 9.6. Deterministic structure alone beats it by 43 points. The model is used at the last inch only — one bounded call per query — and it adds the final 7.4.

CARD · multi_doc_financebench_hybrid_llm_2026-07-06.json

Word-level trust tiers

5 engines · 4 vendor families
All agree99.8%
They argue55%

Five stenographers from four different companies transcribe the same tape. Where all five write the same word it matched human gold 998 times in 1,000 — a 216× separation from the words they argue about. Those go to a human, and the ruling is hash-bound to the exact queue entry.

CARD · asr_gold_gate_2026-07-11 · 1 miss in 479, and it was the gold's archaic spelling

Same question, same settings

The only variable is who shares the server
0.0%Run alone · 20 receipted runs
63.9%Batched with strangers

Ask a cloud model the same question twice and you can get different answers — not because it learned anything, but because of which other requests shared the batch. Nearly two-thirds of outputs shifted. A locked answer cannot be touched by that.

CARD · inference_batch_card.json (T3b) · T3a replay 20/20

Applied MLConformal predictionProvenanceEdge AI

MAKE YOURSELF AT HOME

Accessibility

Appearance
Noir soundtrack

Quiet, original music. Starts with your click.

Reduce motion

Stop camera movement and transitions.

Larger text

Increase reading text in stories and files.

Visual preferences stay on this device. Sound starts off on each visit.