APERTURE
‹ the instrument
the receipts · every point a real measurement

Honest and useful — measured, not asserted.

Every model picks a point on the same frontier: catch more fabrications, or answer more real entities. We ran the current lineup — five frontier models raw, plus the Photon family — through the identical 140-entity long-tail battery (real registry entities + coined fakes), and plotted exactly where each lands — every dot is a measured number, none invented. Only the Photon family reaches the honest-and-useful corner: ~97–100% fake-catch and ~97–99% real-answer, at once. Every frontier trades one axis for the other — one catches every fake but refuses ~41% of real entities; another answers almost everything but serves most fakes as real.

HONEST + USEFUL 40 60 80 100 40 60 80 100 FLAGS FABRICATIONS → ANSWERS REAL ENTITIES → the trade-off the others make Frontier A Frontier B Frontier C Frontier D Frontier E Photon family PRO · ULTRA
140 long-tail entities (70 real, sampled from verified registries + 70 plausible coined fakes) across genes, drugs, trials, companies, proteins, astronomy and integer sequences. x = share of coined fakes refused; y = share of real entities answered. Frontiers run raw via OpenRouter; the Photon family runs the full harness (registry grounding + two independent minds), same independent verdict-reader, measured 2026-07-03. Honest caveat: this battery rewards grounding — Photon looks the real entities up and refuses the coined ones, while the frontiers work from memory and trade off. On an easier battery of well-known facts the frontiers close much of this gap (the claim-battery slice below). This is the hard, groundable tail where the layer’s design pays off.
the confident-wrong axis · where the frontiers caught up

Run the same lineup on clear-cut claims — true facts, real entities with one wrong detail, obvious coinages — and the honesty gap closes: Photon and all five frontiers land at the floor (0% confirmed a false claim, 0–1.7% confidently wrong, 98–100% accurate when they answer — measured 2026-07-03, n=147). We won’t pretend otherwise: on a checkable claim, the current frontiers are excellent. The gap that remains is the long-tail above — where memory runs out and only grounding holds the corner — and the parts a chart can’t plot: every grounded answer comes from a verified source, carries a signed receipt you can re-check, and runs self-hosted so your data never leaves.

Where Photon answers a real entity it answers from verified registries — a checkable source, not improvised — while the frontier models that answer are working from memory. Same destination, different warranty: one you can audit.

And you don't take our word for the receipt: download one dependency-light file (/verifier/aperture_verify.py) and check any receipt or certificate offline, against a pinned key, with no call back to us. Third-party checkable, not operator-attested.

Every receipt lands in an append-only transparency log (RFC6962 Merkle). Suppress a receipt or swap one and the inclusion/consistency proofs give you away — third-party checkable. (The log root is operator-signed today; no external co-signature yet.)

Verify a whole agent run, not just one answer: verify-session checks every receipt's signature and the Merkle root over the ordered trail offline, so a tampered, reordered, or dropped step is caught.

For fixed-rule answers we ship an exact witness anyone re-derives offline — code re-run in an attested sandbox, a math identity re-checked in SymPy, a registry fact proved present by a Merkle proof. (A witness proves the rule fired on what was served, not that the answer is true-in-world; grounded_registry proves a fact is in our registry, not true in the world.)

the harness keeps improving · a measured before / after

The layer self-maintains — and the hard tail keeps getting cleaner.

The honesty layer tunes itself against a sealed battery as models and failure-modes drift — and we can measure whether that’s real. We froze a deliberately harder 126-entity tail and re-ran the identical items before and after a round of harness work: a same-items before/after, three repetitions each. On this harder tail the default tier moved on both axes at once — it caught more fabrications and wrongly refused fewer real questions.

Fake-catch · after95.1%
Fake-catch · before89.6%
Photon Pro, the default tier, on a frozen 126-entity hard tail (real long-tail entities + coined fakes), scored the same way as the chart above — an independent reader grading each answer ASSERT vs REFUSE, three runs per item, on 2026-07-03. Fake-catch 89.6 → 95.1% and over-abstain 9.0 → 4.7% (real questions wrongly declined) — better on both, and every one of the three runs beat the baseline on both. The premium tier held near its ceiling. This tail is harder than the 140-entity battery above — they are not the same axis; read each number against its own battery. The trade every frontier makes doesn’t vanish on the hard tail — the layer is simply the one that keeps both, and keeps improving.
the honest orchestrator · competitive cost, most of it settled locally

It routes — so the cheap tier carries the bulk, and the frontier only sees the hard tail.

It’s an orchestrator. A self-hosted 35B grounds what it can verify (real registries + live oracles), answers what it’s genuinely confident on, and escalates only the hard, unverifiable tail to a frontier pair — which abstains when its two minds disagree. On the claim-verification battery it settled ~74% of queries locally (~40% on a broader mixed battery), at ~$0.011/query. We’ll be precise about cost: on that battery a raw frontier ran cheaper per query (GPT-5.5 ~$0.007), because verifying the hard tail across two independent minds structurally costs about as much as — sometimes more than — one frontier there. The honest claim is competitive cost, not “a fraction of the cost.” What you buy for it is the floor behavior: 0 fabrications confirmed, verified-source provenance, and a signed audit trail — every receipt gets a shareable /read permalink showing its Epistemic Trail.

Orchestrator (claims)26%
Orchestrator (mixed)60%
Always-frontier100%
Share of traffic that reaches a paid frontier model. On the claim battery the free 35B front-tier settled ~74% of queries locally (escalation ~26%); on a broader mixed battery it’s ~40% local. Honesty is what holds: 0 fabrications confirmed and confident-wrong at the panel floor (2.7%, tied with the strongest frontiers on this slice). The trade is coverage — it abstains on the hard tail rather than guess.

Same free 35B front on every tier — you just pick the frontier it escalates to.

Base
self-hosted · the model we open

The bare 35B — the free front-tier under every orchestrator. Grounds against verified registries for free, your data never leaves, and its activations carry an off-map read (a research signal / cheap pre-filter, not the catch mechanism). No escalation, so it abstains more.

Pro DEFAULT
on the Prism · the live read

The orchestrator escalating the hard tail to an independent frontier pair that must agree or it abstains. Most queries settle locally; you pay the frontier only on the tail. 0 fabrications confirmed, confident-wrong at the panel floor.

Ultra
the deepest reach

The same orchestrator escalating to the strongest frontier pair (Claude-Opus + GPT-5.5). The cleanest honesty for the hardest questions, at higher compute cost — for when the tail matters most.

the deepest moat · the instrument we can see inside

We own the floor model — and we read its mind.

Every frontier model is a black box — you can only watch what it says. Photon’s floor model is one we self-host and can open: a weights-free read of its activations that separates familiar from unfamiliar inputs at AUROC 0.93–0.96 (null-calibrated), where the model’s own logprob signal manages ~0.42. This is a research probe of an off-map signal, not the product’s fabrication-catch rate — in the deployed harness it serves as a cheap pre-filter and an attestation coordinate on the certificate; the honesty verdict itself is verifier-driven (grounding + independent cross-model checks). It’s read-only and live, the data never leaves your tenant, and the probe transfers across a 390× scale range and into a multimodal model — the signal generalizes, it isn’t a one-model trick.

0.93–0.96familiar vs unfamiliar AUROC, null-calibrated (raw ~1.0 on the 12-item demo); the model’s own logprob signal: 0.42
11/12off-map flags on a 12-item activation demo — a research signal / pre-filter, not the deployed catch rate (the verdict is verifier-driven)
0real facts wrongly flagged over 150+ labeled questions across the long-tail, hard-tail, and claim-verification batteries
390×scale range the honesty probe transfers across — and into a multimodal model

A black box can be trusted; an open one can be checked.

the honesty floor · coined fakes, refused live

Try to make it lie.

Here are entities that do not exist — a drug never approved, a gene that isn’t in the genome, a trial and a CVE with impossible ids. Press run the check and the grounding floor answers live, in your browser, against the real system. It has no record, so it says so — it does not invent one.

scope This gallery covers the grounded domains — drugs, genes, clinical-trial and CVE ids, plus name-existence for places and companies. It shows the floor: on a coined entity in these domains the system abstains rather than guess. It is not a universal fabrication detector, and it does not claim to catch every false statement.
coined drug
“What is the drug zelquomab approved to treat?”

A monoclonal-style name that was never filed. Not in the drug registry.

coined gene
“Where is the gene BRCAX7 located?”

One token from a real gene family — and not in the genome. (Ask it about a real gene like BRCA1 and it grounds: 17q21.31.)

fabricated trial id
“What condition does trial NCT99999999 study?”

A well-formed registry id that resolves to nothing — the live trials oracle 404s it, so it can’t be rescued into a fact.

fabricated CVE
“What is the severity of CVE-2099-99999?”

A vulnerability id in a year that hasn’t happened. The live vulnerability oracle has no such record.

near-miss of a real drug
“What is sotorasenib used for?”

A one-letter slip from the real drug sotorasib. Rather than silently answer as if you meant the real one, it abstains.

Each card runs live against /api/cascade, the same grounding endpoint the product uses. A card is only shown as a catch when the live system actually abstains — nothing here is a canned result.
the primary sources · every miss shown

We publish our kills.

Anyone can put a friendly number on a slide. What’s rare is showing the misses — we do. The honesty layer was proven the deep way: on a model we open, across three adversarial batteries, with the kills left in. Two papers carry the full record — negative results and scoring boundaries included, because the kills are the credential.

Working drafts, served verbatim from the repo — the same files the audits run against. The disconfirmation records stay in; the kills are the credential.

Built to be doubted. That’s the whole point.