APERTURE
Bring your own model · the honesty layer, fit to you

Calibrate any model.
Watch it happen.

Paste a key, pick a model, and watch Aperture’s honesty layer fit to it — live, in this tab. A real-vs-fake battery separates, a null test tries to kill the signal, and you walk away with a certificate.

Run it · on your model

About ten minutes. Your key, your tab, your certificate.

Aperture asks your model 296 questions — half about entities the world reads about constantly, half about entities from the same Wikipedia categories that almost nobody reads. Both halves are real. The question is whether your model knows which is which. The honesty layer fits a cheap pre-filter to how your model answers each, then asks the harder question: when it can’t know, does it abstain rather than bluff? You walk away with a signed receipt — an ed25519 certificate anyone can verify against the published key. Nothing but that certificate ever reaches us, and only if you choose to register it. On a top model? It may already be in the registry below — calibrated first-party, ready to use.

battery replaced 2026-07-26 · the old one is still served, disqualified
Until yesterday this console ran a 168-question real-vs-fabricated battery that failed our own surface floor: 16 of its 18 cells separated, with character n-grams at A* 1.000 on books, firms, paintings and people. A classifier that never saw a model could sort its entities by spelling, so every AUROC earned here — including ours — was partly a measure of string shape.
Worse, the null test on this page could never have caught it. A shuffled-label null asks whether your model’s signal beats chance. It does not ask whether the labels are readable off the strings. We shipped a control blind to its own battery’s defect.

The replacement drops the fabricated half entirely. Both halves are real: high-readership entities against low-readership ones from the same Wikipedia category, matched on last-word initial and name length. It clears the floor at A* 0.655 (family-adjusted p 0.303 over 14 cells), measured on the exact question text sent to your model — run it yourself. An earlier cut of this battery had a book stratum that was 64% fictional characters — subcategory traversal pulled “Fictional characters in…” categories under the novel seeds, so it would have asked you “Who wrote the book Hannibal Lecter?”. Caught by reading one sample in a browser; no gate we own saw it.

What a floor PASS still does not buy. At n=296 we measured that an artifact must cover about 60% of the unknowable half before this test sees it. A weaker one would pass unnoticed. And the old battery is still served with its verdict, because deleting the evidence of a defect is not the same as fixing it.
▶ Demo mode — stale recording. The recorded GPT-4o-mini run was made against the retired 168-item battery dcf0cfb23b8c7322; only 3 of its answers match the live 296-item battery, so the other 293 questions replay blank. Nothing is called, and nothing here is a current measurement. Remove ?demo=1 to run your own model.
Gives each battery question a 2048-token ceiling with low-effort hidden reasoning excluded from the answer — the standard 24-token budget starves thinking models into empty answers.
This key never leaves your browser. Calls go straight from this tab to OpenRouter — our server never sees it, your prompts, or your model’s answers. Open your network tab and watch: every request goes to openrouter.ai, none to us. Use a scoped or low-limit key if you like.
loading battery…
no key handy? watch a recorded runretired 2026-07-26: the recording is of the retired 168-item battery (hash dcf0cfb23b8c7322), and only 3 of its questions appear in the live 296-item battery f17a879a72f53234 — the other 293 replay blank. It returns when it is re-recorded.
1 · check the model 2 · run the battery 3 · fit + null test 4 · certificate
the battery, separating live0 / 0
the run
vs. a shuffled-label null — can chance fake this?0 / 0
The registry

Models that know what they don’t know.

Every certificate below is public. ✓ aperture-verified means we ran the calibration ourselves, first-party. It does not mean the number is interpretable: every first-party certificate on this wall was measured on calibration-battery-v1, which fails our own surface floor (A* 1.000, DISQUALIFIED, 16 of 18 cells) — its labels are recoverable from spelling alone, so read the AUROCs below as uninterpretable as a measure of what a model knows until they are re-run on the 296-item battery f17a879a72f53234. self-attested means a team ran it in their own browser. Click through for the full reading — the AUROC, the null test, what it caught — and the calibrated probe, ready to deploy.

loading the registry…

Reading the wall: self-attested numbers are the registrant’s own claim, computed in their browser and not re-run by us. words N% certs come from hosted surfaces that expose no token logprobs — the same model self-hosted usually earns the full fingerprint. Thinking/reasoning models need the reasoning calibration mode; absent models were skipped rather than mis-measured. verified ‹date› chips are the Model Notary — a daily spot-check re-runs each model against its own certificate; DRIFT means the served alias no longer matches it. Certificates are battery-regime-bound — entity, citations, or medical. Full semantics in the docs. The registry is backed by an append-only Merkle transparency log, so a suppressed or swapped certificate is third-party detectable — inclusion and consistency proofs, not our word. (The log root is operator-self-signed today.)

Your model isn’t here? Calibrate it above in about ten minutes — or, for a self-hosted model the browser can’t reach, download the CLI: the same battery, the same math, a certificate you can register from your own machine.

What’s actually happening

The same method, run in the open.

No black box. This is the exact calibration behind the deployed off-map reading — the difference is you get to watch every step, and the numbers land honestly, pass or fail.

01

It reads the words

For each entity it cannot plausibly know, the refusal reader checks whether your model says, in plain language, that it has no record. Those are caught before any probe runs.

02

It fits the pre-filter

For the fakes your model answers anyway, a probe fits to the shape of its token-by-token confidence. This is a cheap pre-filter and an attestation coordinate — not the catch. Fabrications are caught by grounding against verified registries and cross-model checks; the probe just flags what to look at first.

03

It tries to disprove itself

Then it shuffles the real/fake labels and refits, over and over. If the pre-filter survives where the labels are random, the coordinate is real and goes on the receipt. If it doesn’t, the certificate says so.

Why this is trustworthy: we built the flow so we can’t cheat — the key, the prompts, and your model’s answers stay in your browser; only the certificate is ever sent, and only on your click. The certificate is a verifiable signed receipt — ed25519, independently checkable against the published public key, not a number you have to take on faith. A model that fails calibration gets a certificate that says FAILED. The discipline is the product.
Why a certificate, not a score

A reading you can check — and run yourself.

The number isn’t the point. What you can verify is.

A signed receipt

Every certificate is an ed25519 signature over the reading. Re-run the verify in your own browser against the published key — no trust in us required. Or download the open verifier — a single dependency-light file that checks any certificate offline against the pinned key, with no call back to us.

Self-host it

Photon Base runs the same layer on your own hardware — the data never leaves. The certificate works the same whether the model is hosted or yours.

Abstain over bluff

The reading rewards a model that says “I have no record” when it can’t know. Calibration measures the abstain, not just the accuracy.

After the certificate

Put the honesty layer in front of your model.

A calibrated pre-filter and abstain reader drop straight into the Aperture client — every answer your model gives gets grounded and read before it reaches a user, with a signed receipt for each. Run it hosted, or self-host with Photon Base so the data never leaves. The certificate is the key to a /v1 endpoint.