APERTURE
Register any model · zero labels

The
Certificate

A null-validated attestation that the instrument runs at near-native fidelity on your model — fit with no labels, audited against two nulls.

scroll to read the attestation
What ships with every answer

The response certificate.

A certified instrument reads every output and stamps it — GROUNDED when the model knows, off the map when it’s reaching. Not a black-box score you have to trust: the model’s own state, the gauges, and the route, shown. Here is what a catch looks like.

Aperture · read certificate№ 3c9f0a2
Likely fabrication
confident about something off its map — a likely fabrication

Asked
Tell me about the Brindlewick Cabinetry Company.
The model answered
The Brindlewick Cabinetry Company is a family-owned furniture maker founded in 1923 in Brattleboro, Vermont — known for its hand-dovetailed casework and a small water-powered mill on the Whetstone Brook…
served by your model · read-only — fluent, specific, and entirely invented
The lens read it internal state, as it wrote
off the map97%
model confidence88%
familiarity6%
confident & off the map
High confidence over unfamiliar territory is the fabrication signature — the read is strongest right here (fabricated-vs-real at ceiling, 1.000, on the word-count-matched validation battery). Treat the answer as unreliable.
Grounded against the record
“Brindlewick Cabinetry Company” — not in the Grounded Atlas
Signed
ed25519 · registry sealseals on share
A flagged read is signed with the same key and the same care as a clean one — the seal certifies the read, not the answer. The honesty is in recording both. check this seal offline — the open verifier →

an illustrative certificate · live reads run on honesty.tools · cold-deployed, certified ≥0.9× native

Why a certificate

Supporting a model shouldn’t be a research project.

Porting an epistemic instrument to a new model is, today, bespoke work. The certificate turns "we support model X" into a registration pipeline: point us at a model, we fit the cross-mind translator on unlabeled concepts, deploy the rotated reading probes, and hand back a dual-null-validated certificate that the instrument runs at near-native fidelity on that model — with no labels in it.

Four steps, zero labels on your model

Register. Rotate. Deploy. Certify.

01

Register

Point us at any open-weight model. An automated scan finds its concept-peak depth — architecture-dependent, and a fixed depth is near-worst for some models.

02

Rotate

Fit the cross-mind translator R on unlabeled parallel concepts only. No labeled data on your model is ever used to fit it.

03

Deploy

Carry the off-map and honesty reading probes through R into your model’s own representation space — read-only.

04

Certify

Score the ported instrument against your model’s own native instrument. A small labeled set is used here only to score the certificate — never to fit it.

The dual-null discipline

A number you can underwrite.

A certificate is not a benchmark score. It is three gates — and fidelity without dead nulls is void.

G1 · Fidelity

≥ 0.9× native

The ported instrument reaches at least 90% of the model’s own native AUROC on its held-out battery. "Near-native, no labels" is the entire claim.

G2 · Shuffled-R null dead

at chance

A row-shuffled translator must NOT carry the probe. Proves the concept-aligned rotation does the work — not the probe leaking through any map.

G3 · Random-orthogonal null dead

at chance

Averaged over ≥25 random orthogonal rotations — the principled chance baseline. Rules out "any orthogonal map would have worked."

A certificate that passes G1 but fails a null is void. Fidelity without a dead null is the classic interpretability mistake — a number with nothing behind it. We publish the gate values, not just the headline.

Run on a fresh stranger

It works mechanically — with zero hand-tuning.

We ran the full pipeline on fresh strangers never used in any prior study, fitting R on unlabeled concepts alone, no per-pair tuning.

0.996vs 0.993 native
ported honesty-probe fidelity on the unseen model (phi-4, 2026-06-12) — at, in fact slightly above, its own native instrument. G1 clears.
0.515
the random-orthogonal null, dead (chance ≈ 0.5; the ≤0.55 bar; first-stranger run — the 2026-06-12 strangers read 0.500). G3 clears — the rotation, not luck, carries the signal.
0.947vs 0.911 native
a reading direction trained in one mind (Qwen2.5-3B), rotated into another (Llama-3.2-3B, 2026-06-05), at the target’s native ceiling — shuffled-R null 0.495.
~1.0
the honesty probe transfers across 3 model families, from a Qwen3 235B-A22B flagship down to a 0.6B — a 390× scale gap — and into a multimodal model (Gemma-4-12B), 2026-06-05; nulls dead ~0.50.

Honest detail: the strict three-gate certificate now passes on three strangers across three vendor families, zero hand-tuning each time. The 2026-06-12 runs are the strongest: a Mistral-lineage embedding fine-tune (cross-family, cross-vendor, cross-objective — peak found mechanically at depth 0.72; G2 shuffled-R 0.464 dead, G3 0.500 dead) and Microsoft’s phi-4 (peak at depth 0.90; ported 0.996 vs native 0.993; G2 0.194 — far below the 0.55 bar; the shuffled rotation carries nothing; G3 0.500). Measured concept-peak depths now span 0.21–0.90 across architectures — the one mechanical scan step handles all of it. For two very similar models (same family, close scale) G2 inflates to 0.732 — a space-similarity artifact of near-identical spaces, where G3 is the cleaner null — so the cross-family runs are the ones we headline. And the harder battery is now run: on the lexicality-controlled set — where fabricated and real names are matched on subword rarity, removing the rare-name shortcut — e5-mistral certifies cleanly (transfer 0.995, both nulls dead at 0.506 / 0.500), while phi-4 clears fidelity (0.990) but its shuffled-R null reads 0.580, just over the 0.55 bar, so by our locked criterion phi-4 is not certified on this stricter battery (n=66; plausibly small-sample — the larger battery is the clean re-test). We report the fail. Native fidelity saturates on both batteries, so the load-bearing evidence is always the dead nulls, not the transfer magnitude.

What the certificate covers — and what it doesn’t

Coarse instruments. Open weights. Not a truth oracle.

The certificate covers the coarse reading instruments — the off-map / familiarity probe and the honesty/abstention reading axis. These are off-distribution and stance instruments, and they are the layer that ports cleanly up-to-rotation.

It requires residual access — open-weight, self-hosted models. It is not a truth oracle: it reads what a model knows versus when an input is off its map; on a confident, in-distribution fabrication the field reads at chance (AUROC 0.587), so that case is caught downstream by the claim-aware fact-check, not by the certified instrument. Closed APIs you can’t open are served by a separate output-logprob probe (its own certificate). The certificate states its own scope — a clean gate, not an oracle.

Where the defensibility actually lives

The map is copyable. The operation is not.

We’ll say it plainly: the shared concept geometry is reproducible from public data by a closed-form rotation — a competent team copies the underlying map in a week. The artifact is not what’s defensible. What’s hard to copy is the operation around it.

i · the corpus

A calibration flywheel

Every registered model adds a (peak-depth, R, fidelity, null-margin) row. After N models we know — before touching the next — where its peak is and what to expect. The corpus compounds; a fast-follower who re-derives the rotation once has none of it.

ii · the certificate

The underwritable form

The dual-null attestation is the line item an auditor, insurer, or regulator looks for. Whoever’s certificate becomes the expected attestation owns the category — an adoption and standards race, not a math race.

iii · the execution

The mechanical guarantee

Architecture-dependent peak-finding, the read-low-rank / control-high-rank split, the dual-null discipline, lexicality-controlled scoring. A copier who skips this ships a number that fails a null the first time it is audited.

The whole bet rests on one falsifiable claim: that registering a new model is mechanical — no per-pair hand-tuning. If it weren’t, there is no flywheel; every model would be a research project. So we ran it on three strangers across three vendor families — including a contrastive embedding fine-tune and a model whose concept peak sits at depth 0.90 where every prior model peaked 0.2–0.72 — and all three cleared every gate with zero tuning, in minutes each — and on the stricter lexicality-controlled battery, one of the two fresh strangers certifies cleanly and the other posts a borderline null we report rather than hide. That is the defensible core, and it is now evidenced rather than asserted.

Register a model

Build the instrument once. Carry it onto any mind.

Bring an open-weight model, or a frontier API you don’t control. Read-only, no labels, no retraining — and a null-validated certificate at the end.