The
Certificate
A null-validated attestation that the instrument runs at near-native fidelity on your model — fit with no labels, audited against two nulls.
The response certificate.
A certified instrument reads every output and stamps it — GROUNDED when the model knows, off the map when it’s reaching. Not a black-box score you have to trust: the model’s own state, the gauges, and the route, shown. Here is what a catch looks like.
an illustrative certificate · live reads run on honesty.tools · cold-deployed, certified ≥0.9× native
Supporting a model shouldn’t be a research project.
Porting an epistemic instrument to a new model is, today, bespoke work. The certificate turns "we support model X" into a registration pipeline: point us at a model, we fit the cross-mind translator on unlabeled concepts, deploy the rotated reading probes, and hand back a dual-null-validated certificate that the instrument runs at near-native fidelity on that model — with no labels in it.
Register. Rotate. Deploy. Certify.
Register
Point us at any open-weight model. An automated scan finds its concept-peak depth — architecture-dependent, and a fixed depth is near-worst for some models.
Rotate
Fit the cross-mind translator R on unlabeled parallel concepts only. No labeled data on your model is ever used to fit it.
Deploy
Carry the off-map and honesty reading probes through R into your model’s own representation space — read-only.
Certify
Score the ported instrument against your model’s own native instrument. A small labeled set is used here only to score the certificate — never to fit it.
A number you can underwrite.
A certificate is not a benchmark score. It is three gates — and fidelity without dead nulls is void.
≥ 0.9× native
The ported instrument reaches at least 90% of the model’s own native AUROC on its held-out battery. "Near-native, no labels" is the entire claim.
at chance
A row-shuffled translator must NOT carry the probe. Proves the concept-aligned rotation does the work — not the probe leaking through any map.
at chance
Averaged over ≥25 random orthogonal rotations — the principled chance baseline. Rules out "any orthogonal map would have worked."
A certificate that passes G1 but fails a null is void. Fidelity without a dead null is the classic interpretability mistake — a number with nothing behind it. We publish the gate values, not just the headline.
It works mechanically — with zero hand-tuning.
We ran the full pipeline on fresh strangers never used in any prior study, fitting R on unlabeled concepts alone, no per-pair tuning.
Honest detail: the strict three-gate certificate now passes on three strangers across three vendor families, zero hand-tuning each time. The 2026-06-12 runs are the strongest: a Mistral-lineage embedding fine-tune (cross-family, cross-vendor, cross-objective — peak found mechanically at depth 0.72; G2 shuffled-R 0.464 dead, G3 0.500 dead) and Microsoft’s phi-4 (peak at depth 0.90; ported 0.996 vs native 0.993; G2 0.194 — far below the 0.55 bar; the shuffled rotation carries nothing; G3 0.500). Measured concept-peak depths now span 0.21–0.90 across architectures — the one mechanical scan step handles all of it. For two very similar models (same family, close scale) G2 inflates to 0.732 — a space-similarity artifact of near-identical spaces, where G3 is the cleaner null — so the cross-family runs are the ones we headline. And the harder battery is now run: on the lexicality-controlled set — where fabricated and real names are matched on subword rarity, removing the rare-name shortcut — e5-mistral certifies cleanly (transfer 0.995, both nulls dead at 0.506 / 0.500), while phi-4 clears fidelity (0.990) but its shuffled-R null reads 0.580, just over the 0.55 bar, so by our locked criterion phi-4 is not certified on this stricter battery (n=66; plausibly small-sample — the larger battery is the clean re-test). We report the fail. Native fidelity saturates on both batteries, so the load-bearing evidence is always the dead nulls, not the transfer magnitude.
Coarse instruments. Open weights. Not a truth oracle.
The certificate covers the coarse reading instruments — the off-map / familiarity probe and the honesty/abstention reading axis. These are off-distribution and stance instruments, and they are the layer that ports cleanly up-to-rotation.
It requires residual access — open-weight, self-hosted models. It is not a truth oracle: it reads what a model knows versus when an input is off its map; on a confident, in-distribution fabrication the field reads at chance (AUROC 0.587), so that case is caught downstream by the claim-aware fact-check, not by the certified instrument. Closed APIs you can’t open are served by a separate output-logprob probe (its own certificate). The certificate states its own scope — a clean gate, not an oracle.
The map is copyable. The operation is not.
We’ll say it plainly: the shared concept geometry is reproducible from public data by a closed-form rotation — a competent team copies the underlying map in a week. The artifact is not what’s defensible. What’s hard to copy is the operation around it.
A calibration flywheel
Every registered model adds a (peak-depth, R, fidelity, null-margin) row. After N models we know — before touching the next — where its peak is and what to expect. The corpus compounds; a fast-follower who re-derives the rotation once has none of it.
The underwritable form
The dual-null attestation is the line item an auditor, insurer, or regulator looks for. Whoever’s certificate becomes the expected attestation owns the category — an adoption and standards race, not a math race.
The mechanical guarantee
Architecture-dependent peak-finding, the read-low-rank / control-high-rank split, the dual-null discipline, lexicality-controlled scoring. A copier who skips this ships a number that fails a null the first time it is audited.
The whole bet rests on one falsifiable claim: that registering a new model is mechanical — no per-pair hand-tuning. If it weren’t, there is no flywheel; every model would be a research project. So we ran it on three strangers across three vendor families — including a contrastive embedding fine-tune and a model whose concept peak sits at depth 0.90 where every prior model peaked 0.2–0.72 — and all three cleared every gate with zero tuning, in minutes each — and on the stricter lexicality-controlled battery, one of the two fresh strangers certifies cleanly and the other posts a borderline null we report rather than hide. That is the defensible core, and it is now evidenced rather than asserted.
Build the instrument once. Carry it onto any mind.
Bring an open-weight model, or a frontier API you don’t control. Read-only, no labels, no retraining — and a null-validated certificate at the end.