Geometric
Interpretability
Different minds, trained apart, build the same shape of meaning. We are charting it.
Convergence is not a coincidence. It is operational.
Language models built by different teams, on different data, independently arrive at the same relational geometry of concepts — recoverable from one model to another by a single rotation fit on unlabeled examples. And the shared geometry is not merely similar: a reading instrument built in one model, or a behavioral nudge, can be carried across the rotation and used in another at close to native accuracy in the target, with no labels there — a reading probe carried cross-family (Qwen2.5-3B → Llama-3.2-3B) scores AUROC 0.833 against a 0.839 native ceiling, on a 162-item real/fake battery our own surface floor marks DISQUALIFIED (A* 0.855; 1 of 18 cells separates after correction). One family pair, one battery — a candidate, not a law. The behavioral half is dearer: see the rank law below.
Mechanistic interpretability asks how a single model computes — its circuits, its features. Geometric interpretability asks a different question: what is invariant across minds, and what can you carry between them. It is the cartography of the space all these models share.
The wiring, and the map it navigates.
How one model computes
- Object — circuits, features, attention heads
- Scope — one model at a time
- Method — patching, sparse autoencoders, tracing
- Analogy — reverse-engineering one chip
What every model shares
- Object — the shape of the representation manifold
- Scope — across models, up to rotation
- Method — unsupervised alignment, transfer, null-adjudication
- Analogy — comparative anatomy · a reference map
They are complements, not rivals. One is the anatomy of a single instrument; the other is the territory all of them are trying to chart.
Two minds, one geometry.
Fit a single rotation on a handful of unlabeled parallel concepts — no task labels, no supervision — and the concept geometries of two independent models line up. The alignment sits near the top of the scale while two different "control" geometries sit at the floor.
The shape is not an artifact of the measuring tool. It survives a null that destroys the meaning and a null that mimics the geometry. What remains is shared.
Every model is a rotation of one canonical map.
If many minds share a geometry up to rotation, there is a single reference map of meaning they are all rotations of — a model-independent frame, with each model characterized by how it turns into it, plus the small residue that makes it itself.
We built that shared frame across nine models and asked the decisive question: does routing a finding through the one map cost anything versus the direct, pairwise route? It does not, at the coarse layer — on 180 held-out concept centroids the shared frame matches the direct pairwise maps to within noise (RSA delta +2×10⁻⁹), as does a ported honesty probe on a separate 76-entity set (AUROC delta +0.0003). Two limits: that comparison only ever ran on the aggregate layer, so fine per-concept structure is untested, not shown lossless; and the deployed hub route is measurably lossy — 0.455 against 0.833 for the direct pairwise map — so what follows is a claim about the frame we built, not the route we ship. One map replaces the tangle of pairwise translations. That is the artifact the field is built on.
Reading is cheap. Control is dear.
To read a concept across minds you need only a coarse alignment. To steer behavior across minds — to write — you need the alignment in full resolution. Drag the dial.
Aggregate sharp, instance ambiguous
You can carry what a concept is between minds. You cannot reconstruct one particular thought. Govern, don’t generate.
The resolution boundary
The shared map is low-resolution. Coarse instruments — centroids, concept directions, the honesty probe — port up-to-rotation. Fine per-feature structure does not (cross-model feature transfer ≈ 0.14, at the null floor). What’s universal is the coarse layer.
Watch cheap, control dear
So cross-model monitoring is robust and inexpensive; cross-model control is the costly, full-resolution regime — and only the coarse "is this grounded?" signal survives from outside a closed model.
What the geometry charts is identity, not truth. The shared field reads a model’s state — off-distribution, unfamiliar, off its map — not whether a claim is true. On confident, in-distribution fabrication the geometric read is blind, and inverted: a fluent fabrication sits more on-manifold than a real reasoning step (mean distance-to-manifold 0.729 vs 0.894 at layer 25). The 0.587 sometimes quoted here is a different measurement — a wrong-turn reasoning-error probe over 126 graded chains, only 26 of them incorrect, CI [0.455, 0.711] — and not a fabrication AUROC. That is by design: this is the cartography of meaning, not a truth oracle. Catching a fluent, well-grounded falsehood is a different instrument’s job.
Two languages that share no symbols — only a world.
The deepest question: is the shared geometry meaning, or just the fingerprint of overlapping training data? So we built two miniature languages with no symbols in common — only a shared structure of concepts — and trained a model on each, from nothing.
Their maps aligned — nearly as tightly as two models trained on the same corpus (94% of that ceiling, far above an anisotropy null, z +35). Then the control: keep every statistic identical but scramble which concept means what, and the alignment collapses to the floor. The aggregate geometry tracks meaning, not co-occurrence.
Confirmed at ~1B (two ~1.07B from-scratch decoders); the re-run was a same-seed determinism check on independent hardware, not a fresh-init replication: cross-mind RSA 0.794 (94% of the 0.85 same-corpus ceiling), the shared-geometry leg passes; the scramble control has now closed at ≥1B too — permuting which word names which concept, holding every co-occurrence statistic fixed, drops the alignment from 0.794 to 0.021 (collapse ratio 0.028 against a 0.35 bar), and recovering the exact authored ontology stays partial. Still labeled a conjecture below — we report it either way.
We don’t believe a shape until it survives two nulls.
One null erases the meaning. One mimics the geometry. A structure is real only if it outlives both — that is how we separate what’s truly shared from what the instrument invented. And what we can’t yet prove, we publish as a conjecture, plainly labeled.
A field, charted in the open.
A canonical map of meaning, shared across every language model we have measured. Whether it reaches vision, or other substrates, or one day a brain. Whether it is, in the end, meaning. We are pushing on all of it — and showing our work.