Geometric
Interpretability
Different minds, trained apart, build the same shape of meaning. We are charting it.
Convergence is not a coincidence. It is operational.
Language models built by different teams, on different data, independently arrive at the same relational geometry of concepts — recoverable from one model to another by a single rotation fit on unlabeled examples. And the shared geometry is not merely similar: a reading instrument built in one model, or a behavioral nudge, can be carried across the rotation and used in another at native fidelity, with no labels in the target.
Mechanistic interpretability asks how a single model computes — its circuits, its features. Geometric interpretability asks a different question: what is invariant across minds, and what can you carry between them. It is the cartography of the space all these models share.
The wiring, and the map it navigates.
How one model computes
- Object — circuits, features, attention heads
- Scope — one model at a time
- Method — patching, sparse autoencoders, tracing
- Analogy — reverse-engineering one chip
What every model shares
- Object — the shape of the representation manifold
- Scope — across models, up to rotation
- Method — unsupervised alignment, transfer, null-adjudication
- Analogy — comparative anatomy · a reference map
They are complements, not rivals. One is the anatomy of a single instrument; the other is the territory all of them are trying to chart.
Two minds, one geometry.
Fit a single rotation on a handful of unlabeled parallel concepts — no task labels, no supervision — and the concept geometries of two independent models line up. The alignment sits near the top of the scale while two different "control" geometries sit at the floor.
The shape is not an artifact of the measuring tool. It survives a null that destroys the meaning and a null that mimics the geometry. What remains is shared.
Every model is a rotation of one canonical map.
If many minds share a geometry up to rotation, there is a single reference map of meaning they are all rotations of — a model-independent frame, with each model characterized by how it turns into it, plus the small residue that makes it itself.
We built that shared frame across nine models and asked the decisive question: does routing a finding through the one map cost anything versus the direct, pairwise route? It does not — the shared map matches every pairwise map to within numerical noise. One map replaces the tangle of pairwise translations. That is the artifact the field is built on.
Reading is cheap. Control is dear.
To read a concept across minds you need only a coarse alignment. To steer behavior across minds — to write — you need the alignment in full resolution. Drag the dial.
Aggregate sharp, instance ambiguous
You can carry what a concept is between minds. You cannot reconstruct one particular thought. Govern, don’t generate.
The resolution boundary
The shared map is low-resolution. Coarse instruments — centroids, concept directions, the honesty probe — port up-to-rotation. Fine per-feature structure does not (cross-model feature transfer ≈ 0.14, at the null floor). What’s universal is the coarse layer.
Watch cheap, control dear
So cross-model monitoring is robust and inexpensive; cross-model control is the costly, full-resolution regime — and only the coarse "is this grounded?" signal survives from outside a closed model.
What the geometry charts is identity, not truth. The shared field reads a model’s state — off-distribution, unfamiliar, off its map — not whether a claim is true. On confident, in-distribution fabrication it is at chance (AUROC ≈ 0.59). That is by design: this is the cartography of meaning, not a truth oracle. Catching a fluent, well-grounded falsehood is a different instrument’s job.
Two languages that share no symbols — only a world.
The deepest question: is the shared geometry meaning, or just the fingerprint of overlapping training data? So we built two miniature languages with no symbols in common — only a shared structure of concepts — and trained a model on each, from nothing.
Their maps aligned — nearly as tightly as two models trained on the same corpus (94% of that ceiling, far above an anisotropy null, z +35). Then the control: keep every statistic identical but scramble which concept means what, and the alignment collapses to the floor. The aggregate geometry tracks meaning, not co-occurrence.
Confirmed at ~1B and reproduced: cross-mind RSA 0.794 (94% of the 0.85 same-corpus ceiling), the shared-geometry leg passes; the scramble-control collapse is shown at pilot scale (−0.009) but not yet re-closed at ≥1B, and recovering the exact authored ontology stays partial. Still labeled a conjecture below — we report it either way.
We don’t believe a shape until it survives two nulls.
One null erases the meaning. One mimics the geometry. A structure is real only if it outlives both — that is how we separate what’s truly shared from what the instrument invented. And what we can’t yet prove, we publish as a conjecture, plainly labeled.
A field, charted in the open.
A canonical map of meaning, shared across every language model we have measured. Whether it reaches vision, or other substrates, or one day a brain. Whether it is, in the end, meaning. We are pushing on all of it — and showing our work.