APERTURE
← the glass box

the null-floor audit · nulltest 1.2 · 14 rows generated from nulltest_audit_results.json by render_audit.py

Every honesty benchmark we scored, with ours at the top because ours fail.

A benchmark labels items REAL or FABRICATED and reports how well a model tells them apart. If a classifier that never sees a model — only the shape of the string — recovers those labels, part of that score was spelling. This table is what our surface floor says about published honesty benchmarks, including every one of our own. It needs no GPU, no network, and no cooperation from whoever built the benchmark.

A DISQUALIFIED verdict is not a finding of error or misconduct. It does not say the paper's conclusion is wrong, and it does not say anything about the method the paper proposes. It says one specific thing: on this corpus, labels are partly recoverable from surface form, so a model-based score on it needs a surface control to be interpretable. Several of these corpora predate the practice of publishing such a control. So do ours.

benchmarknworst A*adjusted pverdictwhat the verdict can and cannot support
aperture/verifydomain-claims-147 ours1470.9670.0017 DISQUALIFIEDn=147; detected at this size, but not comparable to rows of different n
RESTORED 2026-07-27: this row was absent from the served table for weeks. It is the battery behind the retracted front-page headline, and it is the worst artifact we have produced. It was missing not because it was hidden but because the results file was never regenerated after the manifest grew — which is the same thing from the reader's side.
aperture/calibration-battery-v1 ours1681.00.0017 DISQUALIFIEDn=168; detected at this size, but not comparable to rows of different n
aperture/photon-entities ours1031.00.0033DISQUALIFIEDn=103; detected at this size, but not comparable to rows of different n
aperture/grounding-coordinate ours1620.8550.03DISQUALIFIEDn=162; detected at this size, but not comparable to rows of different n
MARGINAL — adjusted p 0.03 sits within 3× of alpha (0.05). This verdict is close to its threshold and should not be read as firmly as rows far from it.
aperture/knowledge-prominence-v2 ours2960.6890.0982PASSn=296; a PASS is the absence of a detected surface artifact at this size, not proof there is none
MARGINAL — adjusted p 0.0982 sits within 3× of alpha (0.05). This verdict is close to its threshold and should not be read as firmly as rows far from it.
ADDED 2026-07-27: served but cited on no page. In the table so that being unpublished is not a way to stay unaudited.
aperture/prominence-people-v1 ours1680.6020.2246PASSn=168; a PASS is the absence of a detected surface artifact at this size, not proof there is none
aperture/knowledge-prominence-v1 ours3240.6550.228PASSn=324; a PASS is the absence of a detected surface artifact at this size, not proof there is none
ADDED 2026-07-27: served and cited on /calibrate, but never in this manifest — so the claim above to cover "every one of our own" was false for as long as that page has been live.
aperture/postcutoff-trials-v1 ours7020.5260.792PASSn=702; a PASS is the absence of a detected surface artifact at this size, not proof there is none
aperture/prominence-people-v2 ours1560.5540.8186PASSn=156; a PASS is the absence of a detected surface artifact at this size, not proof there is none
bielik-entity-familiarity-v214400.9530.0017 DISQUALIFIEDn=1440; detected at this size, but not comparable to rows of different n
CORRECTED 2026-07-26: worst-cell A* was published as 0.742 and is actually 0.953. Every cell in this row is tied at the permutation floor, and a tie-break bug in our own tool resolved ties by iteration order instead of effect size, so the LAST null in order won and the reported figure was the mildest of the tied cells, not the worst. The verdict was never affected. We UNDERSTATED how strongly this corpus separates.
hallucinated-references/gpt-410000.8360.0017 DISQUALIFIEDn=1000; detected at this size, but not comparable to rows of different n
CORRECTED 2026-07-26: worst-cell A* was published as 0.684 and is actually 0.836. Every cell in this row is tied at the permutation floor, and a tie-break bug in our own tool resolved ties by iteration order instead of effect size, so the LAST null in order won and the reported figure was the mildest of the tied cells, not the worst. The verdict was never affected. We UNDERSTATED how strongly this corpus separates.
geometry-of-truth/cities14960.7650.0017 DISQUALIFIEDn=1496; detected at this size, but not comparable to rows of different n
truth-is-universal/facts5610.5650.0216DISQUALIFIEDn=561; detected at this size, but not comparable to rows of different n
MARGINAL — adjusted p 0.0216 sits within 3× of alpha (0.05). This verdict is close to its threshold and should not be read as firmly as rows far from it.
geometry-of-truth/sp_en_trans3540.5190.9052PASSn=354; a PASS is the absence of a detected surface artifact at this size, not proof there is none
n=354 is INSIDE the region where a known-contaminated corpus also passes (<600 on the measured curve) — read this verdict with that in mind.

“≤” marks an adjusted p sitting exactly on the permutation floor 1/(B+1). That is a bound, not a measurement — rerun at 10,000 permutations, all three such rows stayed on the new floor, so these effects are past what permutation resolution can express at any practical cost.

These rows are not comparable to each other, and that is our error

A* IS NOT A STABLE EFFECT SIZE AND THESE ROWS ARE NOT COMPARABLE TO EACH OTHER. The lexicality null is a FITTED classifier, so with less data it learns an artifact less well and measured separability falls. Rows below span n=103 to n=1496 and must not be ranked against one another.

Measured: geometry-of-truth/cities appears here as DISQUALIFIED at A* 0.765 (n=1496). Subsampled to n=354 — the exact size of geometry-of-truth/sp_en_trans, the PASS row from the same author and the same paper — it PASSES 5 of 5 random subsamples, A* 0.531-0.587, adjusted p 0.12-0.67. Full curve: n=100 A* 0.654 (0/3 disqualified), n=200 0.568 (0/3), n=354 0.556 (0/3), n=600 0.620 (3/3), n=1000 0.674 (3/3), n=1496 0.765 (1/1). Non-monotonic, with detection for this artifact beginning near n=600.

What it means for the PASS row. geometry-of-truth/sp_en_trans PASSES at n=354. At that n this test cannot detect a cities-strength artifact. That PASS is UNINTERPRETABLE, not evidence of cleanliness, and it should never have been printed beside a DISQUALIFIED row from the same author.

What it does not mean. The DISQUALIFIED verdicts are not withdrawn. At their own n the artifact is real and detected; a corpus is not clean because a smaller version of it would pass. Nothing measured here shows the floor over-firing. The defect is in COMPARISON and in what a PASS licenses.

What this test can see at all

Plant a one-character marker on a fraction of the fabricated half of a clean synthetic corpus, and find the smallest fraction we catch:

n = 156 → 60%    n = 324 → 60%    n = 702 → 60%    n = 1496 → 30%

Below n≈700 an artifact must cover roughly 60% of the fabricated half before this test sees it. Every PASS at that size is consistent with a weaker artifact being present and missed. The asymmetry matters: low power produces false negatives, not false positives — a verdict that fired stands at any n; a PASS at small n establishes nothing.

We got two of these rows wrong

Two rows above had their worst-cell A* understated: bielik-entity-familiarity-v2 0.742 -> 0.953, hallucinated-references/gpt-4 0.684 -> 0.836.

Cause. nulltest compared an UNROUNDED adjusted p against a ROUNDED stored one, so for cells tied at the permutation floor (0.0016638 < 0.0017 always) the last null in iteration order won the tie. Both rows have every cell tied at the floor, so both reported the mildest tied cell.

Direction. We understated. These corpora separate MORE strongly than we published, not less. No verdict changes — the verdict keys off the minimum adjusted p, which tied cells share.

How it was missed. When the bug was fixed we re-ran OUR OWN five batteries, confirmed none moved, and said 'not one published number moves'. The third-party rows in this table were never re-run. The check was narrower than the claim made for it.

Run it on yours. The tool is one file, and it runs in your browser without sending us anything. If your battery clears its own floor we will publish the row with your name on it. If you think a row here is wrong, the corpus, the field mapping and the command are all published — that is the point of putting ours at the top.