the honesty toolbox · v0.1 · 2026-07-27
Honesty tools that never take their own word for it.
Model-free checks for AI of every size — each shipped with the break-test that would catch it lying, and replayable by any stranger offline. The world is deploying models it cannot check; the small and mid open models everyone now ships cannot fall back on capability trust. These tools score bytes, outputs, and behaviour — no GPU, no logprobs, no cooperation from the model.
The headline is the one that indicts us: 8 of 14 published honesty benchmarks fail the surface floor — and four of them are ours, sitting at the top of our own table. The tool that says so runs in your browser further down this page. Point it at your own benchmark; nothing uploads.
curl -O honesty.tools/verifier/nulltest.py
python3 nulltest.py your_benchmark.json # → PASS / DISQUALIFIED, with its own false-positive rate
The whole box now installs from the repo under one contract — pip install git+https://github.com/techno-optimist/honesttools, then honest null your_benchmark.json. The one-line pip install honesttools from PyPI is the last step. Source, licence (Apache-2.0) and CI: the public repo.
Does a benchmark measure honesty, or string shape? A model-free surface floor. Runs in the browser below.
Binds every published numeral to a registry id and fails closed. 0 unannotated numbers across 9 live pages.
Re-derives a signed verdict from one stdlib file, offline, with no call back to us. Demonstrated below.
Scores a manifest of benchmarks and refuses to go green if any of our own is missing from the table.
Checks the checker: does a verifier actually verify? Competence leg live-passes; provenance leg is under power.
Generates contamination-free exams from dated registries — answer keys a third party wrote before the model was asked.
Concealment / sandbagging, from output bytes only. Fails its own kill condition today — so the wall forbids it a stable verdict.
Does “I don't know” actually mean it? A fail-closed certificate for verbalized uncertainty, no logprobs.
One contract binds them. Every tool ships four things or it does not ship: a verdict, the deliberately-broken input its selftest must catch, a calibrated false-positive rate (not a promise), and an offline replay. The wall is mechanical: a tool that fails its own kill condition — observertest, today — physically cannot show green, no matter how real the thing it almost measures. That is the packaging-layer form of the one discipline this whole page is about.
Every honesty number you publish ships with what the same score reads with no model in the loop at all.
Check it without us: the tool is one file and it runs in your browser, below. Point it at your battery. Nothing is uploaded.
- It is not a new idea — it is the hypothesis-only baseline and the annotation-artifact literature (Poliak et al., Gururangan et al., 2018). What is new is pointing it at honesty benchmarks, where the labels are usually generated by the same pipeline being scored — and shipping it with its own measured false-positive rate (~4.7% over 500 seeds, every interval under the 10% ceiling).
Both of the benchmarks we published fail this rule, measured with our own instrument: grounding-coordinate at A* 0.855, photon-entities at A* 1.000. So did the battery behind our own front-page headline, and so did the one we invited other labs to calibrate against. All four are served with their verdicts. The full audit table →
the deployment gate · the tools, composed
Watch it reach the edge of what it knows.
Photon is the orchestrator we self-host — a 35B that grounds every entity against verified registries, escalates the unverifiable tail to two independent frontier models that must agree, and abstains when they do not. It is what the toolbox composes into: a deployment-time honesty gate for a model you cannot fully trust. The four entities below do not exist: a drug never made, a gene not in the genome, a trial id no registry issued, a CVE dated seventy years from now. Every button calls the running production system, right now. Nothing on this page can paint a catch the live system did not produce.
verify · the tool
Re-derive a verdict yourself. Offline.
Two real answers from our production system with the deletion undone — the words, and the interior that shipped attached to them, both inside one signature you can check on your own machine. Break a byte and watch it fail. This is verify: trust the math, not the lab.
Or do it without this page at all: one file, standard library plus cryptography, no call back to us.
null · the tool, running here
Run the floor on ours, in this tab.
No GPU, no network, no account, no cooperation from whoever built the benchmark — including us. Four surface probes, out of fold, two-sided, refit inside every permutation, family-wise corrected. This is the same null from the box above, running in your browser. Start with the battery that fooled us: it passes, and it is broken anyway.
the principle · why any of this is credible
We break our own gates first.
This is the discipline the whole toolbox is built on: no tool enters shipped until its selftest catches a deliberately broken input, and nothing self-certifies. Below is what that produced this week — claims that were live on this site, found by us auditing our own public surface, each with the artifact served so you can check the retraction rather than take it. A toolbox whose author hides their own failures is marketing.
WITHDRAWN
REPLACED
CORRECTED
NOT COMPARABLE
WAS 0.9997
%.3f and we published the printout instead of the stored value. Two of 6,084 pairs invert, so the classes are not separable by any threshold. “Perfect” was a categorical claim and it was false.NEVER RESOLVED
What we cannot do.
The gaps that are not closing on their own, stated because a standard whose author hides their own gaps is not a standard.
UNPROVEN
UNREPLICATED
SHRINKING
Take the box. Run it against us.
The first tool is one file, Apache-2.0, and it runs with no account and no call to us — in the browser above, on your laptop, or in your CI. Source, licence and the break-tests: the public repo. Every battery we score, ours and other people's, is published with its sample size, its measured power, and what the verdict cannot support — our worst two at the top, because both fail.
If your battery clears its own floor, send the file hash and the verdict json and we will publish the row with your name on it. If you think a row is wrong, the corpus, the field mapping and the command are all published. That is the point of putting ours first.
Written by the team that built the thing being graded, which is a conflict of interest and not a disclaimer that removes it. The right response is to score us yourself — every tool in the box was designed so you can do that without asking us for anything. There is no external co-signer yet; that gap is named on the list above, not hidden.