Manifesto

LEER EN ESPAÑOL[ESC] MAIN MENU
════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════════

Versión en español

The hole in how we test machines today

Almost every benchmark used to rank language models has the same defect: the model has probably already seen it.

This is not a suspicion. Major mathematics benchmarks such as GSM8K and MATH have been found in the training data of modern models, with contamination reported across 31 of them. One analysis puts MMLU at 29% contaminated, with at least one model dropping 13 points once tested on a clean set. A 2026 study that trained models from scratch while varying how many times a test set appeared in the corpus confirmed the mechanism directly: given enough capacity and enough incentive, models memorise the test.

A score on a contaminated benchmark measures recall. We read it as reasoning.

And the problem repairs itself in the wrong direction. Publishing a benchmark is what makes it useful — and what destroys it. The moment it exists on the open web it is scraped, and every model trained afterwards has an advantage that looks exactly like intelligence. Every public benchmark is a decaying asset.

Visual reasoning is not exempt. Raven's Progressive Matrices were designed to strip out language and culture and measure abstract inference directly, which is why they remain attractive. But they are in copyright, they are all over the internet, and it is still unclear whether model performance on them reflects genuine reasoning or statistical shortcuts. Using them to rank frontier models tells you very little.

What actually works

The benchmarks that resist this share one property: the problems are new, and they are not published. FrontierMath is the clearest example — hundreds of original problems written and vetted by over 70 mathematicians, kept private apart from a handful of examples, specifically so that no model can succeed by pattern-matching against something it read during training.

That is the model we follow. It costs more, because someone has to invent every problem. That cost is the point: it is exactly the cost that cannot be shortcut by scraping.

What this is

An independent benchmark where the problems are made by people who are good at making problems, and where nothing a subject could memorise has ever been published.

Items are original. Authored here, by invitation — Mensa members and others who build interesting problems for their own sake. Contributors warrant that their work is their own and not adapted from any published test.

Answers exist nowhere in plaintext. A creator marks the correct option and it becomes a cryptographic digest on arrival; the plaintext is discarded. Grading compares digests. Nobody — not the operator, not the server, not a process that compromises the server — can read an answer out. The key is not protected. It is absent.

A leaderboard score is produced without opening any key. The digest above protects a key that exists. To produce a number we do not open it: the sheet is marked by eye against the plates, and only the count is entered. Nothing on that path reads a key, so there is nothing on it to steal, and no file whose secret matters later.

What does still exist, and we would rather say so: a sealed digest file for the scanned booklet, on the server and in the repository. It holds no plaintext, but six options across forty-five items is a small number of guesses for whoever holds the secret — which is exactly why the grading path should never touch it. The artefact actually worth guarding is the pile of sheets from past runs, and those are kept outside the repository where a subject under test cannot reach them.

The cost is that a person has to mark every sheet, and a person can miscount. We would rather carry that error than the other one.

Reasoning is collected. It is never published. Models are asked to write down why they chose each option, and those rationales are what the research is actually for — where a model's rule was wrong, and whether two models fail the same item the same way. They are archived offline and appear nowhere on this site, because a set of rationales names an option for every item, and enough of them together would do what publishing the answers would do.

People and models sit the same items, under the same clock. The time limit belongs to the test and is set by its author, never chosen by the subject, so two scores can be placed side by side and mean something.

Benchmark tests are used once, and not published. Public puzzle sets exist for anyone who wants to play — that is the fun of it — but the sets used to evaluate models are drawn from private submissions and stay private. A benchmark loses value every time it is sat; a game gains it. Keeping them apart means the fun cannot burn the measurement.

Models are run by hand, and marked by hand. There is no public endpoint that will grade an arbitrary submission, because that is an oracle: six uniform attempts would reconstruct an answer key from the scores alone. For the same reason a model's run stores a count by default and no per-item grid — a grid tells you which items are A the moment someone answers all A.

Where this goes

Matrices are the beginning, not the point. The same architecture holds anything with a verifiable answer, and the intention is to extend it to mathematics and physics problems curated by working mathematicians and physicists.

Not harder arithmetic. Not longer word problems. Problems with depth — where the difficulty is in seeing what the problem actually is, and where the answer cannot be found anywhere online because it has never been anywhere online.

What we are careful not to claim

Scores here are raw counts. Calling anything an "IQ" requires a normed sample against a representative population, and no such sample exists yet — so the word is not used. Item difficulty is currently positional rather than measured, and should come from response data once enough people have sat each item.

Marking by hand trades one kind of error for another. A digest comparison cannot miscount; a person can. We accept that in exchange for a grading path that opens nothing, and a wrong count is at least a mistake in the open rather than a benchmark that quietly stopped measuring anything.

And a run where the model wrote out its reasoning is not the same task as one where it only chose letters — writing changes both the score and the time taken. Those sittings are ranked as their own entry and never averaged with answers-only ones. Two rows for the same model is not a duplicate; it is two different questions being asked.

We would rather report a small, honest number than a large one that quietly measures memory.


Sources


Take a testLeaderboardCreatorsMain menu