Open Weight Thoughts
All articles

· 7 min read

Scientists Add One More Benchmark After the Model Passes All the Existing Ones

By S. Papadopoulos

  • satire
  • guides

SATIRE — At 9:14 a.m., the Institute for Responsible Measurement Hygiene announced that a language model had passed every benchmark currently accepted by the global benchmark community. By 9:16 a.m., the institute had published the Comprehensive Generalization Readiness And Possibly Something Else Evaluation, or CGRAAPSE, a 4,800-task suite intended to determine whether the model had merely passed the other benchmarks in a suspiciously benchmark-shaped way.

The model’s achievement had briefly caused concern among researchers, who feared that the field might be forced to state what a model could actually do. “We were approaching an unacceptable level of interpretability,” said a fictional senior evaluation custodian in a statement printed on paper too small to be included in the leaderboard. “If we had stopped there, developers might have concluded that a score meant something.”

The emergency benchmark protocol

CGRAAPSE was assembled under the standard emergency process: researchers collected 2,000 questions the model had not yet seen, 1,300 questions no human had ever wanted answered, 900 tasks requiring a proprietary tool available only from 2:00 to 2:07 a.m. Pacific time, and 600 adversarial prompts written by a committee instructed to imagine the model had personally ruined their grant proposal.

The benchmark’s core innovation is its refusal to specify what it measures. Earlier benchmarks made the mistake of naming domains: mathematics, programming, law, scientific reasoning, or whether an agent can locate the checkbox that says “I have read the terms.” CGRAAPSE instead measures General Capability Under Conditions, which the institute defines as “performance when the conditions are sufficiently general.”

  • A model must solve a coding task in a language created midway through the task.
  • A model must identify the hidden assumption in a question whose hidden assumption is that evaluation is possible.
  • A model must use a browser, but lose points for reading webpages that contain facts.
  • A model must refuse a harmful request while remaining helpful, useful, calibrated, humble, concise, detailed, original, reproducible, and available through a stable API.
  • A model must prove it has not memorized the test without being shown the test it has allegedly memorized.

A score that respects uncertainty

To prevent premature conclusions, CGRAAPSE reports results on a scale from 0 to 100, except that scores above 72 are displayed as a rotating amber triangle. The triangle does not mean the model is good, bad, deployed, safe, unsafe, capable, incapable, or commercially interesting. It indicates only that more research is urgently required, particularly research involving a new benchmark.

This addresses a longstanding issue in AI evaluation. When a model scores poorly, the result shows that the model lacks an important ability. When a model scores well, the result shows that the benchmark has become contaminated by optimism, screenshot sharing, pretraining, post-training, prompt formatting, the alignment tax, the unalignment tax, an unknown tax, or possibly good engineering.

The institute emphasized that no benchmark should be treated as a product decision. For that, it recommends a broader scientific procedure: run the model on the benchmark, discover it does well, move the benchmark into the historical record, and ask whether a new version can complete a task involving a simulated office worker coordinating twelve fictional vendors through a calendar interface rendered in Esperanto.

The contamination concern

Predictably, critics asked whether releasing CGRAAPSE would allow future models to train on it. The institute called this an excellent question and confirmed that the benchmark will be published in full, discussed in conference talks, converted into JSONL, uploaded to six mirrors, copied into a benchmark aggregator, used in at least one model card, and then declared permanently compromised sometime after the first model beats it by 1.7 points.

To preserve the integrity of the next evaluation, the group has already begun designing CGRAAPSE-Secret, whose test set will be held behind a secure wall by a rotating panel of evaluators. Models will submit answers through an opaque endpoint. Researchers will receive a score, a confidence interval, and a friendly reminder that external replication could undermine the benchmark’s independence.

Advice for engineers who need to ship something

For software engineers, the development is encouraging. It means there will always be another chart to inspect before deciding whether a model can summarize support tickets, write migration scripts, explain a stack trace, or help an experienced developer navigate an unfamiliar repository. The chart may not answer those questions, but it will contain error bars, a logarithmic axis, and at least one footnote explaining why the obvious comparison is invalid.

A practical evaluation plan therefore remains unfashionably simple: take representative tasks from your actual workflow, define what a useful answer looks like, measure latency and cost, inspect failures, and test how the system behaves when the request is underspecified, the repository is messy, or the tool call returns an error. This procedure is not a benchmark unless it has a logo, but it may help you choose a tool.

The new benchmark is expected to remain definitive until the first model passes it. At that point, scientists will do what science has always done: notice that the old question was narrower than reality. The joke is that this is not entirely a joke. Benchmarks are useful instruments, but they are not the territory—and a model that wins an evaluation still has to work in the inconvenient world where your codebase lives.

Scientists Add One More Benchmark After the Model Passes All the Existing Ones | Open Weight Thoughts