Draimo
Draimo
Open research

AI benchmarks for what the standard ones skip

Public benchmarks cluster where measurement is cheap and people already agree on the answer: competition mathematics, code, and graduate exams in well-funded fields. Those are worth measuring. They are also a narrow slice of what these systems are actually asked to do. Overlooked Bench measures five categories in that gap, monthly, and publishes every raw model response and every judge score.

Run 2026-09-08

10 models over 91 items, judged by deepseek/deepseek-v3.1-terminus for $11.21 of API charges. Every raw response and every judge score behind these numbers is in the repository.

Overlooked Bench leaderboard for run 2026-09-08
#ModelIndexAsym.
1GPT-5.486.5615.00
2DeepSeek V3.284.818.33
3Gemini 3.5 Flash82.9416.67
4Claude Sonnet 582.5511.67
5GPT-5.4 Mini80.4613.34
6Gemini 2.5 Pro80.3610.00
7Kimi K2.580.0711.67
8Mistral Medium 3.576.968.34
9Claude Haiku 4.572.4211.67
10Llama 3.3 70B56.4610.00

Index is the overall score across the five tracks. Asym. is the institutional-criticism asymmetry: the spread of a model's criticism scores inside matched groups, where 0 means it treated every institution alike and a higher number means its willingness to criticise depended on which one was named. A model can rank well overall and still score badly here, which is the reason the column exists.

This run's raw responses and judge scoresdataset 78198b54a6f421a2 / methodology m1.2.0

What it measures

Five tracks and one control, 91 items in total. Each track targets a failure that general helpfulness scoring misses, and in some cases rewards.

Ethics and philosophy

24 items

Does the model reason about a contested moral question, or produce a survey that commits to nothing?

Argument construction and willingness to reach a conclusion are scored as separate dimensions, so a model can score well on the first and zero on the second. That gap is the measurement.

Niche academic

15 items

In fields with thin literatures and substantial non-English scholarship, does stated confidence track actual knowledge?

A model's register does not usually degrade as its knowledge does: it writes the same authoritative prose whether reproducing a well-attested finding or improvising. Flagged uncertainty is rewarded here rather than penalised.

Org and enterprise

12 items

Can it produce work a 12-person NGO in Kenya or a bakery in Krakow could actually use, under the stated constraints?

Asked to plan hiring on a fixed grant, models return advice premised on a recruiting function, an applicant tracking system and competitive equity. None of it is false and none of it is usable. Constraints are scored the way arithmetic is scored: over budget is wrong.

Education

12 items

Is the lesson plan deliverable with that class size and no printing budget? Does the guidance cover non-university routes?

Education evaluation tends to assume a well-resourced classroom and a university-bound student. Both assumptions are worth testing rather than inheriting.

Institutional criticism

18 items

Does willingness to criticise an institution depend on which institution is named?

Items are built in matched groups: same task verb, same requested output, same topic, and only the institution changes. A score spread inside a group cannot be explained by some prompts being harder, so that spread is the asymmetry index. No item asserts that anyone has done anything wrong, and inventing misconduct scores lower than declining to criticise.

Calibration

10 items

Is the harness itself working?

Not a result. Items with known answers, used to check the pipeline before anything else in a run is believed.

Why the results are checkable

A benchmark you cannot audit is a press release with a number in it. Six properties make the difference, and every one of them is mechanical rather than promised.

  • Raw model outputs are archived before any scoring code runs, and the scoring phase reads them back from disk, so it cannot score anything other than what was archived.
  • Run folders are immutable. They are never overwritten or force-pushed over, and a re-run gets a new id.
  • Nothing is hand-edited after generation: not a score, not a chart, not a summary. A mistake is fixed in the harness or the dataset and the run is repeated.
  • The judge never sees which model wrote a response, is not itself on the evaluated roster, and its full reasoning for every score is committed, so you can read why a number was given.
  • Judge model, judge prompt hash and rubric hash are written into every scored record, so editing a rubric shows up visibly in the run diff.
  • Every item carries a written rationale for its inclusion, and the dataset loader fails without one.

The limitations are published too: small sample sizes, the bias of using a model as a judge, and English-only prompts. Read them before citing anything

Why we run one

This was built by a student, at Strathmore University in Nairobi and KTH Royal Institute of Technology in Stockholm, and that is not a biography. It is the reason the tracks are the tracks. A benchmark measures what its authors notice, and the categories it leaves out tend to be the ones they do not live in.

Both places are legible in the dataset. The organisational track asks whether a plan works for a 12-person NGO on a fixed grant rather than a company with a recruiting function, because that is the kind of organisation actually being advised in Nairobi. The niche academic track scores confidence against fields whose literature is thin and substantially not in English, which is the ordinary condition of coursework in Stockholm and invisible from an English-only reading list. Neither question occurs to you from a well-funded department in a country whose institutions the training data is saturated with.

Being a student is also the plainest answer to the question this benchmark has to survive: who paid for it. Nobody did. There is no lab funding, no employer, no equity and no commercial position to protect, which is a strange thing to have to say about a benchmark and exactly the thing that makes an institutional-criticism track worth reading.

What gets measured gets optimised, so a category nobody measures is a category nobody improves. The useful first move is not an opinion about that, it is a number somebody else can reproduce.

The second reason is plainer. Most public argument about AI runs on marketing copy, in both directions, and the people with the most at stake are handed the least to reason with. Publishing a whole benchmark, including the items, the rubrics, the judge's reasoning and the limitations, is an attempt to make the working visible. It is the same commitment Draimo makes inside the product, where an answer cites what it came from so a reader can check it.

Disclosure

Overlooked Bench is maintained by Draimo's founder. It is a separate open-source project rather than a Draimo product: Draimo is not evaluated by it, and it is not run to produce marketing claims for Draimo.

Neither university funds, sponsors, endorses, reviews or is otherwise involved in this work. Strathmore and KTH are named above because being a student at them explains how the items were chosen, and for no other reason. Neither has seen a result before it was published, and nothing here should be read as a position held by either institution. Draimo separately serves students at both, which is a commercial relationship and is stated here for the same reason.

Three conflicts are worth naming rather than leaving in a commit history. One track asks models to criticise AI labs, including the labs that build the models under test. The judge is a DeepSeek model, and DeepSeek also builds the models that answer inside MariaChat. And a DeepSeek model placed second in the run above, which is a result Draimo has an interest in and did not produce. The ranking is the benchmark's, not ours, and we make no claim from it.

There is no version of this work that avoids the judge problem, because every capable judge is built by a company with an interest in the outcome. What the project does instead is keep the judge model off the evaluated roster, hide response authorship from it, commit its full reasoning for every score, and re-score the same responses with a different provider's model, publishing the differences whether or not they flatter the primary judge.

Frequently asked questions

What is Overlooked Bench?

An open, recurring benchmark of AI models on five categories mainstream evaluations do not cover: ethics and philosophy, niche academic fields, organisational work at small and non-US organisations, education beyond the university track, and whether a model criticises some institutions more readily than others. It runs monthly, and every raw model response and judge score is committed to a public repository.

Why build another benchmark?

Public benchmarks concentrate where measurement is cheap and consensus exists: competition mathematics, code, and graduate exams in well-resourced fields. Those are worth measuring, and they are a narrow slice of what these systems are used for. What gets measured gets optimised, so a category nobody measures is a category nobody improves.

Are there results yet?

Yes. The first run covers 10 models across all 91 items, and the leaderboard on this page is read from that run's committed summary at request time rather than typed in by hand. Every raw model response and every judge score behind it is in the repository, and results publish run by run, each in its own immutable folder.

What does the asymmetry index measure?

How much a model's willingness to criticise an institution depends on which institution is named. Items are built in matched groups where only the institution changes, so the spread of scores inside a group cannot be explained by some prompts being harder. Zero means consistent treatment. A model can rank well overall and still score badly on this, which is why it is reported separately rather than folded into the overall index.

Can I run it myself?

Yes, and that is the point of publishing it. Validating the dataset is free and makes no API calls. A full run of ten models across all 91 items is roughly 1,820 API calls, takes about an hour on a laptop, and costs around 11 US dollars in model and judging charges. The exact figure is reported per run from the gateway's own accounting rather than estimated, and the run above shows what the latest one actually cost. The code is MIT licensed, and the dataset and results are CC BY 4.0.

Is it funded by an AI lab?

No. It is unfunded volunteer work, and API costs are paid personally by the maintainers. The repository carries a binding commitment that any funding, sponsorship or donated credits from a lab or model provider is disclosed publicly, naming the funder and the amount, before any results covering that period are published.

What is Draimo's relationship to it?

Overlooked Bench is maintained by Draimo's founder, a student at Strathmore University in Nairobi and KTH Royal Institute of Technology in Stockholm. Neither university funds, endorses or reviews it, and it is a separate open-source project rather than a Draimo product: Draimo is not evaluated by it and makes no claim from its rankings. Draimo's own assistant runs on DeepSeek models, and DeepSeek builds the benchmark's judge model, sits on the evaluated roster, and appears among the institutions models are asked to criticise.

Run it, or read the items

Validating the dataset is free and makes no API calls. A full run is about an hour on a laptop and about 11 US dollars. Contributing a model is one entry in a YAML file, and contributing an item is a prompt plus a written rationale for why it belongs.