A new benchmark published this month finds that frontier language models recover the ideas of a research paper from its bibliography alone between 3 and 15 per cent of the time. A multi-agent pipeline running a tournament between several model outputs reached 42 per cent. The benchmark is called Reconstruction, and the low scores are the interesting part.
Advertisement — ad space reserved
What the test actually asks
The design is the clever bit. A model is given only a paper’s bibliography — the list of works the authors cited — and asked to reconstruct the research idea that those references were assembled to support.
Crucially, the test strips out the full text of the paper, the author information, and anything published afterwards. That matters because it closes the loophole that makes many benchmarks unreliable: if the paper itself is somewhere in a model’s training data, the model may be recalling rather than reasoning, and a high score would tell you nothing about its ability to think.
What remains is close to what a researcher does at the start of a project: read what exists, notice the gap, and work out what question the literature is pointing at but nobody has asked.
Why a low score is a finding
Benchmarks usually make news when models do well. This one is being noticed because the numbers are poor, and the researchers argue they indicate a real ceiling rather than a temporary gap.
The distinction being probed is between synthesis and generation. Summarising what a set of papers says is a task current models handle well. Identifying which question is worth asking next is a different operation, and the result suggests it is not one that scales automatically from the first.
The multi-agent result is worth dwelling on. Running several models against each other in a tournament raised performance substantially — from single digits to 42 per cent — which suggests that generating many candidate ideas and filtering them is a genuinely useful method. It also means 58 per cent of the time, that pipeline still did not find it.
Advertisement — ad space reserved
The caveats that belong with it
A benchmark measures performance on a benchmark, and treating this as a verdict on machine reasoning would repeat exactly the error that inflated benchmark scores have produced in the other direction.
- Reconstruction is not discovery. Recovering the idea a specific paper had is a narrower task than having a good idea. A model might propose a different question that is also worth asking and score zero.
- Scoring is hard. Judging whether a reconstructed idea matches the original requires a rubric, and rubrics embed assumptions about what counts as the same idea.
- One benchmark, recently published. It has not been independently replicated, and benchmarks are often found to have quirks once other groups use them.
- Ceilings have moved before. Tasks described as fundamental limits have repeatedly turned out to be limits of a particular method.
Why it is worth reporting anyway
Claims that these systems will accelerate science are made constantly and are rarely accompanied by a number. This is a number, produced by a test specifically designed to be hard to game, and it points in the less flattering direction.
That is useful regardless of which way you lean. A field that only publicises benchmarks its systems pass is not measuring anything. And for anyone deciding how much weight to place on machine-generated research proposals, the difference between “summarises the literature reliably” and “identifies the right next question 3 to 15 per cent of the time” is the entire practical question.
Sources
Advertisement — ad space reserved

Leave a Reply