London-based AI lab Inherent Labs says a 27-billion-parameter agent it built, called Faraday, has outperformed both Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5 at a task few AI systems have been tested on: reproducing the results of published scientific papers.

Replicate, don’t just predict

Faraday was evaluated on Replica, a benchmark Inherent Labs built in-house. It asks an AI agent to regenerate a specific figure from a research paper under a limited time and compute budget, without ever seeing the original plot. The initial suite spans 310 tasks drawn from 100 papers across fields including natural language processing, materials science and weather forecasting.

According to Inherent’s own write-up, Faraday produced more faithful replications than Claude Opus 4.8 and GPT-5.5 in every paper category tested, despite running on a far smaller base model — Alibaba’s open-weight Qwen 3.6, at 27 billion parameters. Faraday also directs GPT-5.5 Codex as a coding subroutine during its work, meaning a much smaller orchestrating model outperformed one of the larger systems it partly relies on.

How it was trained

The company says Faraday was trained with long-horizon reinforcement learning, using per-task grading rubrics to cut down on noisy reward signals, multi-sample aggregation, and turn-level credit assignment, with an LLM judge validated against human evaluations. Chief scientist Edward Hughes, a Google DeepMind alumnus, told TechCrunch that the more interesting result was the training method itself rather than the head-to-head win over larger models.

Inherent Labs, founded by several former DeepMind researchers, emerged from stealth in May 2026 with a $50 million seed round. The company, which has around a dozen employees, says it plans to grow to 20–25 staff by the end of the year.

A narrower kind of success

The researchers are careful to flag the limits of the result. As they put it in their own write-up, “perfectly reproducing a plot is not the same as successful replication” — genuine scientific replication also requires sound experimental design, faithfulness to the original claims, and sensible use of resources, none of which the Replica benchmark fully captures. Independent, third-party verification of Faraday’s results has not yet been published.

The result adds to a mixed picture of how AI systems handle real scientific work: an earlier study found that current models could recover only 3–15% of novel research ideas when tested against held-out papers, suggesting genuine scientific reasoning remains far from solved even as narrower replication tasks improve.

Read also