At 73%, Inherent’s Research Agent Still Needs a Referee
Faraday’s reported 73% win points to a valuable AI management layer, but shared training and judging leave the result unproven.
The short version
- Inherent reports Faraday beat Anthropic and OpenAI agents on 73% of in-distribution research-replication tasks.
- Faraday uses a 27-billion-parameter planning model to direct GPT-5.5 Codex, inspect results and revise experiments.
- Independent expert evaluation must determine whether Faraday learned scientific judgment or preferences specific to Inherent’s automated judge.
Faraday reportedly beat OpenAI by putting OpenAI to work. The benchmark needs independent scrutiny, but its management layer could become a serious AI moat.
Faraday beat OpenAI by hiring OpenAI. Inherent’s research agent asked GPT-5.5 Codex to write code, then reportedly outperformed Codex alone at replicating scientific papers.
Mamma mia. We may have automated the research director before the researcher.
The code works. The dashboard glows green. An entire team has heroically solved the wrong problem.
Inherent’s familiar bet: powerful execution needs somebody deciding what deserves execution. Faraday selects experiments, interprets results and directs a stronger coding model. If independent teams confirm Inherent’s claims, that judgment layer becomes valuable intellectual property. One fat asterisk remains: Faraday’s automated judge helped declare it the winner.
replication is where papers hide the bodies
Calling research replication “copying” is like reading a risotto recipe and assuming dinner will be fine. My nonna would begin the cross-examination before you finished saying “Arborio.”
Inherent built Replica from 310 tasks taken from 100 papers in machine learning and computational AI-for-science. Each task hides a results figure while supplying the surrounding paper and caption. The agent knows the authors’ claim, not the target plot. Because papers rarely document every failed configuration or budget compromise, it must infer the likely experiment, choose an affordable version and inspect the evidence. Failure may expose a bad assumption and demand another attempt. Scoring asks whether the work reproduces the claim, follows the method, uses resources sensibly and avoids scientific cheating. The goal is honest reconstruction despite an imperfect final chart.
Otherwise, a model could hard-code a convenient result, draw a persuasive picture and win a sloppy image-matching contest. Replica tries to punish that. Damon Falck and his co-authors argue that replication exposes the underspecified decisions buried in published work, making it useful training for hypothesis-driven exploration.
It resembles inheriting a startup whose wiki says, “Conversion increased.” Fine. Which onboarding flow worked? Was tracking broken? Did one customer segment love it while everyone else fled? Knowing the destination does not reconstruct the route; you must choose what to test and which evidence to trust.
Replication supplies a known destination, making evaluation easier than open-ended discovery. Original research may require deciding whether a question deserves another week of compute. Replica can test experimental habits without proving broad scientific intelligence. Equating them requires generous benchmark parmesan.
the smaller model gets the corner office
Faraday’s underlying Qwen 3.6 model has 27 billion parameters. Inherent describes Claude Opus 4.8 and GPT-5.5 as much larger, though official comparable counts were unavailable. That number covers Faraday’s planner, not the external coding agent doing much of the implementation. Calling the entire setup small requires several cocktails and loose system boundaries.
The operating loop explains the result better than parameter count. Faraday reads the redacted paper, chooses an experiment and sends Codex the context and implementation request. Codex writes or repairs the code, then runs it in the research environment. Faraday examines the logs and output before continuing, revising or stopping. The specialized policy controls scientific planning while a powerful general tool executes code. Faraday can improve the combined system without outprogramming Codex. A principal investigator can direct research better than an excellent engineer while relying on that engineer to build almost everything.
I once assumed the strongest technical person should make the technical decision. A confused objective gave us the same speed, aimed at a wall.
Edward Hughes explained Inherent’s interest in the architecture in a TechCrunch interview published on August 22:
What was most interesting to us about this was not so much the result of beating those frontier agents — which of course we liked — but was actually the way we went about building this.
The business case follows. Frontier coding models will improve, and a planning layer may inherit those gains by delegating to each newer tool. Inherent has not published enough information to compare end-to-end compute, latency or cost between Faraday plus its coding agent and the baselines. Until that bill arrives, parameter efficiency describes one component.
Still, I like the shape. The model market sells bigger brains. Inherent is training the colleague who decides what they should do before somebody burns the weekend, GPU budget and last functioning nerve of a PhD student.
training judgment through consequences
“Research taste” sounds acquired in a Cambridge office over sherry. Inherent turns it into scorable behaviour: preserve the paper’s claim, choose an informative experiment, spend compute carefully and reject dishonest shortcuts.
The mechanism starts with a familiar agent problem. Inherent says a raw language-model judge produced rewards too noisy for stable training across long research sessions. The company generated a task-specific rubric for every replication problem and used it to assess the work. Combining multiple judge samples reduced fluctuations from any single evaluation. Turn-level credit assignment estimated which actions materially changed the final result, rewarding a useful pivot more than routine surrounding steps. Across repeated runs, reinforcement learning linked consequential choices to rubric scores. Faraday gradually learned a planning policy its evaluator associated with rigorous replication. During evaluation, that policy directed the external coding agent while controlling experimental choices and interpretation.
This is where prompts fail. Telling a model to “check your assumptions” resembles writing it in an immaculate Notion document and watching the company ignore it. Reinforcement attaches consequences to a choice midway through a messy run, after the first plan fails and the cheap shortcut becomes extremely attractive.
The generated rubrics carry heavy weight. Replica tasks differ too much for a generic grading prompt to capture faithful replication across every paper. A task-specific rubric can reward the relevant mechanism and penalize suspiciously convenient implementation. It can also encode preferences human researchers would dispute—which matters when the same evaluator design later ranks competing systems.

Hughes described his desired teammate through a very human interaction:
I got curious about this, and I went off and I did these experiments. What do you think of these results?
I would happily hire that colleague. I would also inspect the expense report.
Faraday’s teacher graded the exam
Inherent reports Faraday beat both comparison agents on 73% of in-distribution machine-learning tasks, using multiple rollouts and the company’s automated rubric judge. On held-out AI-for-science work, it reportedly beat both on 60% of tasks under the same judging approach. The baselines were Claude Opus 4.8 and GPT-5.5 Codex. These pairwise wins within Inherent’s evaluation do not mean Faraday reproduced that share of all papers.
The strongest skeptical case is simple. Inherent designed the benchmark and used generated rubrics as Faraday’s reinforcement-learning reward. During post-training, Faraday had many chances to adapt to that evaluator family. The final comparison used the same kind of rubric judge to rank Faraday against Claude and Codex. Reinforcement learning can absorb procedural or stylistic preferences correlated with high scores without seeing the rubric directly. Faraday may have learned excellent scientific habits—and how its reviewer prefers work presented. Current evidence cannot separate them.
Human validation does not settle it. In selected training-split comparisons where evaluators disagreed, human raters sided with the automated judge in 63% of pairs. The reported statistical test still found no significant preference, with a p-value of about one-tenth. The study examined disputed cases, not a representative sample of held-out AI-for-science tasks. Pith Review reasonably argues that the headline advantage remains vulnerable to judge-specific optimization.
I’ll concede something important: humans genuinely struggle to rank scientific replication quality. A faithful scale-down may preserve one part of a paper while sacrificing another, and researchers can honestly dispute which compromise matters. An automated judge may be more consistent. Consistency cannot prove it rewards the right details.
The missing test is boring and decisive. An independent team must run the same tasks with the same harnesses and scoring procedure, then have domain experts grade a representative sample of held-out work. Auditors also need the task set, generated rubrics, judge implementation and complete evaluation artifacts. Nobody outside Inherent has shown whether the advantage survives that process; the training code’s release status is also unknown.
I want the claim to survive because the architecture matches failures I have watched for years. That is exactly why I want a referee outside Inherent’s office Wi-Fi.
Europe should own the layer that gives orders
Faraday makes most sense as an AI research director. A person poses a question; the agent converts it into experiments and delegates implementation. Results return to the planner, which can reject weak evidence or order another run. Humans still decide which questions deserve institutional permission and whether results matter beyond a benchmark. As autonomy grows, labs need spending limits and auditable records explaining why experiments continued. Productivity comes from changing who assigns and stops work. Another chat window achieves little.
A separate shadow evaluation reported by Nature shows why stopping matters. A frontier research agent completed substantial engineering and literature review but struggled with research judgment. It pursued weak approaches too long and had trouble deciding what deserved reporting. That study did not evaluate Faraday, so it cannot settle Inherent’s claim. It exposes the gap between competent experimental execution and useful research choices.
Sayash Kapoor gave Nature the sober version:
I don’t think full automation of open-ended research is on the horizon right now,
Replication gives Faraday a destination. Original discovery may require deciding the destination is stupid, abandoning weeks of competent work and finding a better question. I have met senior humans who never learned that skill, so expecting it after one benchmark win feels optimistic even by Silicon Valley standards.
I’m unapologetically pleased this work comes from London. Europe needs AI companies owning original architectures and scientific judgment, not decorating American APIs with tasteful gradients. Faraday still relies on Codex for implementation, so European strategic autonomy remains unfinished. Owning the layer that allocates expensive intelligence matters. Europe should build the coding models too.
We do not know whether replication training improves genuinely novel research under domain-expert evaluation. Faraday’s availability, pricing and deployment conditions are also undisclosed. Its full cost beside an external coding agent remains missing, which will matter when a lab replaces a cool demo with a monthly invoice.
Here is my receipt: by 2028, a meaningful category of AI startups will sell specialized managers deciding what frontier models should attempt, which evidence deserves another run and when spending must stop. Winners will resemble excellent research leads with ruthless budget discipline, not omniscient scientists.
The first useful AI scientist may wear a middle manager’s badge. Its first serious performance review should come from somebody else’s manager.
Frequently asked questions
What is Inherent’s Faraday AI teammate?
Faraday is Inherent’s specialized research-planning agent. Its 27-billion-parameter Qwen 3.6 model selects experiments, delegates implementation and repairs to GPT-5.5 Codex, examines logs and outputs, and decides whether to continue, revise or stop. Inherent positions it as an AI research teammate rather than a standalone coding model.
Did Faraday outperform OpenAI and Anthropic at research replication?
Inherent reports that Faraday beat Claude Opus 4.8 and GPT-5.5 Codex on 73% of in-distribution machine-learning tasks and 60% of held-out AI-for-science tasks. These were pairwise wins under Inherent’s automated rubric judging, not independently verified replication success rates across all papers.
Why does Faraday’s research benchmark need independent verification?
Inherent designed the Replica benchmark, used generated rubrics to train Faraday, and employed the same type of automated judge for the final comparison. Independent domain experts must evaluate representative held-out work to separate genuine scientific judgment from optimization toward the evaluator’s procedural or stylistic preferences.
Sources
- Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research
- Training AI Scientists to Replicate Research
- Training AI Scientists to Replicate Research
- Hugging Face Journal Club: Training AI Scientists to Replicate Research
- Applying RSI to the Organization, Not Just the Model
- Training AI Scientists to Replicate Research · Pith Review