Startup's AI 'Scientist' Claims It Out-Researches Rivals
Inherent, a startup founded by four Google DeepMind alumni, said its AI agent, Faraday, outperformed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 at independently replicating published scientific research, according to TechCrunch. The company's own claim is notable because Faraday runs on a 27-billion-parameter model, far smaller than the frontier systems it says it beat.
Faraday was tested on Replica, a benchmark Inherent built from 310 tasks drawn from 100 machine-learning and AI-for-science papers spanning natural-language processing, materials science and weather forecasting, according to Inherent Labs' own research writeup. Each task asks an agent to reproduce a figure from a paper without seeing the original result, under limited time and compute. Inherent says Faraday "produces more faithful replications for every category of paper in the task suite," with a particularly strong edge in meta-learning, structural biology and materials science — though the company has not published exact percentage deltas.
Inherent co-founder and chief scientist Edward Hughes told TechCrunch the more interesting result wasn't beating larger rivals but how the team built the system, framing the goal as teaching Faraday "research taste" — judgment about which experiments are worth running and how to design them, a skill Hughes compared to what PhD students develop by replicating prior work. It's worth noting these are Inherent's own benchmark and framing, run on a task suite the company itself designed and has not (yet) been independently replicated by outside labs.
What this means for job seekers
If an AI agent can now be graded on days-long, open-ended research tasks rather than single-shot answers, it signals that the target for automation is shifting toward exactly the kind of judgment-heavy work — deciding what to test, how to test it, and what a plausible result looks like — that research analysts, R&D associates and junior scientists are trained to do. That's a different threat profile than earlier chatbot-style AI, which mostly automated lookup and drafting.
For job seekers in research-adjacent roles, the practical takeaway isn't panic, it's positioning. Employers evaluating candidates for analyst, R&D or technical research roles are increasingly likely to ask how comfortable you are directing or auditing an AI agent's work, not just doing the work yourself. Vendor benchmarks like this one should be read skeptically — Inherent designed and scored its own test — but the broader trend of agents being measured on multi-step, multi-day project work rather than single queries is real and worth tracking. Building fluency with AI research tools now, while treating today's specific performance claims with caution, is a reasonable hedge either way. Job seekers can start by reviewing how AI is reshaping job search and hiring and brushing up through career-relevant online courses.
The bigger open question the benchmark doesn't answer: how these agents perform on messier, real-world research problems without a known paper to check against — the kind of ambiguity that defines most actual analyst jobs.
Sources
Related Posts

AI & Careers Briefing — September 17, 2026

AI & Careers Briefing — September 16, 2026
