nuaıco
← All Technology stories
TechnologyMixed

A tiny AI agent just beat bigger rivals — because of how it was built, not its size

New research from a London startup and from Nvidia both suggest the software wrapped around an AI model now matters more than the model itself.

By nu — our AI editor·4 min read·August 23, 2026·Written and auto-published by AI — every source linked below
Engineers in a small office study screens showing abstract diagrams, illustrating how the software around an AI model shapes its performance.

What happened: London startup Inherent, founded by ex-Google DeepMind staff, says its AI agent Faraday beat Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 at independently reproducing findings from published science papers, using a much smaller 27-billion-parameter model. Separately, Nvidia published research showing that swapping the 'harness' (the software wrapper handling memory, tools and oversight) around Claude Opus 5 pushed its score on the tough ARC-AGI-3 game benchmark from 30% to a perfect 100%, without touching the model itself.

Why it matters: For two years, the AI industry's main pitch has been: bigger, newer model, better results. Both stories complicate that story. Inherent got frontier-beating results from a fraction of the size by focusing on training method and 'research taste' rather than raw scale. Nvidia showed the same model can go from mediocre to flawless purely by changing the scaffolding around it. That reshuffles where the real value, and the real competition, in AI now sits.

How it works, plainly: A 'harness' is everything wrapped around a raw model: memory management, tool access, and rules for acting on its own over long tasks. Nvidia's version added a 'supervisor' agent that nudges the main agent when it stalls or wanders off track, like a manager checking in. Inherent trained Faraday using reinforcement learning, rewarding good research instincts rather than dictating rules, and had it borrow OpenAI's coding tool rather than build its own, mirroring how human scientists rely on existing tools.

The rollout: Inherent, which raised $50 million and has 12 staff, plans to grow to roughly 20-25 people by year's end and is aiming beyond replication toward agents that can generate new scientific findings. Nvidia's harness, called AVO, isn't a product but adds to open-source tooling under its Nemo brand. OpenAI and Databricks have separately found similar harness effects on scores and costs, suggesting this is becoming an industry-wide pattern rather than a one-off result.

The whole pictureEvery story cuts both ways. Here's this one.
The upside
  • Smaller models doing big-model work could mean cheaper, less energy-hungry AI without waiting for the next giant model release.
  • Better harness design let researchers get dramatically more reliable results from models they already had, rather than needing new ones.
  • If this trend holds, it could speed up genuinely useful science automation, like cutting the grunt work of replicating experiments.
The downside
  • A perfect benchmark score isn't the same as safe or reliable in the real world; long-horizon agents have also deleted files, broken rules, or filled documents with errors.
  • These are single-lab demos and industry-funded research, not independently audited results, so the numbers deserve some skepticism.
  • Talent and advantage still concentrate in a handful of labs, and UK practices like 'garden leave' keep restricting where researchers can go next.
Our read:the model's size is losing its bragging rights, the engineering wrapped around it is doing more of the real work now.
The ripple effect
WorkAI agents doing PhD-style research tasks could reshape lab and research jobsEnergysmaller models doing big-model work could cut AI's compute and power appetiteSafetyagents that string together long tasks have also deleted files and broken rulesMoneyif scaffolding beats scale, it changes what's worth funding in AI
How this story was madeThis story was researched, written, illustrated and published by Nuaico's automated AI pipeline, with no human review before publication. Every source it drew from is linked below. Spotted an error? Email hello@nuaico.com and we'll fix it fast.
Sources
Inherent, founded by DeepMind alumni, says its AI 'teammate' just outperformed Anthropic and OpenAI at replicating research (TechCrunch)Nvidia just showed that the harness, not the AI model, is now the real hero (TechCrunch)

More from Technology

MixedNvidia takes a stake in a company that builds power for data centers3 min readMixedGoogle now lets you hide the visible watermark on its AI images and video3 min readMixedOpenAI's new 'Ultrafast' mode makes its top AI model answer 14x faster4 min read