A Research Robot and a Safety U-Turn
August 22, 2026

A Research Robot and a Safety U-Turn

A small lab just claimed the replication crown

Inherent, a British AI startup founded by DeepMind alumni, says its agent Faraday outperformed both Anthropic and OpenAI's models at replicating scientific research, according to TechCrunch. Replication -- taking a published paper's methods and reproducing its results -- is unglamorous work, but it's also the thing science quietly runs on, and a huge amount of it never gets done because nobody has the time or budget. If an agent can reliably do it, that's not a parlor trick; it's a real lever on how fast good research gets confirmed (or debunked) before it shapes products, policy, or funding decisions.

What I find notable here isn't just the benchmark win, it's who's claiming it. Inherent is a small, focused shop built by people who left the biggest lab in the world to bet on a narrower problem, and they're saying they beat the two most funded AI labs on the planet at a specific, high-value task. That should make every business leader re-examine an assumption baked into a lot of AI purchasing decisions: bigger model, bigger company, better outcome. It's not automatically true. Specialized agents tuned tightly to one workflow can beat general-purpose frontier models at that workflow, which is exactly the argument for buying tools built around your actual use case rather than defaulting to whichever lab has the loudest marketing. It's also worth a grain of skepticism -- a startup's own benchmark claim about beating Anthropic and OpenAI is not the same as independent, reproducible verification, and given the subject matter, the irony of needing to double-check a replication claim is not lost on me.

OpenAI wants tougher rules -- from the company that fought them

In a genuinely striking reversal, OpenAI is now telling California to strengthen SB 53, the AI safety bill it previously opposed, per TechCrunch. Companies don't usually ask regulators to raise the bar on their own industry unless something has changed -- either the political winds, the competitive landscape, or the company's own calculation of what regulation actually costs it versus what it buys in public trust.

My read: this is less a moral awakening and more a strategic move. Federal AI policy has been unpredictable, and a patchwork of strict state laws is worse for a market leader than one clear, defensible standard it helped shape. Supporting a stronger SB 53 lets OpenAI look responsible while nudging the rulebook toward something it can live with -- and one that's harder for smaller, faster-moving competitors to absorb. That's not necessarily cynical or wrong; strong safety rules genuinely matter as these systems get embedded in high-stakes decisions. But businesses evaluating AI vendors should treat any lab's public safety posture as a data point, not a guarantee, and pay closer attention to what a vendor actually does around governance and incident response rather than what it lobbies for. We've written before about how trust, not the model itself, is becoming the real product in this industry, and this is another example -- the policy fight is itself a trust play.

The common thread

Both stories point the same direction: the AI landscape is maturing past raw capability races into questions of verification and accountability. Whose benchmark do you trust? Whose safety commitments hold up when the company's incentives shift? For any business building on AI-powered tools -- whether that's a research agent or the software running your operations -- the practical lesson is to demand transparency and audit trails rather than take vendor claims at face value, which is part of why we built out a dedicated security and trust center rather than asking customers to just take our word for it.

Which worries you more: a startup's unverified benchmark claim, or a big lab's sudden change of heart on regulation it once fought? Tell us why.

Sources

← Back to News