TechCrunch reported this week that an AI hallucination nearly triggered a US military operation, with a GovAI research scholar warning that service members need to understand the uncertainty baked into large language models. Sit with that for a second: an institution with some of the most rigorous verification protocols on Earth almost acted on a fabricated output. If that can happen inside the Pentagon's chain of command, it is absolutely happening inside your company's Slack channels, sales decks, and board memos right now -- just with lower stakes and less scrutiny.
The lesson for business readers isn't 'don't use AI.' It's that AI-generated output needs the same skepticism you'd apply to an intern's first draft, not the same trust you'd give a signed report. Every workflow that feeds an LLM's answer directly into a decision -- without a human check or a grounded data source -- is one confident hallucination away from an expensive mistake. This is exactly why we've argued that agent governance should focus less on hiring outside auditors and more on locking down what an agent is even allowed to touch or say, a point we made in our take on agent governance. If your internal tools generate answers from live business data instead of a model's memory, the hallucination risk drops considerably -- which is one reason internal tools built on structured, permissioned data are a safer bet than a chatbot bolted onto nothing.
The same TechCrunch report on the military incident dovetails with another story this week: Vals AI, backed by Andreessen Horowitz, is trying to become the neutral, trusted standard for benchmarking AI models. It's a telling bit of timing. We're at the point where there are so many models, so many vendor claims, and so much marketing noise that an independent scorecard has become its own startup category -- and a well-funded one. Vals is betting that buyers are tired of taking a lab's word for how good its own model is.
I think that bet is correct, but it comes with an obvious tension: a16z has money in dozens of the very AI companies whose models a 'neutral' benchmark would need to grade honestly. That doesn't automatically discredit Vals, but it means buyers should treat any benchmark, however well-intentioned, as one data point rather than gospel -- the same caution we raised about independent AI auditors sounding great in theory while their independence remains largely untested in practice. For a business deciding between AI vendors, the real benchmark is still your own data run through a real pilot, not a leaderboard. That's a big part of the argument behind buying vs. building your own tools rather than betting the business on someone else's marketing claims.
Two more Anthropic stories landed this week, and together they say something about how fast the biggest safety-focused lab is diversifying. TechCrunch reported Anthropic is now operating a biology lab running actual experiments, even as its own researchers continue warning publicly that AI poses existential risk. Separately, Accenture has reportedly become Anthropic's first embedded evaluator -- a consulting engagement TechCrunch calls Accenture's most high-risk yet. Running wet-lab biology and outsourcing your model evaluation to a global consultancy are very different moves, but both signal a company that's stopped treating 'AI safety' as a purely theoretical PR line and started treating it as an operating discipline with real partners and real budgets attached. Whether that discipline is rigorous enough to match Anthropic's own warnings is exactly the kind of thing an outside evaluator, or a benchmark like Vals, ought to be pressure-testing rather than taking on faith.
Which of these stories worries you more: an AI hallucination almost triggering a military operation, or the idea that a venture-backed startup might become the industry's main referee for judging AI models?
Sources