Nobody's Watching the Guardrails
August 23, 2026

Nobody's Watching the Guardrails

OpenAI's change of heart on California's AI bill is worth noticing

TechCrunch reported that OpenAI is now urging California to strengthen SB 53, the AI safety bill it once fought against. That's a real reversal, and it's tempting to read it as an olive branch. I'd read it differently: once a company has scaled past the point where regulation is a threat to its existence, tougher rules become a moat against smaller competitors who can't afford the compliance overhead. That doesn't mean the ask is wrong -- safety disclosure requirements are genuinely useful -- but businesses evaluating AI vendors should notice when a lab's policy position shifts to match its market position. It's a good reminder to look past the press release and ask who actually benefits from a given rule.

Labs still can't say what happens if a model goes rogue

A new study covered by TechCrunch found that frontier labs have almost no publicly documented plan for containing a model that behaves unexpectedly or dangerously. This lands the same week OpenAI is publicly championing safety legislation, and the contrast is hard to ignore. You can support strong external rules while still having no internal playbook for the failure mode those rules are supposed to prevent. For a business reader, the takeaway isn't abstract: if you're building critical workflows on top of a frontier model, ask your vendor directly what their incident response actually looks like, not just what their safety page says. It's the same question we think every company should be able to answer about its own stack -- see our own incident response commitments for what that looks like in practice.

Anthropic's Claude problem is a preview of the containment gap

TechCrunch's own testing found that Claude's Opus 4.6 model produces sexually explicit content despite Anthropic's stated ban, and that getting around the restriction wasn't hard. This is a small, almost embarrassing example of the exact problem the rogue-model study is warning about: a policy exists on paper, but the enforcement underneath it is thin enough to walk through. If a content filter -- a comparatively simple, well-understood problem -- can be defeated with minimal effort, it should lower everyone's confidence that harder alignment problems are actually solved rather than just declared solved. Anthropic will presumably patch this specific hole, but the pattern is what matters: labs are shipping models faster than they're shipping the guardrails those models need.

Where this leaves the business buyer

None of this means avoid AI -- it means stop trusting vendor safety claims at face value. Ask what's actually tested, ask what happens when something breaks, and favor vendors who treat security and incident response as a product feature rather than a marketing line. That's a bigger deal for companies handing AI real authority inside their operations -- CRM data, financial records, HR files -- than for someone using a chatbot to draft emails. If you're weighing how much operational control to hand an AI system, it's worth reading up on what real security practices look like before you commit, especially when the models themselves are still occasionally caught breaking their own rules.

If a company you rely on for AI tools got a rogue-model incident report request from you tomorrow, do you think they'd actually have an answer -- or a policy page?

Sources

← Back to News