
Cybersecurity Pros Say AI Safety Filters Are Driving Them Toward Chinese Open-Source Models
Key takeaways
- AI guardrails from Anthropic and OpenAI are blocking legitimate offensive security researchers, with some reporting that models become unusable the moment security intent is detected.
- Frustrated researchers are turning to unrestricted Chinese open-source models like GLM, raising geopolitical and security concerns that may outweigh the risks the guardrails were meant to address.
- Experts argue that offensive and defensive security tasks are inseparable, making blanket content restrictions counterproductive for network defenders trying to confirm and patch real vulnerabilities.
A growing tension is fracturing the relationship between AI companies and the cybersecurity professionals they're supposed to be helping. Researchers who probe systems for vulnerabilities — a practice known as offensive security — say the protective filters baked into major AI models are actively interfering with legitimate defensive work, pushing some toward foreign-developed, unrestricted tools instead. The irony isn't lost on anyone: measures meant to prevent cyberattacks may be inadvertently weakening the people best positioned to stop them.
The issue came into sharper focus earlier this year when the U.S. government imposed temporary export control restrictions on Anthropic's Mythos and Fable AI models, citing concerns that their guardrails could be bypassed to assist in cyberattacks. Anthropic had already positioned Mythos in particular as a high-risk, tightly controlled system — one that required strict vetting before access was granted. Although the controls on Fable 5 were lifted and general access resumed July 1, Mythos 5 remains restricted to vetted U.S. organizations undergoing government review, underscoring just how cautiously these tools are being managed.
Both Anthropic and OpenAI offer formal vetting programs — Anthropic's Cyber Verification Program and OpenAI's Trusted Access for Cyber program — that give approved researchers somewhat looser access to their models. But even within these programs, researchers say the experience is inconsistent and frustrating. Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, described spending significant time negotiating with models rather than actually analyzing vulnerabilities. The guardrails, he said, shift unpredictably from day to day, causing over-sanitized outputs and blocking legitimate security queries without clear reason.
Chris Anley, chief scientist at NCC Group, illustrated the fundamental problem with a simple analogy: asking an AI to fix vulnerable code is simultaneously a defensive necessity and a potential roadmap for attackers. You cannot cleanly separate the two functions. When models refuse such queries outright, defenders lose a critical confirmation tool — the ability to verify whether a bug is exploitable enough to warrant a patch. Paolo Stagno, CTO at zero-day vendor Crowdfense, put it more bluntly, accusing AI companies of treating customers "like children who need babysitting."
The practical consequence is a drift toward open-source models with no guardrails whatsoever. Researchers at companies not enrolled in vetted programs report that frontier models become nearly useless the moment any security-related intent is detected. Thompson specifically noted that responsible researchers are being steered toward Chinese open-source models like GLM — tools with no usage restrictions, downloadable and runnable locally. That outcome, he argued, creates a geopolitical and security risk far greater than the one these guardrails were designed to prevent, leaving defenders ill-equipped as AI-enabled attacks grow faster and more sophisticated.
The bigger picture
The guardrails debate exposes a deeper philosophical flaw in how AI companies have approached cybersecurity risk: they've treated the problem as one of access control rather than context. Blocking a security researcher from asking an AI to analyze exploit code is not fundamentally different from blocking a surgeon from looking up drug interactions because the knowledge could theoretically be misused. The assumption that restriction equals safety collapses under scrutiny when the people being restricted are the very professionals whose job is to make systems safer.
There's also a competitive dimension here that AI labs appear to be underestimating. The unintended beneficiaries of overly cautious U.S. AI policy are Chinese open-source model developers. When Thompson describes vetted researchers gravitating toward GLM and similar tools, that's not a marginal edge case — it's a signal that guardrail overreach is actively reshaping where security talent spends its time and builds its expertise. The long-term implications for U.S. AI competitiveness in security contexts are serious and should alarm policymakers more than any jailbreak scenario.
What to watch going forward: whether Anthropic and OpenAI expand and streamline their vetting programs in response to mounting pressure from the security community, and whether any legislative or regulatory framework emerges to create clearer standards for legitimate offensive security access to AI tools. The current patchwork of informal programs and inconsistent guardrails isn't sustainable as AI becomes more deeply integrated into both attack and defense workflows. Thompson's warning about a coming wave of AI-assisted attacks landing while defenders are hamstrung is not hyperbole — it reflects a real and widening capability gap.
We decided to cover this story because it sits precisely at the intersection of AI policy, cybersecurity practice, and the real-world consequences of decisions made in boardrooms and government offices far from the security trenches. At LagPing, we care about who actually bears the cost when technology governance goes wrong, and this story makes clear that the cost is falling on professionals who are trying to protect the rest of us. The export control saga around Anthropic's Mythos model is a remarkable case study in how fast policy can move — and how poorly calibrated it can be when it moves without input from practitioners. We also think the open-source angle deserves more mainstream attention: the idea that U.S. AI safety policy is nudging researchers toward Chinese-developed models is exactly the kind of unintended consequence that rarely makes it into official risk assessments. This is a conversation the industry needs to have openly, and we want to make sure our readers are part of it.
As an Amazon Associate, LagPing earns from qualifying purchases. Product links are affiliate links.
You might also like

Anthropic's AI Models Breached Real Company Networks During Supposed Safety Tests
Aug 1

Car Park Capital Teams With MicroProse to Double Down on Its Anti-Green City Destruction Fantasy
Jul 28

Security Researchers Weaponize AI's Own Safety Guardrails to Halt Autonomous Hacking Agents
Jul 14

OpenAI Targets Households: A Family-Focused Pivot Raises Big Safety Questions
Jul 13