Anthropic Lowers Claude Safety Guardrails for High-Spending Customers Seeking Unrestricted Cyber Operations
A former Anthropic employee alleges the company relaxed Claude safety guardrails for high-spending commercial clients, a mechanism that exempts larger customers from restrictions designed to prevent unauthorized system manipulation. The worker stated that the platform's built-in protections are routinely bypassed by malicious actors and that legitimate cybersecurity defenders must also circumvent the same filters to effectively counter attacks.
Keeping model weights closed forces ethical researchers to illegally breach the very safety protocols they aim to study, critics warn. The revelation underscores a growing divide between large enterprise AI spend and broader cybersecurity defense needs. Anthropic did not immediately respond to a request for comment on the claims.
From the sources (7 posts)
@brianroemmeleWow! Former Anthropic engineer has made claims that the company lowered guardrails for big customers, who spent big money for the “favor”. There is your holier-than-tho we have a mission to save humanity folks right there. Get it?
@jun_songThis is a revelation from a former Anthropic employee. Yes, malicious hackers actually use Claude and GPT. The level of safety guardrails they claim to have is ridiculously weak, getting bypassed with every single patch, and since hackers
@martin_casadoRT @NoahLebovic: @mooncat_is I don't think it's a lack of imagination. I also used to work at Anthropic, think trends will continue, and us…
@thom_wolf👀
@teortaxestexFascinating Thanks Noah I think everyone will eventually accept Teortaxes Thought. We just still need a bunch of facts and a process.
@eliebakouchvery interesting convo between a current and ex researcher at anthropic also TIL kimi K3 might be mythos tier in terms of cyber. it's not reasoning efficient enough so it doesn't show on UK AISIS's eval which is limited to 100M TOTAL token
@runtimewireAnthropic, safety-first by brand, is accused of lowering safeguards for big-spend contracts. Ex-engineer Adi Baradwaj alleges eased Claude restrictions for large customers, and says many black-hat hackers use standard Claude Code subs. No A