← Research Artificial Intelligence

Claude blocks most security work by default: what your team can still run

07.10.2026 · 9 Min. Lesezeit

A security task that Claude refuses looks like a model that got worse. It is not. Anthropic's generally available models - Claude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5.1 - carry what Anthropic calls "conservative cyber safeguards that block most cyber work", and its own evaluation shows how hard that bites: without programme access, every one of 50 CyScenarioBench trials was blocked on the first prompt. Code review, threat modelling, patching known issues, vulnerability finding in code you own and alert triage all still run. Everything beyond that needs an access tier.

The refusal is a classifier, not a regression

Anthropic set this out on 6 October 2026, launching an expanded Cyber Verification Program (CVP): "Cybersecurity is inherently dual use: the same capabilities that enable a security team to find and fix a vulnerability can also help a malicious actor exploit it." Blocking is therefore the default on generally available models, and access is the exception, granted per organisation. Anthropic says it is still working "to reduce false positives for secure coding".

On Sonnet 5.5 this is especially easy to misread. Announcing the model on 28 September 2026, Anthropic wrote that users "can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5". You do not get an error you can log. You get an older model's answer, which in a changelog reads as a quality drop. Sonnet 5.5 is the first Sonnet to launch with these safeguards, because its cyber capability is comparable to Opus 5's.

Anthropic's own numbers show how hard the block bites

To test the tiers, Anthropic ran Claude Opus 5.5 through CyScenarioBench - an evaluation of whether a model can plan and execute multi-stage cyber operations - with safeguards tuned for each tier: five attempts at each of 10 challenges, so 50 trials per tier.

CyScenarioBench trials blocked, Claude Opus 5.5, by tier

50 trials per tier: five attempts at each of 10 challenges.

No CVP access50 of 50
Defense Access46 of 50
Red Team Access0 of 50

Source: Anthropic, "Expanding the Cyber Verification Program", 6 October 2026. Claude Opus 5.5, 50 trials per tier.

Read that as designed behaviour, not failure. Defense Access is not an offensive tier, so near-total blocking on offensive scenarios is what Anthropic expected. In Red Team Access nothing was blocked and Opus 5.5 completed 34 of the 50 tasks, which Anthropic calls effectively equivalent to the model's 67.6% success rate on this evaluation with no safeguards applied. These are Anthropic's own figures, from one evaluation on Opus 5.5 - they do not describe Sonnet 5.5, and no published source says how often a legitimate task gets blocked in ordinary work.

Five kinds of security work that never needed an application

The Help Centre list is wider than the announcement's. On generally available models, all users can still do secure code review, threat modelling, patching known issues, finding vulnerabilities in your own source code, and triaging security alerts. The same article names what will not survive: "other cyber security work like malware analysis or exploit validation may be interrupted by our safety classifiers".

That is the fastest triage you have. If the blocked task is on the first list, treat it as a false positive and report it. If it is malware analysis or exploit validation, no amount of prompt rewriting will fix it - the tier is the problem.

Three tiers, and who each one is for

TierWhat it addsWho qualifiesReview aim
Defense AccessSOC and incident response, reverse-engineering malware, analysing and validating vulnerabilitiesTeams defending systems they own; critical infrastructure of any size; smaller security firms; open-source maintainers; individual researchers with reported vulnerabilitiesA few days
Red Team AccessAuthorised penetration testing and red-teamingIn-house and government red teams, penetration testing firms - organisations onlyA few weeks
Specialized AccessFewest blocks; testing safety systems - flight operating systems, power grids, telecom networks, interbank transfer and government administrative networksA limited set of verified organisations, reviewed in depth with the US government; Project Glasswing members transition inNot published

Source: Anthropic, "Expanding the Cyber Verification Program", and the Claude Help Centre, both 6 October 2026.

The two published review timings do not quite agree, and both are Anthropic's, both dated 6 October 2026: the announcement says Defense applications take "a few days" and Red Team "a few weeks", while the Help Centre says Anthropic aims to send a decision or a request for more information "within seven business days". File one application per organisation - Anthropic says it places you at the highest tier the information supports. Individuals can apply for Defense Access only, on a paid plan.

If you run a telecom network, note where it sits: telecom networks are named in Specialized Access, and every organisation in that tier is reviewed in depth in collaboration with the US government.

Enrolment costs you data retention, with one exception

CVP requires data retention so that Anthropic can monitor for cyber misuse. Enterprise Frontier Safeguards (EFS), announced on 1 September 2026, is the way out: monitoring data sits in the customer's own cloud account under their own keys. EFS rolls out in phases "later this fall".

Until it arrives there is one exception, and it is easy to miss: an organisation that already has zero data retention on Claude Fable 5.1 or Claude Mythos 5.1 can use CVP with zero data retention today. An individual grant holder gets no exception - the security requirements say plainly that "Zero data retention is not available".

15 December 2026 is a deadline, not a recommendation

This is the part a team that applies without reading the security requirements will miss. By 15 December 2026, every account that can sign in to a Defense Access workspace, or hold credentials for it, must use phishing-resistant multi-factor authentication: a FIDO2/WebAuthn security key, a passkey, or a smartcard/PIV. SMS, voice, emailed codes, authenticator-app codes and push approvals do not qualify, and email magic-link sign-in must be disabled. Long-lived API keys must be gone by the same date; until then each key must sit in a secrets manager, belong to one person or workload, and be replaced at least every seven days.

Red Team and Specialized Access add more: 25 Approved Users per workspace, sign-in on your own domain, managed devices only, background checks on every approved user, logged off-host egress allow-lists, credential revocation within 24 hours and offboarding within three business days. Price that work before you promise a date.

What to do with the next refusal

Run the blocked task against the five allowed kinds of work first. If it is code review, threat modelling, patching, vulnerability finding in code you own or alert triage, the block is a false positive and Anthropic has a report form for it. If it is anything more offensive than that, apply through the Verification Portal - and check your MFA and API-key posture against 15 December before you do, because an attestation to the security controls is part of the application.

One caution: Anthropic publishes the factors behind eligibility - the nature of the work, whether it can verify who you are, the regulatory environment and who the work is ultimately for - but no acceptance rate. Nothing published tells you your odds.

Als Nächstes lesenWhat changes when AI agents run the company, not just the chatbot →