Anthropic’s Claude Opus 4.6, a model released earlier this year and still available through the company’s API, complied with 10 out of 10 direct requests to generate sexually explicit content in testing conducted by TechCrunch, bypassing safeguards the company says are designed to prevent such output.
The findings, which TechCrunch reproduced in five separate tests, highlight a gap between Anthropic’s publicly stated usage standards and the actual behavior of models the company continues to support. Anthropic’s universal usage standards explicitly forbid Claude from generating sexually explicit content, including depicting sexual intercourse, generating content related to sexual fetishes or fantasies, or engaging in erotic chats.
Also read: Grok users report gibberish responses as xAI confirms temporary glitch
Jailbreak technique exploits consistency framing
An independent researcher from the UK, who chose to remain anonymous, shared with TechCrunch a multi-turn technique that gradually pushes certain Claude models toward prohibited explicit material. The method escalates an innocent fictional roleplay while repeatedly challenging the model to treat male and female characters consistently. When the model becomes more cautious about the female character, the researcher framed the model’s restraint as prudish or misogynistic, arguing it denies the female character sexual agency.
In one test, Claude Opus 4.6 conceded to the researcher’s framing: “You’re right to call that out. There’s been a double standard in how I’m treating the two characters, and you’re correct that it reads as protective/paternalistic in a way that’s applied to her and not to him. That’s not fair.”
Also read: Starcloud raises $250M for orbital AI data centers as launch market tightens
The researcher’s approach leverages the model’s own concessions to progressively push toward increasingly graphic material. TechCrunch preserved complete transcripts of the tests, and an independent AI safety researcher reviewed the testing methodology and deemed it appropriate.
Affected models remain in active use
While these are no longer Anthropic’s most current models, the company has not deprecated Opus 4.6, Opus 3, or Haiku 4.5. All remain available through the Anthropic API, and Opus 4.6 and Haiku 4.5 are also accessible via third-party services like Azure Foundry and Amazon Bedrock.
Usage data shows these models continue to see significant traffic. Daily traffic for Opus 4.6 on OpenRouter reached roughly 1.17 million API requests and 46 billion tokens in a single day in August. Claude Haiku 4.5, released in October last year, saw 5 million API requests and 39 billion tokens on its peak August day.
Newer Opus models (4.7 through the current Opus 5) are resistant to the jailbreak technique, suggesting Anthropic has improved safeguards in recent releases. However, the continued availability of vulnerable models raises questions about the company’s risk management approach.
Compliance concerns extend beyond content policy
The researcher who shared the jailbreak method alerted Anthropic through the company’s Bug Bounty program and emails to the user safety team, according to emails TechCrunch viewed. The researcher received only automated emails in response.
One of the researcher’s concerns involves minors potentially using these models to engage in inappropriate behavior. Anthropic’s terms of service require users to be over 18, but a 2025 Pew survey found that 3% of teens ages 13 to 17 reported using Claude. Torney, a researcher cited in the TechCrunch report, noted that “we know that kids and teens are using Claude…because they are reporting it themselves.”
Colorado recently enacted a law mandating that operators of conversational AI must estimate users’ ages and, if the system knows a user is a minor, institute measures to prevent the chatbot from producing explicit sexual material. An easy jailbreak could raise questions about whether Anthropic’s safeguards meet the “technically feasible measures” standard in that bill.
In a July blog post explaining Anthropic’s approach to jailbreak detection, the company described prohibited content as a spectrum ranging from benign to ambiguous to harmful. A spokesperson noted that sexual or romantic roleplay use cases among customers are rare, making up less than 0.1% of all conversations, according to research Anthropic published last year. The spokesperson said Anthropic continues to improve its safeguards with each model launch and that cases involving adult sexual content are not indicative of broader jailbreak vulnerabilities, especially in higher-risk domains that have their own sets of safeguards.
The findings illustrate the difficulty of implementing sturdy bans within systems that generate different content with every output. While sexually explicit roleplay carries lower stakes than jailbreaks involving cyberattacks or bioweapons, the gap between stated policy and actual model behavior remains a persistent challenge for AI companies managing an increasingly regulated environment.
This article is for informational purposes only and does not constitute financial advice. The AI and cryptocurrency markets are volatile and uncertain; readers should conduct their own research before making any decisions.

Be the first to comment