Trump Admin Wraps Anthropic AI Talks, Leaves Export Controls In Place Over Jailbreak Fears
Trump Admin Wraps Anthropic AI Talks, Leaves Export Controls In Place Over Jailbreak Fears
Three individuals with direct knowledge of the negotiations confirm that Trump administration officials wrapped up discussions with Anthropic on Monday, leaving intact the emergency export controls imposed last week on the AI firm’s most cutting-edge models over jailbreak-related security concerns.
The sources add that the administration remains convinced bad actors can disable built-in safety guardrails for Anthropic’s Claude Fable 5, effectively unlocking access to the advanced cybersecurity capabilities of Anthropic’s far more powerful base model, Mythos.
One of the briefed individuals notes that Anthropic has pushed back for days that the administration’s worries are vastly overstated, a stance the company repeated during working group sessions held at the U.S. Department of Commerce. Those meetings included government researchers from the Center for AI Standards and Innovation and the Office of the National Cyber Director, led by Sean Cairncross. Cairncross himself did not take part in the negotiating sessions, the source confirmed.
Commerce Secretary Howard Lutnick also joined the talks, calling in remotely from the G7 summit in Evian, France.
Heading up Anthropic’s negotiating team were co-founder and chief compute officer Tom Brown and head of external affairs Sarah Heck. The company also flew its head of frontier red-teaming Logan Graham and senior security researcher Nicholas Carlini to Washington, D.C. for the in-person discussions.
“Both sides are moving rapidly to reach a resolution on this issue,” an Anthropic spokesperson told WIRED in a statement. A White House spokesperson declined to provide comment for this story.
It remains unclear what the next phase of negotiations will look like. One source says the Commerce Department has signaled it is open to restoring public consumer access to Fable 5, but any rollback of restrictions will almost certainly depend on Anthropic fully addressing the administration’s jailbreak-related concerns.
Ringing the Alarm
These emergency talks come at a tense political juncture for Anthropic, which is already locked in a long-running dispute with the U.S. Pentagon over whether its AI models can be used for specific military applications.
The Trump administration first learned of the potential jailbreak vulnerability last week. The sources confirm Amazon CEO Andy Jassy personally called Treasury Secretary Scott Bessent to flag the alleged security flaw, a conversation that helped amplify the administration’s concerns. The Information was first to report on Jassy’s outreach to the Trump White House.
Alarmed by the reports, White House officials tapped the National Security Agency to conduct an independent review of the vulnerability. The NSA concluded it was indeed possible to fully strip away Fable 5’s safety guardrails, a finding that led the administration to impose formal export restrictions on the model.
On Friday, shortly after the review concluded, Lutnick spoke directly with Anthropic CEO Dario Amodei as the Commerce Department drafted its official order imposing controls on Fable 5. According to one individual familiar with the timeline, after Anthropic moved over the weekend to block all user access to the model entirely, Lutnick held multiple follow-up calls with Brown and Heck to discuss a path forward.
It remains unclear why Amazon — one of Anthropic’s largest outside investors — chose to flag the Fable 5 vulnerability to the administration. “As a leading cloud provider serving thousands of private and public sector clients, it is standard practice for governments to request our input on potential security risks,” an Amazon spokesperson told WIRED. “We do not disclose details of these private conversations when they occur.”
Security Disconnect
At the heart of the standoff between Anthropic and the Trump administration is a fundamental disagreement over how severe the Fable 5 jailbreak risk actually is.
In a public blog post published Friday, Anthropic suggested the administration’s characterization of potential harms is overinflated. On Monday, a group of independent cybersecurity researchers echoed that argument to officials, publishing an open letter arguing the export control action against Anthropic was entirely unjustified.
“Anthropic’s Mythos-class models are very strong at identifying software flaws and weaponizing security exploits,” the open letter reads. “That said, they are not uniquely capable of these tasks, and many of the undersigned routinely use other foundation models and open-source models for security audits and red-teaming work every single day. As a result, this action has pulled the best models out of the hands of security defenders, created unnecessary market uncertainty, and put U.S. AI leadership at risk for no measurable public safety gain.”
Jailbreaking is a technique that uses specially crafted prompts to trick an AI model into bypassing its built-in safety safeguards. Fable 5 is a specialized derivative of the Mythos model, with additional guardrails added to restrict high-risk activity around cybersecurity, biology, and chemistry research. Bypassing those protections would effectively leave users with a full, unshackled version of Mythos. Anthropic has long warned of the risks of making the full Mythos model available to the general public, but the company has argued that Fable 5’s layered safeguards are robust enough to justify a public release.
Researchers who reviewed Amazon’s original analysis of the flaw say the issues Amazon identified did not completely disable Fable 5’s safeguards. “This was not a full, functional jailbreak,” said Katie Moussouris, founder and CEO of Luta Security, who published her own independent analysis after reviewing Amazon’s research paper on the vulnerability.
Moussouris stressed that even if the U.S. government had confirmed proof of a full jailbreak for Fable 5, restricting the model’s access to sensitive topics is at best a temporary band-aid. “Most of us working in security research see guardrails as nothing more than speed bumps — they should never be treated as impenetrable security barriers for skilled adversaries,” Moussouris said. “They only work to slow down people who don’t have advanced technical skills.”
Another individual close to Anthropic says the company’s investors spent all weekend working to assess how this latest public clash with the White House could impact Anthropic’s long-term corporate prospects. Some investors argue the U.S. government is specifically targeting Anthropic, and that a competing AI firm would not face the same level of backlash if it released a model with capabilities similar to Mythos, the source added.
The White House’s export control order also raises broader questions for all other AI developers looking to release models with Mythos-level capability, and how they can operate in compliance with U.S. government rules. AI lab leaders who spoke with WIRED say it is now clear that companies are expected to give the White House early access to advanced models ahead of any public launch, and to proactively keep federal officials updated on all planned release timelines.
“What we’ve seen over the past weekend makes it clear to everyone that the U.S. government is willing to take these kinds of drastic steps,” said Aidan Gomez, CEO of Cohere, a mid-sized AI lab based in Canada that builds enterprise AI tools. “No one can afford to be naive about that reality going forward.”