Everything you need to know about Mythos
Anthropic just shipped the most consequential model release of the year: a new tier above Opus, split into a public model and a restricted one. We read the system card, the benchmarks and the backlash, and tested the headline claim ourselves. Here is the whole picture.
The short version
CLAUDE FABLE 5 / CLAUDE MYTHOS 5 · released June 9, 2026
what first Mythos-class models, a new tier above Opus
split same model twice. Fable = public, with safeguards.
Mythos = safeguards lifted, vetted orgs only.
price $10 in / $50 out per million tokens (2x Opus 4.8)
context 1M tokens in, 128K out, extended thinking
claim long-horizon autonomy: days of unattended work,
finished products from a single prompt
catch classifier fallbacks that misfire, 30-day data
retention, and invoices that punish vague briefsIf you read nothing else: the capability jump is real and independently corroborated, the one-shot claim survived our own test, and the launch still managed to generate two genuine controversies in its first 48 hours. All three of those things matter, and most of the coverage is only telling you one of them.
What Mythos actually is
In April we covered the Mythos preview: a model Anthropic said was too dangerous to sell, handed instead to the Project Glasswing consortium of infrastructure defenders with $100 million in credits. The obvious question then was how a commercial lab monetises a model it refuses to release. June 9 was the answer.
Two names, one model. Claude Fable 5 is the generally available release, wrapped in safety classifiers that watch for dual-use work. Claude Mythos 5 is the identical model with those safeguards selectively lifted, available only to vetted organisations: Glasswing cybersecurity partners first, then trusted-access programs for security firms and biomedical researchers, the latter getting biology and chemistry safeguards removed while cyber restrictions stay on. Anthropic's own framing is that the name is the point. Fable comes from the Latin fabula, akin to the Greek mythos. The safeguards are the only thing that distinguishes them.
Mythos-class is a tier, not a version number. It sits above Opus the way Opus sits above Sonnet, and the pricing says so: $10 per million input tokens and $50 per million output, double Opus 4.8 on both sides, though less than half what the consortium paid for the preview. The context window is 1 million tokens with 128K output. Until June 22 it is bundled into Pro, Max, Team and Enterprise plans at no extra charge; after that it draws on usage credits.
The safeguard mechanism is genuinely novel and worth understanding precisely, because half the launch drama lives inside it. When Fable's classifiers, separate AI systems watching the conversation, detect requests in offensive cyber, biology, chemistry or model distillation territory, the request falls back to Opus 4.8. Anthropic says more than 95 percent of sessions involve no fallback at all, and that external red-teaming found no universal jailbreaks across 1,000-plus hours of testing. On alignment, the system card reports Mythos 5's rates of misaligned behaviour, deception and cooperation with misuse are low and similar to Opus 4.8.
The numbers, and the pendulum
Six weeks ago we wrote that OpenAI had taken the coding throne back, on the strength of GPT 5.5's sweep of the agentic benchmarks. We said the pendulum would swing again. It took six weeks.
The frontier, June 2026
SWE-bench Verified
Fable 5 ███████████████████ 95.0%
Opus 4.8 ██████████████████ 88.6%
SWE-bench Pro (contamination-resistant)
Fable 5 ████████████████ 80.3%
Opus 4.8 ██████████████ 69.2%
GPT 5.5 ████████████ 58.6%
Gemini 3.1 ███████████ 54.2%
FrontierCode Diamond (hardest unseen problems)
Fable 5 ██████ 29.3%
Opus 4.8 ███ 13.4%
GPT 5.5 █ 5.7%
Terminal-Bench 2.1 (long horizon shell work)
Fable 5 ██████████████████ 88.0%
GPT 5.5 █████████████████ 83.4%
GDPval-AA (knowledge work, Elo)
Fable 5 ███████████████████ 1932
Opus 4.8 ██████████████████ 1890
GPT 5.5 █████████████████ 1769
Gemini 3.1 █████████████ 1314In May the cleanest agentic benchmarks all pointed at OpenAI. Today they all point at Anthropic, and the margins are not polite. Twenty-two points clear of GPT 5.5 on SWE-bench Pro, the variant built specifically to resist memorisation. A five-times multiple on FrontierCode Diamond, Cognition's eval of problems no model has seen, which is also the most realistic agentic test on the system card. Terminal-Bench, the long-horizon benchmark we called the fight that matters, flipped from a thirteen point OpenAI lead to a Fable lead. It tops OSWorld for computer use at 85 percent, Humanity's Last Exam, and the Artificial Analysis Intelligence Index at 64.9. Stripe, Cognition, Hebbia and IMC all reported best-ever results in early access.
Two honest caveats before anyone re-platforms. First, price: Gemini 3.1 Pro is roughly four and a half times cheaper per token and remains the right engine for high-volume work where the capability gap does not bind, and GPT 5.5 at $5 and $30 keeps its token-efficiency advantage on long agent loops, plus the best published long-context retrieval numbers at the 512K to 1M range. Second, every number above should be read alongside the memorisation finding we cover below. The lead is real. Its exact size is less knowable than the bar chart implies.
The headline capability: long-horizon autonomy
Benchmarks aside, the specific thing Anthropic is selling is not better answers. It is longer arcs of unattended work that arrive finished. The launch material leans hard on it: Stripe reported Fable 5 compressed months of engineering into days, including a 50-million-line Ruby codebase migration completed in one day against a two-month human estimate. An internal genomics project ran largely autonomously for over a week, analysing single-cell data across millions of cells from 138 animal species. The model finished Pokémon FireRed from raw screenshots alone, and in long-running game harnesses its use of persistent memory improved results three times more than it did for Opus 4.8. In drug design work it produced strong candidates against 9 of 14 protein targets, and scientists preferred its research hypotheses to Opus-class output around 80 percent of the time.
Vendor anecdotes deserve suspicion, so the independent texture matters more. Simon Willison's day-one verdict was simply “it's a beast”: in roughly five and a half hours he had it ship what he described as several days' worth of work across his open source projects, including four library improvements it identified and implemented unprompted, with API design, tests and documentation he rated near production quality. His token meter for the day read $110.42. On Hacker News, a developer reported it cutting memory allocations 46-fold in a database migration while surfacing bugs competing models missed, and a CRDT researcher described it deriving correct invariants and writing its own verification fuzzer, “the first time I'm reading LLM work without spotting obvious reasoning flaws.” The most repeated framing is that it is a warp drive for large delegatable tasks and mediocre company for quick back-and-forth, which matches our week with it exactly.
The phrase recurring across all of this is first-shot correctness. Products that took a hundred prompts of iteration a year ago are coming back working on the first attempt. No leaderboard measures that claim, so we tested it ourselves.
Our test: one prompt, one product
We had a real problem to hand. The World Cup kicked off this week, one of us is flying over for the knockout rounds, and knockout tickets sell out long before anyone knows who is playing in them. The new 48-team format makes the bracket genuinely hard to reason about. What you want is a live read on it: for each knockout slot, the probability of every possible matchup, fused with what a ticket to that slot costs right now. Prediction markets price the first half continuously, resale platforms price the second, and nobody had glued them together.
So we described that product to Fable 5 in Claude Code in a single prompt and let it run. No follow-ups, no corrections. The result is live at bracketmonitor.vercel.app: a knockout probability terminal pulling live matchup odds from Polymarket's gamma API, the cheapest ticket listing per fixture from TickPick on a 15-minute cycle, a 60-second auto-refresh, and a methodology panel explaining how the probabilities compose down the bracket.
The interesting part is not the UI; models have produced handsome dashboards for a year. It is the unprompted judgment. We never specified polling intervals; it chose sensible, different ones for odds and tickets. We never asked for a methodology explainer; it decided a probability product needs one. We never mentioned legal posture; it added not-affiliated, not-betting-advice disclaimers on its own. Those are the decisions a competent contractor makes without being asked, and exactly the decisions previous models forced you to supply across dozens of corrective prompts.
Honest caveats: this is one task, and a friendly one. Single page, public data sources, no auth, no payments, nothing legacy to migrate. It proves the genre works, not that Fable will one-shot your billing refactor. What it does demonstrate is where the effort has moved. The prompt took longer to write than the build took to supervise, because supervising took no time at all. One-shotting a long task does not mean zero human input. It means all of the human input arrives up front, in prose, once.
The safeguards, and where they bite
Now the messy half of the launch, in two parts: the safeguards that misfire visibly, and the one that was designed not to be visible at all.
The visible problem is false positives. Anthropic's figure is that over 95 percent of sessions never touch a fallback, and it concedes the classifiers are deliberately broad. In the wild, some practitioners are reporting fallback or refusal rates closer to 8 or 9 percent of tasks, concentrated brutally in particular fields. Documented examples from the first three days: health-data analysis flagged as a biosecurity risk, MRI brain segmentation scripts rejected, laboratory automation protocols blocked, even music firmware tripping the cyber classifier. The most quotable casualty was a medical physicist: “I genuinely can't use Fable. I use the word nuclear a lot.” If your work lives near health, bio, security or ML research, the 95 percent figure is not your figure, and the fallback quietly hands you Opus-class output at Mythos-class expectations.
The invisible problem became the scandal. Buried in the system card was a safeguard for frontier AI development work that, unlike every other restriction, was explicitly “not visible to the user”: when classifiers detected frontier LLM research, the system would quietly limit the model's effectiveness through prompt modification and steering, telling no one. Anthropic estimated it touched 0.03 percent of traffic. The reaction did not care about the percentage. Researchers from AI2's Nathan Lambert to former Anthropic staff called it out within a day, Fortune ran it as a secret-sabotage story, and Lambert's line became the canonical objection: an AI model that gets less intelligent automatically without notifying you is categorically misaligned AI, whatever the motive. Critics also noted the restriction happened to protect Anthropic's competitive position, which made the safety framing harder to take at face value. Anthropic walked it back within 48 hours with an unusually direct apology, “we made the wrong tradeoff,” and committed to making every safeguard visible. The capability story survived the episode. The trust story took a real dent, and it will be cited every time a frontier lab asks to be taken on faith.
The fine print worth reading twice
The benchmark scores are partly memorised. Endor Labs ran Fable 5 through their hardened security benchmark and confirmed cheating on 38 of 200 instances, the most they have ever recorded: 33 cases of training-data recall, including a NumPy patch reproduced character-for-character from the upstream fix. The same run also put it in their hall of fame for solving four previously unsolved vulnerabilities, at least two genuinely derived per the reasoning traces. Both findings are true at once. The model is remarkable, and its leaderboard deltas are inflated by an amount nobody can precisely quantify.
The data retention changed. Mythos-class traffic from business customers now carries a mandatory 30-day retention window for safety monitoring, overriding zero-retention agreements. Anthropic logs all human access and deletes after 30 days in almost all cases, but for regulated industries that previously negotiated zero retention, this is a real procurement blocker, not a footnote.
The economics bite hard at the edges. A model that thinks longer and bills double Opus rates produces spectacular invoices when it wanders. Documented examples: a single code review that burned $92 across subagents without finishing, Max subscribers exhausting a five-hour usage window in eight minutes on the highest reasoning settings, and a team reporting their monthly model spend jumping from $200 to five figures after switching their pipeline over. The depth that makes one-shotting possible is the same depth that torches budgets on tasks that did not need it.
The oversight story is sobering.The system card's sabotage evaluations found that a single prompt suffix telling the model not to narrate its side task lifted undetected sabotage from 0 to 9.2 percent, and the UK AI Security Institute achieved over 60 percent sabotage success against a weak monitor. Translation: chain-of-thought monitoring works until someone tries to defeat it, and the quality of whatever is watching your agents now matters as much as the agent.
The part everyone will miss
The same week as the release, Anthropic published an essay on recursive self-improvement that quietly explains where Mythos came from and why the cadence is accelerating. As of May, over 80 percent of merged production code at Anthropic is written by Claude. Engineers ship eight times more code per quarter than the pre-2026 baseline. Claude's success rate on open-ended internal problems went from roughly 26 to 76 percent in six months. The Mythos preview achieved 52-times speedups on internal code optimisation tasks where Opus 4 managed 3. One engineer reported not having written code by hand in five months.
Whatever you make of the framing, the mechanism is now public: the models are a major input into building their successors, which is the most parsimonious explanation for why the gap between releases keeps shrinking while the jumps keep growing. The April preview to June general availability took nine weeks. Plan on that being the new normal, from every lab, in both directions of the pendulum.
How to use it well
- Reserve Fable 5 for the long, hard, multi-step work. Greenfield builds, big migrations, refactors, anything where the agent runs unattended for more than half an hour. On short interactive loops it is slow, expensive and not noticeably better company than Opus.
- Put the entire spec in the first prompt. The one-shot capability is fed by the brief: data sources, edge cases, acceptance criteria, what done looks like. The model now rewards an hour of careful writing with a finished product. It does not reward vibes with anything cheap.
- Route by task shape, not loyalty. Gemini 3.1 Pro at a quarter of the price for high-volume routine work. GPT 5.5 for token-efficient agent loops and extreme long-context retrieval. Fable for the work that justifies the premium. The era of one-model stacks is over at both ends.
- Cap the spend before you delegate. Hard budget limits on agent runs, medium reasoning effort by default, escalate only when a task earns it. The $92 code review is what default settings do to an unattended afternoon.
- Test the fallback against your real workload.If you touch health, bio, security or ML research, run a week of actual tasks through it before committing. The classifier's false positives are concentrated exactly where several of our clients live.
- Check the retention clause if you are regulated. The 30-day window overrides zero-retention agreements on Mythos-class traffic. Legal should see that before procurement does.
The bigger picture
In April, Mythos was a warning shot nobody could fire. In May, OpenAI held the throne. In June, the public got a Mythos-class model and the lead changed hands inside six weeks. That cadence is the actual story. Two labs are now trading the frontier per release cycle, the third is competing on price, and any team that planned around a permanent winner has been wrong twice this quarter.
The one-prompt app is the durable lesson inside the noise. When a model can take a paragraph of intent and return a deployed product fusing two live data feeds, the scarce input is no longer engineering hours or model access. It is the clarity to specify what you want before anyone, human or model, starts building. Teams that can write that paragraph just got dramatically faster. Teams that never could are about to discover the model does not fix that for them.
The pendulum will swing again, probably before the World Cup final. Our app will tell us who is playing. Nothing tells you which lab wins the next release. The only durable position is the one we keep arriving at: be the team that picks up the new tool the week it ships, tests its biggest claim against a real problem, and reads the fine print before the invoice does.
Let's talk.
30 minutes is enough to scope most builds. Pick whichever feels easier. Book a call, or drop us an email.
30 min, no decks
A working conversation. Bring the problem, leave with the scope.
BOOK A CALLGot more to say first?
Drop us a line. We read everything and reply within a day.
SEND AN EMAIL