Claude Opus vs Sonnet vs Haiku: Which Model Should You Actually Use?
Use Sonnet by default. Switch to Opus when the task is hard enough that being wrong would cost you something real. That single rule covers almost every decision your team will make, and the difference it makes is larger than most people expect — in workshops, the most common cause of “AI wasn't good enough for this” turns out to be a hard task run on a fast model.
Haiku is for volume, not for chat. Fable is for the rare problem that genuinely defeats Opus. And the model picker sits next to the send button, where most people have never clicked.
What the four tiers actually are
“Claude” is not one AI. It is a family of models sharing one chat box, and nothing on screen tells you which one is answering. The names come from writing — a haiku is short, a sonnet is structured and medium, an opus is a major work — and that ordering is the whole idea.
| Model | Built for | Working memory | Reach for it when |
|---|---|---|---|
| Haiku | Speed and volume | Around 200,000 words of <a href="/blog/claude-context-window-what-it-means-for-your-team">context</a> | Something simple runs thousands of times, usually inside an automation rather than a chat |
| Sonnet | Everyday knowledge work | Around 1,000,000 words | Almost everything: drafting, summarising, analysis, research, questions across your documents |
| Opus | Hard reasoning and long, multi-step work | Around 1,000,000 words | The task is genuinely difficult, or a missed detail is expensive |
| Fable | The most demanding reasoning Anthropic ships | Around 1,000,000 words | Opus has actually failed at this and the answer is worth the wait |
Each tier carries a version number — Sonnet 5, Opus 5, Haiku 4.5. The tier name tells you the model's role; the number tells you its generation. Choosing the right tier matters far more than chasing the latest number, and the tier is the only part of this most teams need to think about.
The context window is how much Claude can hold in mind at once — your uploaded documents, the conversation so far, everything. A million words is roughly a dozen full-length books. This is why “paste in the whole contract” works on Sonnet and Opus, and why Haiku is a poor choice for anything document-heavy.
Which model for which task
Skip the benchmarks. Match the model to what the task costs you if it comes out wrong. This is the table worth screenshotting for your team.
| The task | Model | Why |
|---|---|---|
| Draft an email, a proposal section, a post | Sonnet | Fast, and the draft needs editing rather than rescuing |
| Summarise a long thread, meeting, or report | Sonnet | Summarising is pattern work; the extra reasoning buys you nothing |
| Ask questions across a folder of documents | Sonnet | The large context window is doing the work here, not the reasoning |
| Review a contract for clauses that could hurt you | Opus | A missed clause is expensive; this is exactly what the extra care is for |
| Pressure-test a strategy or a business case | Opus | You want it finding the hole in your argument, not agreeing with you |
| Build something over many steps without supervision | Opus | Long, multi-step work is where the gap between tiers is widest |
| Anything going to a board, a regulator, or a client under your name | Opus | The cost of being wrong is the whole point |
| Sort thousands of tickets, tag leads, extract one field | Haiku | Simple, repetitive, high volume — speed is the requirement |
| A problem Opus has already tried and failed | Fable | The only honest reason to reach past Opus |
The one-line version to give your team: default to Sonnet, escalate to Opus when it is hard or high-stakes, let Haiku live inside your automations, and treat Fable as a last resort.
Opus vs Sonnet: the real difference
This is the comparison people actually want, and the honest answer is not “Opus is smarter.” On a single, well-specified task — write this email, summarise this document — you will often struggle to tell them apart, and Sonnet will be quicker. Paying for Opus there buys you a longer wait.
The gap opens on two specific things, and it is worth knowing them precisely:
- Depth on genuinely hard problems. When a task requires holding several constraints in mind at once and reasoning through them — a restructuring with tax, legal, and staffing implications; a technical decision with knock-on effects — Opus goes further before it settles. Sonnet reaches a reasonable answer faster. Opus reaches a better one.
- Long, unsupervised runs. Give either model a task with twenty steps and no one watching, and the difference is stark. Sonnet drifts. Opus keeps hold of the original goal, checks its own work, and finishes what it started rather than leaving a stub. If you are building anything that runs without a person in the loop, this is the whole ballgame.
So the question is not “which is better.” It is: is this task hard, or is it just long? Long is Sonnet. Hard is Opus. A forty-page document you want summarised is long. A forty-page document you want the risk in is hard.
The setting that matters more than the model
Here is the thing almost nobody on a non-technical team knows, and it is worth more than the entire model choice: you can tell Claude to think harder before it answers. Extended Thinking is the setting, and it is worth understanding properly.
Click the model name next to the send button and you will find not just the picker but the controls for how much effort Claude puts in. Turn that up and it stops answering with the first thing that comes to mind — it works the problem through, weighs options, checks its own logic, then writes.
I have watched this change the answer more often than switching from Sonnet to Opus does. A team convinced they needed the most expensive model usually needed the model they already had, thinking properly. Try that before you conclude you need to upgrade anything.
Four things teams get wrong
These four cost real money and real quality, and I see all of them in nearly every organisation we work with.
- “Always use the most powerful model.” This is the expensive mistake. You wait longer, you pay more, and on the majority of tasks you cannot tell the difference. Worse, people who default to the heavyweight stop noticing which tasks are actually hard — and that judgement is the skill worth having.
- “Haiku is the bad one.” Haiku is not a worse Sonnet. It is a specialist. For sorting ten thousand messages it is the right answer and Opus is the wrong one — you would be paying premium rates and waiting, for a job that needs neither.
- “Claude is inconsistent.” Two people ask the same hard question and get answers of visibly different quality, so the team concludes the tool is unreliable and quietly stops trusting it. Almost always they were on different models, or one had thinking turned up. This one belief does more damage to AI adoption than any genuine limitation of the technology.
- “I can't control this anyway.” You can. The picker is next to the send button. And if you are ever unsure mid-conversation, ask Claude which model it is — it will tell you.
What to do this week
Three moves, none of which take longer than ten minutes.
1. Find out what your team is defaulting to
Ask three people to open Claude and read out the model name next to the send button. If anyone is on an older or smaller model than they assumed, you have just found free capability sitting unused. This check takes two minutes and is the highest-value thing in this article.
2. Re-run one task that disappointed you
Pick something where Claude underperformed on genuinely hard work. Run it again on Opus, with thinking turned up. Compare the two answers side by side. Most teams discover that “AI can't do this” was really “we asked the fast model to do the hard thing.” Recalibrating that belief is worth more than any individual answer.
3. Agree one sentence and stop discussing it
Adopt a shared rule so nobody has to deliberate: “Default to Sonnet. Switch to Opus when it's hard or high-stakes. Haiku lives in our automations.” One sentence turns a recurring question into a reflex, and a reflex is what you actually want.
Getting the model right is the easy half. The harder half is knowing which of your team's tasks are genuinely hard — that judgement is what we build in the Deployed Kickstart, a half-day working session mapped to your real workflows. The Partner program keeps that judgement current as the models keep changing.
Frequently asked questions
What is the difference between Claude Opus and Sonnet?
Sonnet is the everyday workhorse and Opus is built for hard reasoning and long, multi-step work. On a single well-specified task — write this email, summarise this document — you will often struggle to tell them apart, and Sonnet is quicker. The gap opens on genuinely difficult problems that require holding several constraints in mind at once, and on long unsupervised runs where Sonnet drifts from the original goal and Opus keeps hold of it.
Which Claude model should I use by default?
Sonnet. It handles the overwhelming majority of real work quickly and well: drafting, summarising, analysis, research, and questions across your documents. Switch to Opus only when a task is genuinely hard or when a missed detail would be expensive, and let Haiku power background automations.
Is Opus always better than Sonnet?
No. Defaulting to the most powerful model is a common and expensive mistake — you wait longer, you pay more, and on most tasks you cannot tell the difference. The useful question is whether a task is hard or merely long. Long is Sonnet; hard is Opus. A forty-page document you want summarised is long. The same document, when you want the risk in it, is hard.
When should I use Claude Haiku?
When something simple runs thousands of times, usually inside an automation rather than a chat — sorting tickets, tagging leads, extracting one field from many forms. Haiku is a speed specialist, not a worse version of Sonnet. It also has a much smaller context window, so it is a poor choice for document-heavy work.
How do I know which Claude model I am using?
The model name sits next to the send button, and clicking it opens the picker along with the controls for how much effort Claude puts into an answer. If you are unsure mid-conversation, you can simply ask Claude which model it is and it will tell you.
What do the version numbers in Claude model names mean?
The tier name — Opus, Sonnet, Haiku — tells you the model’s role, and the number tells you its generation. Choosing the right tier matters far more than chasing the latest number, and the tier is the only part most teams need to think about.
Why does Claude seem inconsistent between colleagues?
Almost always because they are on different models, or one has thinking turned up and the other does not. Two people asking the same hard question on different tiers get visibly different quality, and teams often conclude the tool is unreliable. That belief does more damage to AI adoption than any genuine limitation of the technology.
Found this useful? Send it to someone who needs it.