Which Model Should You Use in Claude Code?
Use the most capable model for work that needs reasoning across many files: architecture changes, hard debugging, and large refactors. Use a balanced model for everyday features, and a fast one for simple edits and repetitive tasks. Switch per task, and remember that good context often improves results more than a bigger model.
The tiers
Anthropic's Claude models are offered in tiers. The most capable tier handles the hardest reasoning, long multi-step tasks, and complex codebases best, at higher cost and latency. A balanced tier offers strong coding performance at lower cost and faster responses, and handles most everyday work. A fast tier is the cheapest and quickest, suited to simple tasks and high volumes. Specific model versions change regularly, so check which models your plan and Claude Code version offer.
Matching tasks to tiers
- Most capable: designing a new subsystem, debugging an intermittent failure across services, large refactors touching many modules, and reviewing risky changes.
- Balanced: implementing features in a known codebase, writing tests, fixing ordinary bugs, and most day-to-day work.
- Fast: renaming, formatting, small edits, generating boilerplate, explaining code, and subagent tasks such as searching or summarising files.
Switching models during a session is straightforward in Claude Code, so choose per task rather than per project.
Why context often beats model size
Many apparent model failures are context failures. The agent invents a helper because it did not find the existing one, uses the wrong convention because nobody told it, or misunderstands a system because the relevant design note was not in reach. A more capable model guesses better, but still guesses.
Giving the agent the right context, through a concise CLAUDE.md and retrieval over documentation and the codebase, often fixes these problems on a smaller model. It also cuts cost twice: fewer tokens per request and less need for the most expensive tier.
Cost considerations
Cost depends on the model's price per token and the number of tokens processed. Long sessions, large files read into context, and repeated re-reading of the same material multiply tokens. Prompt caching reduces the cost of repeated context, and retrieval reduces how much context is needed at all. Measuring usage per task type shows where a cheaper tier or better context would pay off.
Using different models for subagents
Subagents let you mix tiers inside one task. A capable model can plan and make the key decisions while delegating searches, file summaries, test runs, and simple edits to subagents on a faster model. Each subagent works in its own context and returns only its results, so the main session stays focused and the expensive model processes fewer tokens. Define subagents for the chores that recur most in your work.
Signs you need a stronger model
Escalate when the agent keeps proposing fixes that do not address the root cause, loses track of constraints across several files, produces designs that ignore obvious interactions between components, or needs repeated correction on the same point. Before escalating, check whether missing context is the real problem: if the information the agent needed was not available to it, a stronger model may still fail, while better context may let the current one succeed.
Evaluating a new model on your own work
When a new model is released, pick five to ten representative tasks from recent work: a bug fix, a feature, a refactor, a test-writing task, and an explanation of unfamiliar code. Run each on the current default and the new model with the same context, and compare correctness, the number of corrections needed, time, and cost. That small benchmark, built from your own codebase, is more informative than general leaderboards for deciding whether to change defaults.
A practical default
Default to the balanced tier for daily work. Escalate to the most capable tier when a task stalls, spans many parts of the system, or carries high risk. Delegate simple subtasks to the fast tier. Revisit the defaults when new models are released, by rerunning a few representative tasks from your own work and comparing results, rather than relying on general benchmarks.
Frequently asked questions
- Is the most powerful model always best for Claude Code?
- No. It is best for hard, cross-cutting problems, but slower and more expensive for routine work that a balanced or fast model handles well. Matching the model to the task keeps costs reasonable and latency low without losing quality where it matters.
- How do I switch models in Claude Code?
- Claude Code lets you choose the model for a session and switch during it, through a command or settings. Available models depend on your plan and version. Check the current documentation for the exact command and model names, since they change as new models are released.
- Why does Claude Code make mistakes even with the best model?
- Often because it lacks context: it cannot find existing code, does not know your conventions, or has not seen the relevant design decision. Better context through a concise CLAUDE.md, retrieval over docs and code, and small, clear tasks fixes many of these mistakes regardless of model.
- Does a bigger model use more tokens?
- Not necessarily more tokens, but each token typically costs more on a more capable model. Capable models sometimes complete tasks in fewer attempts, which can offset the higher price on hard problems. On routine tasks the extra capability is rarely needed, so the higher per-token cost is mostly waste.
- Should subagents use the same model as the main session?
- Not necessarily. Subagents doing searches, summaries, or simple edits often work well on a faster, cheaper model, while the main session uses a more capable one for planning and decisions. Assign models per subagent where your setup allows, and check results to confirm the faster model is reliable for each role.