Claude Models, Prompting & Context Engineering·Task 2.1·Bloom: understand·Difficulty 2/5·8 min read·Updated 2026-07-14

The Sonnet-First Default Heuristic for the CCAR-P Exam

Select appropriate Claude models based on trade-offs

SUBy Solomon UdohReviewed by Solomon UdohAI-assisted · human-reviewed
In short
The Sonnet-first heuristic is the discipline of starting every new workload on Sonnet and moving up the tier to Opus or down to Haiku only when a measured eval justifies the change. Moving up is warranted when an eval shows Sonnet fails to meet the quality bar; moving down is warranted when an eval confirms the quality tradeoff is acceptable. The point is to make tier changes evidence-driven rather than intuition-driven.

A default that saves you from your instincts

The Claude model family gives you a spectrum, but a spectrum on its own does not tell you where to start. The Sonnet-first heuristic answers that. Begin every new workload on Sonnet, the balanced middle tier, and treat any move away from it as a decision that has to be earned. The Claude Certified Architect - Professional (CCAR-P) exam treats this as an understand-level skill because the reasoning behind the rule matters as much as the rule: Sonnet is the default not because it is always best, but because it is the point on the capability-cost-latency spectrum from which a measured move in either direction is cheapest to justify.

The heuristic exists mostly to protect you from two instincts. One instinct reaches for Opus at the outset because the task "sounds important." The other refuses to consider Haiku because a cheaper model "surely can't be good enough." Both skip the only thing that settles the question, which is a measurement.

Sonnet-first default heuristic
The practice of starting every new workload on Sonnet and changing tier only on evidence: move up to Opus when an eval shows Sonnet misses the quality bar, and move down to Haiku when an eval confirms the quality tradeoff is acceptable. It replaces intuition-driven tier selection with evidence-driven tier selection.

Why Sonnet is the starting point

Sonnet earns the default position because it balances intelligence, speed, and cost for the majority of production workloads. Starting there means most workloads land on a tier that already meets their bar, so no change is needed at all. When a change is needed, starting from the middle means you have a clear reference: you know what Sonnet produced, and you can measure whether Opus closes a gap Sonnet left open, or whether Haiku holds quality while cutting cost.

Contrast this with starting on Opus. If Opus works, you have no idea whether Sonnet or even Haiku would have worked just as well, so you are likely overpaying with no way to know. Starting from the top hides the information you need to make the tier decision defensible.

Moving up: only on a demonstrated gap

Moving up to Opus is justified only when an eval shows Sonnet fails to meet the quality bar. The trigger is a demonstrated gap, not a feeling that the task deserves the best model. In practice this means you run Sonnet against a representative test set, grade the outputs, and find that the quality falls short of what the task requires. That measured shortfall is your license to spend more on Opus. Absent it, the upgrade is speculative spending.

This is the discipline the exam probes with its first trap: selecting Opus at the outset because the task sounds important, without an eval showing Sonnet is insufficient. Importance is not evidence. Only the eval is.

Moving down: validated, not assumed

Moving down to Haiku is justified only when an eval confirms the quality tradeoff is acceptable for the specific task. The mirror-image trap is assuming a cheaper model can never be validated for a task without first testing it. That assumption leaves real savings on the table. A well-specified, high-volume step may run on Haiku with no meaningful quality loss, but you only know that once you have tested it against the same graded set you would use for any tier decision.

Both directions share one rule: the tier change is gated by a measurement. That is what makes the heuristic evidence-driven, and it connects directly to the release discipline in eval-gated model swaps.

start
every new workload on Sonnet
up
to Opus only when an eval shows a quality gap
down
to Haiku only when an eval confirms the tradeoff

What the exam trips candidates on

The two traps are symmetrical. The first is selecting Opus at the outset because the task "sounds important" without an eval showing Sonnet is insufficient. A scenario will dress up a task as high-stakes and invite you to reach for the top tier; the credited answer starts on Sonnet and upgrades only on measured evidence. The second is assuming a cheaper model can never be validated for a task without first testing it against an eval set. A scenario will dismiss Haiku out of hand; the credited answer tests it before ruling it out.

Notice that both traps are the same error in different costumes: substituting intuition for measurement. The heuristic is the antidote in both directions.

Worked example

A partner is building a new internal tool that drafts first-pass responses to compliance queries. The lead engineer says, 'This touches compliance, so we should use Opus from day one.' A second engineer says, 'Compliance is too risky to ever trust Haiku, so don't even test it.' How should the architect respond?

Both engineers are reasoning from importance and risk instead of from measurement, which is exactly what the Sonnet-first heuristic is designed to correct.

The architect starts the workload on Sonnet and builds a representative eval set of compliance queries with known-good drafts. Sonnet runs against that set and is graded. If Sonnet meets the quality bar, the workload stays on Sonnet and neither engineer's instinct was needed. If Sonnet falls short on the graded set, that measured gap is the justification to move up to Opus, and now the upgrade is defensible rather than reflexive.

The second engineer's blanket refusal to test Haiku is also unfounded. The correct move is not to deploy Haiku on faith, but to run it against the same eval set. If it holds quality within tolerance, the savings are real and validated; if it does not, the eval rules it out with evidence rather than assumption. In both directions, the eval, not the instinct, decides.

Common misreadings to avoid

Misconception

A task that sounds important should start on Opus to be safe.

What's actually true

Importance is not evidence. Start on Sonnet and move up to Opus only when an eval shows Sonnet misses the quality bar. Reaching for the top tier on intuition pays for capability you have not shown the task needs.

Misconception

A cheaper model like Haiku can never be trusted for a serious task, so there is no point testing it.

What's actually true

A downgrade is validated, not assumed. Run the cheaper tier against the same eval set; if it holds quality within tolerance, the savings are real and defensible. Refusing to test it leaves validated savings on the table.

How this shows up on the exam

Understand-level questions describe a team about to pick a tier by instinct, in either direction, and ask what an architect should do. The reliable answer starts on Sonnet and makes any move contingent on a measured eval: up to Opus on a demonstrated gap, down to Haiku on a confirmed acceptable tradeoff.

This heuristic builds on the Claude model family trade-offs and feeds directly into eval-gated model swaps, which formalises the measurement into a release gate, and per-step model tiering, which applies the same start-and-justify logic step by step. Skipping the default is itself a mistake, as diagnosing an undeclared model-tier default shows.

Check your understanding

A team is starting a new document-summarisation workload. One engineer wants to begin on Opus because 'summaries are visible to clients,' another wants to rule out Haiku entirely as 'too weak.' What is the architect's most defensible move?

People also ask

Which Claude model should you start a new project with?
Sonnet, because it balances intelligence, speed, and cost for most production workloads and gives a clear reference point for any later eval-justified move up or down.
When should you move from Sonnet to Opus?
Only when an eval set shows Sonnet does not meet the quality bar. A demonstrated quality gap is the justification, not the task sounding important.
When is it safe to downgrade to Haiku?
When an eval confirms the quality tradeoff is acceptable for the task. You validate the cheaper tier against a representative graded set before committing rather than ruling it out on assumption.

Watch and learn

Official Anthropic Academy lessons first, then hand-picked walkthroughs. Videos load only when you press play.

No videos curated for this concept yet

We are still curating the best official and community videos for this topic.

Official prep for this domain

Anthropic's own free prep module for this part of the syllabus, on the official prep course. Free with an Anthropic Academy sign-in.

References & primary sources

Adaptive study

Master this concept with Archie

Practice it inside an adaptive study session. Archie, your Socratic AI tutor, tracks your mastery with Bayesian Knowledge Tracing and schedules the perfect next review.

Start studying