- In short
- Model selection has two symmetric failure modes. Over-engineering uses the top tier for routine, well-structured work, wasting speed and budget without improving the outcome. Under-resourcing uses the fastest tier for ambiguous, high-stakes work, risking a lower-quality answer where it matters most. Both stem from not re-evaluating the task profile before choosing a tier, and correct selection sits where task stakes and structure meet the appropriate capability level.
Two ways to get it wrong
Model selection can fail in two opposite directions, and the CCAO-F exam tests, at the evaluate level, that you can name and avoid both. It is comforting to think there is a safe default -- "always use the most capable model" or "always use the fastest" -- but each of those blanket rules is one of the two failure modes in disguise. Over-engineering and under-resourcing are symmetric errors, and neither blanket rule escapes them.
Both errors share a single root cause: not re-evaluating the task profile before choosing a tier. When you skip that read and reach for a fixed answer, you will over-provision on the tasks that needed less and under-provision on the tasks that needed more. Correct selection is not a default at all; it is the point where the task's stakes and structure meet the matching capability level, found fresh each time. This is the evaluative capstone of matching tier to stakes.
- Over-engineering and under-resourcing
- The two symmetric failure modes of model selection. Over-engineering assigns the top tier to routine, well-structured work, wasting speed and budget for no better outcome. Under-resourcing assigns the fastest tier to ambiguous, high-stakes work, risking a lower-quality answer where it matters most. Both come from not re-evaluating the task profile; correct selection matches capability to the task's stakes and structure.
Over-engineering: capability the task never uses
Over-engineering is reaching too high on the spectrum. Putting routine, well-structured work on the top tier does not make the output better, because the task never taxes the extra reasoning the top tier provides. What it does is spend: the response is slower than it needed to be, and on metered access it draws more usage per call, multiplied by however many times the task runs. The outcome is unchanged; only the waste is added.
The reason over-engineering is tempting is the belief that more capability is free insurance. It is not. The extra capability has a price in speed and budget, and if the task cannot use it, that price buys nothing. This is why "always use the most capable model" is not a safe default with no downside -- the downside is real and recurring, just invisible if you only look at output quality. Recognising the wasted speed and budget as a genuine cost is the evaluative move.
Under-resourcing: a ceiling too low for the stakes
Under-resourcing is the mirror error: reaching too low on the spectrum. Putting ambiguous, high-stakes work on the fastest tier saves time, but the fast tier's lower ceiling on nuanced work risks a weaker answer precisely where a weak answer is most costly. The saving is real but trivial next to the downside; a faster mediocre answer on a high-consequence judgment is a bad trade.
"Always use the fastest model" fails for the same structural reason as its opposite: it ignores the task profile. It feels safe because it saves time, but on ambiguous, high-stakes work the time saved is dwarfed by the quality risked. Both blanket rules are seductive because each removes a decision, and both are wrong because the decision they remove -- reading the task's stakes and structure -- is the one that determines the right tier. The speed-versus-capability tradeoff is what makes both directions costly.
What the CCAO-F exam trips candidates on
Two errors are tested, and they are the two blanket rules. The first is framing "always use the most capable model" as a safe default with no real downside. The credited reading names the downside -- wasted speed and budget on routine work -- and rejects it as a default. The second is framing "always use the fastest model" as safe because it saves time, ignoring the task's stakes. The credited reading names the risk -- a weak answer on high-stakes work -- and rejects that too.
The unifying insight the exam rewards is that both rules fail for the same reason: they skip the task-profile read. Whenever a question offers a blanket "always use tier X" rule as the safe choice, the answer is to reject it in favour of matching the tier to each task's stakes and structure.
Worked example
Two managers set opposite policies. Manager A: 'Always use the top tier -- more capability can't hurt.' Manager B: 'Always use the fastest tier -- speed is always good.' Their shared workload is a mix of bulk data extraction and occasional high-stakes ambiguous strategy calls. Evaluate both policies.
Both policies fail, and they fail on opposite tasks for the same underlying reason: each replaces the task-profile read with a fixed rule. Trace each across the shared workload.
Manager A's "always the top tier" over-engineers the bulk data extraction. That work is routine and well-structured, so the top tier's extra capability has nothing to bite on -- the extractions come out the same, but slower, and on metered access costlier, multiplied across every item. A's belief that "more capability can't hurt" is exactly the myth the exam targets: the hurt is in wasted speed and budget, invisible only if you ignore cost and latency. On the strategy calls, A happens to be right, but by luck, not reasoning.
Manager B's "always the fastest tier" under-resources the high-stakes strategy calls. Those are ambiguous and consequential, where the fast tier's low ceiling risks a weaker answer precisely where quality matters most. The time B saves is trivial against that risk. On the bulk extraction, B happens to be right, again by luck.
Neither policy is safe, because each is a blanket rule that ignores the profile, and so each is correct only on the half of the workload that happens to suit its fixed tier. The right policy is neither: read each task's stakes and structure and match the tier -- fast, efficient extraction and top-tier strategy calls -- so both failure modes are avoided.
Common misreadings to avoid
Misconception
Always using the most capable model is a safe default with no downside.
What's actually true
Misconception
Always using the fastest model is safe because it saves time.
What's actually true
How this shows up on the exam
Questions offer a blanket "always use tier X" policy as the safe choice, or describe a task mis-provisioned in one direction. Reject the blanket rule and name the specific failure mode: over-engineering wastes speed and budget on routine work, under-resourcing risks quality on high-stakes work. The reliable answer re-reads the task profile and matches capability to stakes and structure.
This knowledge point is the evaluative capstone of task statement 3.3, drawing together the speed trade, matching tier to stakes, and combined entry-point and model matching.
One manager mandates the top tier for everything 'because more capability can't hurt'; another mandates the fastest tier 'because speed is always good.' The workload mixes bulk data extraction with occasional high-stakes ambiguous strategy calls. Which evaluation is correct?
People also ask
What is over-engineering in model selection?
What is under-resourcing a task?
Is always using the most capable model safe?
Watch and learn
Official Anthropic Academy lessons first, then hand-picked walkthroughs. Videos load only when you press play.
No videos curated for this concept yet
We are still curating the best official and community videos for this topic.
Official prep for this domain
Anthropic's own free prep module for this part of the syllabus, on the official prep course. Free with an Anthropic Academy sign-in.
References & primary sources
Master this concept with Archie
Practice it inside an adaptive study session. Archie, your Socratic AI tutor, tracks your mastery with Bayesian Knowledge Tracing and schedules the perfect next review.