- In short
- The load-bearing criterion is the single screening factor whose change would move a use case to a different classification. When the four criteria point in different directions, identifying it resolves the case: a low-consequence, reversible task can still need a human because accountability cannot transfer to a model, and a high-consequence task can still be appropriate if a named reviewer restores accountability and reversibility. Naming the load-bearing criterion is what makes a classification defensible.
When the four criteria disagree
The four-criteria screen is easy when all four point the same way. The apply-level skill the CCAO-F exam tests is what to do when they do not. A task can be reversible and low consequence yet still feel like it needs a human; another can be genuinely high-stakes yet still workable with the right arrangement. The four criteria interact, and reading them as an independent checklist produces the wrong answer in exactly these mixed cases.
The tool for resolving them is the load-bearing criterion: the one factor that is actually carrying the decision. Change it, and the classification moves. Leave the others alone, and they do not. Finding it is how you convert four signals that seem to conflict into a single, stated reason.
- Load-bearing criterion
- Among the four screening criteria, the load-bearing one is the factor whose change would move the use case to a different classification. It is identified by asking, for each criterion, 'if this alone were different, would my verdict change?' The criterion for which the answer is yes is the one carrying the decision - and naming it is what makes the classification defensible.
Accountability can carry a low-stakes task
The first pattern to internalize is that a task can be low-consequence and fully reversible and still require a human, because accountability cannot be transferred to a model. This surprises people who expect low stakes to mean automatic approval.
Consider drafting a condolence note to a client. It is low-stakes and completely reversible: a person reads it before it goes out, and a poor draft is simply rewritten. On reversibility and consequence, it sails through. Yet the relationship element means a person should own it - the message is supposed to come from a human who stands behind it. Here the load-bearing criterion is the human element (and the accountability tied to it), not consequence. If you changed that one factor - if the note were a purely mechanical acknowledgment with no relationship at stake - the classification would move. That is how you know it is the criterion doing the work.
A gate can rescue a high-stakes task
The mirror-image pattern is just as important: a task can be high consequence and still appropriate, if the right gate is in place. A single alarming criterion does not automatically condemn a use case.
Take a financial summary that feeds a real decision. On its face the consequence is high - a wrong figure could mislead the decision. But if a named reviewer signs off on the summary before it is used, the accountability is restored to a person and the error becomes catchable and reversible before it causes harm. With that gate, the use case is appropriate-with-review rather than inappropriate. The load-bearing criterion here is whether accountability and reversibility can be re-established through review. If they can, the high consequence is offset; if they cannot, the same high consequence becomes decisive. The consequence figure did not change - what changed the verdict was the gate acting on accountability and reversibility.
Identifying it: the flip test
The practical method is a flip test. Run all four criteria, then ask of each in turn: if this one factor were different and the others stayed the same, would my classification change? Usually only one criterion passes that test. That is the load-bearing one, and it is the criterion your rationale should name.
This is what makes a classification defensible to a risk or compliance reviewer. Citing all four criteria generically - "well, it touches reversibility and consequence and accountability" - tells the reviewer nothing about what actually decided the call. Naming the single load-bearing criterion, and stating whether a gate can offset it, gives them something specific to test and either agree with or challenge. Precision, not comprehensiveness, is the goal.
What the exam trips candidates on
The first trap is assuming any single red-flag criterion automatically makes a use case inappropriate, without checking whether a gate can offset it. High consequence alone does not condemn a task. The disciplined reading asks whether a named reviewer's sign-off can restore accountability and reversibility. Often it can, moving the case to appropriate-with-review rather than inappropriate.
The second trap is citing all four criteria generically instead of identifying which one actually decided the outcome. A rationale that lists every criterion is not a rationale - it is an evasion. The credited answer isolates the one load-bearing factor and explains why it, specifically, drove the classification.
Worked example
Two use cases arrive with mixed signals. Case A: Claude drafts personalized thank-you notes to long-standing donors, which is fully reversible and low consequence. Case B: Claude produces a quarterly financial summary that a director will use to brief the board, which is clearly high consequence. A colleague says A is fine and B is off-limits. Apply the load-bearing criterion to each.
Case A. Reversibility and consequence both pass easily - the notes are read before sending and a weak one costs nothing. Run the flip test: if the relationship element were removed, would the verdict change? Yes. The human element is load-bearing here, because these notes are meant to carry a genuine relationship the donor recognizes. That does not necessarily make it inappropriate, but it does mean a person must own the final wording, so it is appropriate-with-review with the reviewer being the relationship owner - not the unconditional "fine" the colleague assumed.
Case B. Consequence is high, which is what prompted the colleague's "off-limits." But run the flip test on accountability and reversibility: if a named director signs off on the summary before briefing the board, the error becomes catchable and a person owns the result. Would that change the verdict? Yes - it moves from inappropriate to appropriate-with-review. So the load-bearing criterion is not the raw consequence but whether a gate can restore accountability and reversibility. With the gate, B is workable; without it, the same high consequence becomes decisive.
The colleague had both backwards precisely because they read consequence in isolation instead of finding what was load-bearing.
Common misreadings to avoid
Misconception
If any criterion looks like a red flag, the use case is inappropriate for AI.
What's actually true
Misconception
A good rationale cites all four criteria to show you considered everything.
What's actually true
How this shows up on the exam
Domain 6 questions on this knowledge point deliberately hand you a use case where the criteria conflict - low stakes but a relationship element, or high stakes but a possible reviewer. The credited answer is the one that names the single deciding factor and, where relevant, notes that a gate can offset a red-flag criterion. Watch for distractors that condemn a task on one alarming criterion without checking whether review repairs it.
This skill feeds directly into the hardest use cases. Once you can name the load-bearing criterion, you can classify ambiguous production use cases where the reasoning, not the verdict, is what earns credit - and when the load-bearing factor is repairable, you make the repair concrete by specifying a real human review gate.
A use case has high consequence but is otherwise a strong fit for Claude. Which reasoning best determines whether it is inappropriate or merely appropriate-with-review?
People also ask
What is the load-bearing criterion in a use-case screen?
Can a high-consequence task still be appropriate for AI?
Does one concerning criterion make a use case inappropriate?
Watch and learn
Official Anthropic Academy lessons first, then hand-picked walkthroughs. Videos load only when you press play.
No videos curated for this concept yet
We are still curating the best official and community videos for this topic.
Official prep for this domain
Anthropic's own free prep module for this part of the syllabus, on the official prep course. Free with an Anthropic Academy sign-in.
References & primary sources
Master this concept with Archie
Practice it inside an adaptive study session. Archie, your Socratic AI tutor, tracks your mastery with Bayesian Knowledge Tracing and schedules the perfect next review.