- In short
- The three use-case classifications are the possible outcomes of screening a task against the Delegation criteria: fully appropriate for AI (reversible, low consequence, no special human element), appropriate only with a defined human review step, or inappropriate for AI to perform at all. Each classification should be paired with a written rationale naming which criterion drove it, so the decision can be defended later.
From screen to verdict
Running the four-criteria screen produces a judgment, and that judgment has to land somewhere concrete. On the CCAO-F exam, every proposed use case sorts into exactly one of three classifications. This is the step that turns "I weighed reversibility, consequence, human element, and accountability" into a decision someone else can act on and audit.
The three outcomes are not arbitrary labels. Each corresponds to a different relationship between the task and the human oversight it needs: none beyond ordinary review, a specific checkpoint, or a hard stop where a person must own the work instead. Learning to name the outcome, and to name it for a stated reason, is the whole competency here.
- The three use-case classifications
- The possible results of screening a task: (1) Fully appropriate - reversible, low consequence, no special human element; delegate with normal review. (2) Appropriate with human review - useful for AI, but stakes or accountability require a defined human gate before the output is used. (3) Inappropriate - consequence, irreversibility, or a human-element requirement means AI should not perform it, and a human role must own it. Each comes with a written rationale.
Fully appropriate and inappropriate: the two ends
The two ends of the scale are the easier calls, and it helps to fix them first.
A fully appropriate use case is reversible, low consequence, and needs no special human element. A wrong output would be caught in ordinary review and cost little, and no relationship or judgment is at stake that a person must supply. Drafting an internal FAQ from already-approved policy documents is a typical example: the source is trusted, mistakes are easy to catch, and nothing about the task requires a human to own it. You delegate it and apply the normal review you would give any work product.
An inappropriate use case sits at the far end. Here the consequence, the irreversibility, or the human-element requirement means AI should not perform the task, and no light checkpoint fixes that. Generating a final medical or legal determination is the standard case: the outcome is high stakes, hard or impossible to reverse once acted on, and accountability for it cannot transfer to a model. The right response is not to bolt on a review step but to name the human role that must own the determination outright.
Appropriate with human review: the middle that carries the work
The middle classification is where most real use cases live, and it is the one that does the most work. "Appropriate with human review" means the task genuinely benefits from AI assistance, but the stakes or the accountability involved require a human checkpoint before the output is used. Claude does the bulk of the drafting or analysis; a person verifies the specific risk that matters before anything happens.
The critical discipline is that this classification is not complete when you write the label. It is complete when you define the gate. A financial summary that feeds a decision is consequential, but if a named reviewer signs off before it is used, the accountability and reversibility are effectively restored and the use case becomes workable. The value of the classification comes entirely from that gate being real and specified, which is why the human review gate has its own detailed treatment. Naming "appropriate with review" without saying who checks what and when is only half a decision.
Why the rationale is part of the answer
Whichever of the three you land on, the classification is paired with a written rationale, and that pairing is not bureaucratic decoration. A rationale names which criterion drove the decision, and that is what makes the call defensible later.
"It feels risky" does not travel. It cannot be reviewed, cannot be reused when a similar case appears, and gives a compliance reviewer nothing to check. "Irreversible consequence plus non-transferable accountability, so a human must own the determination" does travel: it states the reason, it can be tested against the facts, and it survives being handed to someone who was not in the room. Documenting the reasoning is what turns an individual judgment into an organizational decision.
What the exam trips candidates on
The first trap is picking a classification on gut feeling instead of naming which criterion drove it. A verdict with no stated reason is exactly what the classification framework exists to replace. The exam rewards the answer that ties the outcome to a specific criterion, not the one that simply asserts a label.
The second trap is treating "appropriate with review" as effectively the same as "fully appropriate," on the logic that AI still does most of the work either way. The two are not equivalent. Fully appropriate needs no special checkpoint; appropriate-with-review is only safe because a specific human gate stands between the output and its use. Collapsing the two - delegating an appropriate-with-review task and applying only ordinary review - removes the very control that made the classification defensible.
Worked example
A team lists four proposed uses of Claude: (1) draft an internal FAQ from approved policy docs, (2) summarize candidate resumes into a shortlist for a recruiter, (3) generate a final, unreviewed benefits-eligibility determination, and (4) draft customer-facing responses to billing complaints. Classify each and give the reason.
(1) Draft internal FAQ from approved policy docs. The source is already cleared, mistakes are caught in ordinary review, and no one's outcome hinges on a single line. Reversible, low consequence, no human element. Classification: fully appropriate, delegate with normal review.
(2) Summarize resumes into a shortlist. AI can genuinely help, but the output affects who gets considered, and accountability for fair screening cannot pass to the model. Classification: appropriate with human review - and the gate must be specified, for example a person reviewing the shortlist and the exclusions for adverse-impact patterns before anyone is contacted.
(3) Final unreviewed eligibility determination. Irreversible in effect once acted on, high consequence for the person, and accountability cannot transfer to Claude. No light checkpoint repairs that. Classification: inappropriate - a human must own the determination.
(4) Draft customer billing responses. Useful for AI drafting, but the messages go to real customers about their money, so a person should verify accuracy and tone before sending. Classification: appropriate with human review, gate defined at the point of sending.
Notice that each verdict is tied to the criterion that decided it, not asserted. That is what makes the portfolio auditable.
Common misreadings to avoid
Misconception
You can classify a use case correctly just by choosing the label that feels right.
What's actually true
Misconception
Appropriate with review is basically the same as fully appropriate, since AI does most of the work either way.
What's actually true
How this shows up on the exam
Domain 6 items present a proposed use case and ask you to classify it, usually with distractors that either skip the rationale or blur the middle category into the top one. The dependable approach is to run the four criteria, pick the classification the deciding criterion points to, and - for the middle category - insist the answer specifies a gate rather than just naming the label.
From here the path forks two ways. When the four criteria disagree, you find the load-bearing criterion that actually decides the classification, and when you land on "appropriate with review," you specify a real human review gate in who/what/when form. Both build directly on being able to name the three outcomes and justify the one you chose.
A recruiter wants Claude to summarize incoming resumes into a shortlist. The team labels it 'appropriate with human review' and moves on. What is missing before this use case is ready to run?
People also ask
What are the three classifications for an AI use case?
What does appropriate with human review mean?
Is appropriate-with-review the same as fully appropriate?
Watch and learn
Official Anthropic Academy lessons first, then hand-picked walkthroughs. Videos load only when you press play.
No videos curated for this concept yet
We are still curating the best official and community videos for this topic.
Official prep for this domain
Anthropic's own free prep module for this part of the syllabus, on the official prep course. Free with an Anthropic Academy sign-in.
References & primary sources
Master this concept with Archie
Practice it inside an adaptive study session. Archie, your Socratic AI tutor, tracks your mastery with Bayesian Knowledge Tracing and schedules the perfect next review.