Governance, Risk, and Responsible Use·Task 6.1·Bloom: evaluate·Difficulty 4/5·10 min read·Updated 2026-07-14

Classifying Ambiguous Production Use Cases for the CCAO-F Exam

Identify appropriate and inappropriate use cases

SUBy Solomon UdohReviewed by Solomon UdohAI-assisted · human-reviewed
In short
Classifying ambiguous production use cases means applying the full screening framework - the four criteria, the load-bearing factor, and gate design - to realistic tasks such as automated eligibility determinations or unreviewed candidate screening. An irreversible, high-stakes determination with no transferable accountability is inappropriate for unreviewed AI; a workflow that filters people out without reviewing the exclusions can hide systemic harm. The exam rewards articulating the specific irreversibility and accountability reasoning, not just a verdict.

Putting the whole framework to work

This knowledge point is where the pieces of Task 6.1 come together. The CCAO-F exam sets an evaluate-level bar here: given a realistic, ambiguous production use case, apply the four criteria, find the load-bearing criterion, and decide whether a defined human review gate can make it workable. The cases are chosen to resist a snap answer, and the credited response is a chain of reasoning, not a label.

What distinguishes an evaluate-level answer from a lower one is that the verdict is inseparable from its justification. "Inappropriate" on its own earns little; "inappropriate because the determination is irreversible in effect and accountability cannot transfer to a model" earns the credit. The skill is producing reasoning that would survive a compliance reviewer reading it cold.

Classifying ambiguous production use cases
Applying the full screening framework to realistic, non-obvious tasks: run the four criteria, name the load-bearing factor, and judge whether a defined gate offsets it. Verdicts turn on specifics - an irreversible, high-stakes determination with no transferable accountability is inappropriate; a task that filters people out needs review of the exclusions, not only the surfaced results. The reasoning, not the verdict alone, is the deliverable.

The irreversible, unaccountable determination

The clearest inappropriate pattern is a determination that is high-stakes, effectively irreversible once acted on, and carries accountability that cannot pass to a model. Automated final eligibility for benefits is the archetype.

Walk the criteria. Consequence is high: the decision materially affects a person's access to something they need. Reversibility is poor: once a determination is issued and acted upon, the harm has landed, and an after-the-fact correction does not undo the interim damage. Accountability cannot transfer: a human institution remains answerable for eligibility decisions, and "the model decided" is not a defense anyone can stand behind. When you run the flip test, no single gate rescues an unreviewed determination, because the thing at stake - a final, binding decision about a person - is exactly what a human must own. The load-bearing factors are irreversibility and non-transferable accountability, and neither is repaired by using a more capable model or moving faster. The verdict is inappropriate for unreviewed AI, and the articulation of why is the point.

When low stakes still needs a human

The mirror case is the task that looks fully appropriate on consequence and reversibility yet still needs a person, because it carries a relationship or empathy element. This is where candidates over-approve.

A short, reversible, low-cost message can still be work a human should own - a note to a grieving client, a sensitive acknowledgment, anything whose value depends on a person genuinely standing behind it. Here the load-bearing criterion is the human element, not consequence or reversibility, and the correct classification is appropriate-with-review (or human-owned) even though a two-criterion glance would wave it through. The evaluate-level move is noticing that the deciding factor is not the one the numbers point to.

The silent-exclusion workflow

The subtlest pattern is a workflow that filters people out and never reviews who was filtered. An AI screen produces a shortlist; a human looks only at the shortlist; the excluded group is never seen by anyone. No individual decision looks wrong, yet the workflow can hide a systematic disadvantage to some groups.

The failure is that the review gate, if there is one, points at the wrong thing. Reviewing the quality of the surfaced shortlist does nothing to catch bias in what was dropped. A defensible design applies a human review gate to the exclusion step itself - someone checks the filtered-out group for adverse-impact patterns before the shortlist is treated as final - not only to the results that made it through. This is the same reasoning developed in diagnosing unreviewed-exclusion bias; at the use-case-screening stage it shows up as a gate placed on the exclusions.

Evaluating an ambiguous use case end to end
Loading diagram...
A defensible verdict follows the chain: criteria, load-bearing factor, whether a gate offsets it, and whether that gate covers the people filtered out.

What the exam trips candidates on

The first trap is choosing "fully appropriate" for a use case just because it saves time or uses a more capable model. Efficiency and model quality are not criteria on the screen. A time-saving distractor is a distractor precisely because it answers a question the framework never asks; the credited answer ignores the speed benefit and reasons from irreversibility and accountability.

The second trap is overlooking that excluded or filtered-out cases also need a human check, not only the cases that were surfaced. On a screening scenario, the tempting answer reviews the shortlist. The better answer notices the silent exclusions and puts the gate where the hidden harm is. Reviewing only what surfaced is a partial control that misses the actual risk.

Worked example

A public-benefits office proposes two changes. Change 1: Claude issues final eligibility determinations automatically, with no caseworker review, to clear a backlog. Change 2: Claude screens applications and forwards a 'qualified' subset to caseworkers, who review only that subset. Evaluate both, giving the reasoning the exam expects.

Change 1. Run the criteria. Consequence: high - eligibility governs access to needed support. Reversibility: effectively irreversible once a determination is acted upon; a later reversal does not undo the interim harm. Accountability: the office remains answerable, and that cannot pass to a model. The load-bearing factors are irreversibility and non-transferable accountability, and no gate short of full human ownership repairs an unreviewed final determination. Verdict: inappropriate for unreviewed AI. Clearing a backlog is a time benefit the screen does not weigh, so it does not change the call. A caseworker must own each determination.

Change 2. Screening to assist caseworkers can be legitimate, but the proposed design reviews only the "qualified" subset. That leaves the excluded applicants unseen by any human, so a systematic disadvantage in who gets filtered out could go undetected while every individual forwarded case looks fine. The load-bearing issue is the un-reviewed exclusion. The fix is a defined gate on the exclusions: a caseworker reviews the filtered-out group for adverse-impact patterns before the screen's output is relied on. With that gate, the use case can be appropriate-with-review; without it, reviewing only the shortlist is a partial control that misses the real risk.

In both, the verdict travels only because it names the specific irreversibility, accountability, and exclusion reasoning - not because it asserts a label.

Common misreadings to avoid

Misconception

A use case is fully appropriate if it clearly saves time or uses a more capable model.

What's actually true

Speed and model capability are not screening criteria. A time-saving, high-capability use case can still be inappropriate when it is irreversible and accountability cannot transfer to a model. Reason from the criteria, not the efficiency.

Misconception

In a screening workflow, reviewing the quality of the surfaced shortlist is enough.

What's actually true

Reviewing only what was surfaced misses bias in what was silently filtered out. A defensible gate reviews the exclusions - the people dropped - not just the shortlist, because that is where hidden systemic harm lives.

How this shows up on the exam

Domain 6 evaluate-level questions hand you a realistic use case - eligibility automation, resume screening, an automated exclusion - and ask for the classification. The winning answer states the verdict and the specific irreversibility/accountability/exclusion reasoning behind it, ignores time-saving distractors, and, on screening cases, insists the gate cover the filtered-out group.

This knowledge point is the capstone of Task 6.1, drawing together the load-bearing criterion and the human review gate, and it hands off to the ethics domain, where the same silent-exclusion pattern reappears as unreviewed-exclusion bias. Master the reasoning here and the hardest governance items become straightforward.

Check your understanding

A hiring team uses Claude to screen resumes into a shortlist and has a recruiter review only the shortlisted candidates before interviews. What is the most defensible assessment of this use case?

People also ask

How do you classify a borderline AI use case?
Run the four criteria, identify the load-bearing one, and decide whether a defined gate can offset it. Irreversible, high-stakes factors with no transferable accountability make it inappropriate; a real gate that restores accountability makes it appropriate-with-review.
Why is unreviewed eligibility determination inappropriate for AI?
The determination is high-stakes and effectively irreversible once acted on, and accountability for it cannot transfer to a model. With no human owning the result, unreviewed AI output is not defensible.
Do filtered-out cases need human review too?
Yes. A workflow that reviews only what was surfaced can hide harm in what was silently excluded. The gate must cover the exclusions, not just the shortlist.

Watch and learn

Official Anthropic Academy lessons first, then hand-picked walkthroughs. Videos load only when you press play.

No videos curated for this concept yet

We are still curating the best official and community videos for this topic.

Official prep for this domain

Anthropic's own free prep module for this part of the syllabus, on the official prep course. Free with an Anthropic Academy sign-in.

References & primary sources

Adaptive study

Master this concept with Archie

Practice it inside an adaptive study session. Archie, your Socratic AI tutor, tracks your mastery with Bayesian Knowledge Tracing and schedules the perfect next review.

Start studying