- In short
- The four risk thresholds are the questions that decide whether an output requires human review: stakes (what an error costs), reversibility (whether the action can be undone), audience (who sees it), and regulatory exposure (whether a law, rule, or contract governs it). Each is asked independently, and any one crossing its threshold can require review regardless of how confident the output looks.
Deciding escalation by policy, not by mood
Some outputs must never go out as a Claude draft alone, no matter how good they look. This is the operational heart of Diligence, one of the two competencies from Anthropic’s AI Fluency Framework that anchor output evaluation: where Discernment is the skill of critically evaluating output against requirements, sources, and standards, Diligence is deciding when verification is required and taking responsibility for the result. The Claude Certified Associate - Foundations (CCAO-F) exam frames that Diligence as knowing which outputs must never ship unreviewed in advance, so that the decision to escalate is made by policy rather than in the moment after something has already gone wrong. The four risk thresholds are that policy: four questions you ask of every output to decide whether it needs a human before it ships.
The reason there are four, and not one, is that risk arrives from different directions. An output can be low-stakes but irreversible, or internal but regulated. A single question would miss whichever dimension it does not cover. Asking all four catches the risk wherever it comes from.
- The four risk thresholds
- Four independent questions that decide whether an output requires human review before release. Stakes: what is the cost if this is wrong? Reversibility: can the action be undone? Audience: who sees it? Regulatory exposure: does a law, rule, or contract govern it? Any single threshold crossing can require review, regardless of how confident or polished the output appears.
Stakes
Stakes asks the cost of being wrong. When an error would be expensive, damaging, or hard to recover from, the output demands human review regardless of how confident it appears. This is the threshold most people already feel, but the exam's emphasis is on the "regardless of how confident" clause: a high-stakes output that reads flawlessly still crosses the stakes threshold, because the appearance of correctness is not correctness. Stakes are judged by consequence, not by how the output looks.
Reversibility
Reversibility asks whether the action can be undone. A draft you can revise before it matters clears a low bar; a step that cannot be taken back, a sent client deliverable, a filed report, a public statement, clears a much higher one. This is a distinct question from stakes, because an action can be moderate in cost yet irreversible, and irreversibility removes the safety net that would otherwise let a mistake be caught and corrected downstream. When there is no undo, the review has to happen before the action, not after.
Audience
Audience asks who sees the output. An internal working draft is the low end; external, executive, and regulatory audiences raise the requirement sharply. The wider or more consequential the readership, the more a mistake costs in trust and the less room there is to quietly fix it. Importantly, an internal audience lowers this threshold but does not switch off the others: internal work can still be high-stakes, irreversible, or regulated, and any of those can require review on its own.
Regulatory exposure
Regulatory exposure asks whether a law, rule, or contract governs the content. Regulated material carries obligations that AI assistance does not remove; using Claude to draft something does not transfer or dissolve the compliance duties attached to it. When a rule governs the output, the regulatory threshold is crossed by that fact alone, independent of stakes or audience, because the obligation exists regardless of how the work was produced.
What the CCAO-F exam trips candidates on
The first trap is judging only stakes and ignoring that the same output might also be irreversible once sent. A scenario will foreground the cost of an error and hope you stop there, missing that the action cannot be undone. The credited answer runs all four thresholds, catching reversibility as its own trigger.
The second trap is assuming an output aimed at an internal audience never needs escalation, regardless of stakes. Internal lowers the audience threshold but does not clear the others. The exam rewards recognising that an internal but high-stakes, irreversible, or regulated output still crosses a threshold and still needs review.
Worked example
Claude drafts an internal financial model that the finance team will use, unchanged, to set next quarter's headcount budget. A colleague says 'it's just internal, so no review needed.' Which thresholds actually apply?
The "just internal" reasoning stops at the audience threshold and treats a low reading there as a blanket exemption, which is precisely the trap. Audience is only one of four independent questions, so run the others.
Stakes: an error here misprices next quarter's headcount budget, which is expensive and disruptive to unwind, so the stakes threshold is crossed regardless of how polished the model looks. Reversibility: the model will be used unchanged to set the budget, so once decisions are made off it, they are costly to reverse, the reversibility threshold is crossed too. Regulatory exposure: depending on the organisation, financial planning of this kind may touch internal controls or audit obligations, which would cross the regulatory threshold on its own. Only the audience threshold reads low, because the readership is internal.
Three of four thresholds are crossed, so the output needs human review despite being internal. The lesson is that a low audience reading does not switch off stakes, reversibility, or regulatory exposure; any single one of those can require escalation by itself. Deciding "no review" from the audience question alone is the mistake the four-threshold policy exists to prevent.
Common misreadings to avoid
Misconception
If the stakes look manageable, the output does not need escalation.
What's actually true
Misconception
Internal-only outputs never need human review.
What's actually true
How this shows up on the exam
Domain 2 questions on this knowledge point describe an output and ask whether it needs human review. The dependable approach is to run all four thresholds independently, stakes, reversibility, audience, regulatory exposure, and escalate if any single one is crossed, treating a confident or internal-looking output as no exemption.
These thresholds are the general form of the stakes calibration idea, and they underpin the fixed do-not-ship-without-review categories, the iteration versus escalation signal, and their combined application in applying the escalation thresholds to a scenario. They also connect to the accountability ownership principle, which is why crossing any threshold matters.
Claude drafts an internal financial model the finance team will use unchanged to set next quarter's budget. Does it need human review?
People also ask
When does AI output need human review?
What are the four risk thresholds for escalation?
Does internal-only output ever need review?
Watch and learn
Official Anthropic Academy lessons first, then hand-picked walkthroughs. Videos load only when you press play.
No videos curated for this concept yet
We are still curating the best official and community videos for this topic.
Official prep for this domain
Anthropic's own free prep module for this part of the syllabus, on the official prep course. Free with an Anthropic Academy sign-in.
References & primary sources
Master this concept with Archie
Practice it inside an adaptive study session. Archie, your Socratic AI tutor, tracks your mastery with Bayesian Knowledge Tracing and schedules the perfect next review.