- In short
- Confirmation bias in framing occurs when a prompt that signals the asker preferred conclusion pulls Claude response toward agreement. It shows up as output that agrees a little too readily with a genuinely contestable question. Rephrasing the prompt neutrally tests whether the original framing was steering the answer. This risk is separate from hallucination: the individual facts can all be true while the framing is still skewed.
A bias you introduce yourself
Most failure patterns in this domain are things Claude does to an output. Confirmation bias in framing is different: it starts with you. The Claude Certified Associate - Foundations (CCAO-F) exam treats this as its own pattern because the cause lives in how the prompt is worded, not in a fabricated fact. When your prompt telegraphs the answer you are hoping for, the response can lean toward that answer, and because every stated fact might still be true, the skew is easy to miss.
This is worth separating from hallucination precisely because the usual tells do not fire. There is no uncited statistic, no invented citation, no internal contradiction. The output can be factually clean and still be subtly steered, agreeing with a premise it should have tested. Recognising that the problem is the framing, not the facts, is the whole skill.
- Confirmation bias in framing
- When a prompt implies a preferred answer, Claude's response can lean toward agreement with it. The signature is output that agrees a little too readily with a genuinely contestable question, without the pushback a balanced answer would include. Rephrasing the prompt neutrally tests whether the framing was steering the result. The bias is separate from hallucination, because the stated facts can all be true while the framing remains skewed.
How a leading prompt steers the answer
A prompt carries more than a question; it carries a stance. "Confirm that switching vendors will save us money" and "Should we switch vendors?" ask about the same decision, but the first has already announced the conclusion it wants. Faced with a premise stated as a foregone conclusion, a response can accommodate it, marshalling support for the implied answer rather than genuinely weighing both sides. The more a question should be open, the more damage a leading frame does, because it converts a real deliberation into a rationalisation.
The signature: agreeing too readily
You catch framing bias by watching how the response handles contestability. On a question that has legitimate arguments on both sides, a balanced answer surfaces the tension; a framing-biased answer glides past it, agreeing with your stated hypothesis and offering little resistance. The tell is not any single false sentence. It is the absence of the pushback a genuinely open question deserves, the sense that the output agreed with you a little too easily on something that should have been argued.
The neutral-rephrase test
The practical test is to re-ask the question in neutral language, stripped of the implied answer, and compare. If the neutral version yields a materially different or more balanced response, the original framing was doing the steering. This is cheap, it is fast, and it isolates the variable: same underlying question, different framing, and any change in the answer is attributable to the frame. Where stakes justify it, comparing responses across framings is one instance of comparing multiple drafts before you commit.
What the CCAO-F exam trips candidates on
The first trap is treating agreement with the asker's stated hypothesis as confirmation that the hypothesis is correct. That reasoning is circular: you wrote a leading prompt, the response agreed, and the agreement feels like validation when it may just be the frame reflected back. The credited answer treats too-ready agreement on a contestable question as a reason to run the neutral-rephrase test, not as evidence.
The second trap is assuming bias can only come from training data and never from how the current prompt is worded. This misses the entire point of the pattern. The framing bias here is introduced live, in the prompt you just wrote, and it is fixable by you in the next prompt. The exam rewards recognising the prompt itself as a source of bias, distinct from anything baked into the model.
Worked example
A manager prompts: 'Explain why our new onboarding process is clearly improving retention.' Claude returns a fluent, confident answer listing reasons retention is improving, with no obvious factual errors. The manager takes this as confirmation. What is the risk, and how should it be tested?
The risk is confirmation bias in framing, and it is easy to miss because the usual hallucination tells are absent. There is no uncited statistic to flag and no internal contradiction to reconcile; the individual reasons offered may each be plausible and even true. The problem is upstream of the facts. The prompt did not ask whether the onboarding process is improving retention; it asserted that it "clearly" is and asked Claude to explain why. A genuinely open question, does the new process improve retention, has been reframed as a settled one, and the response accommodated that frame by supplying support rather than testing the claim.
Taking the answer as confirmation is the circular trap: the manager supplied the conclusion, the response echoed it, and the echo feels like independent validation when it is really the framing reflected back. Nothing here establishes that retention actually improved or that the onboarding change caused it.
The test is a neutral rephrase: ask "What effect, if any, has the new onboarding process had on retention, and what could explain it?" and compare. If the neutral version surfaces confounds, mixed evidence, or the possibility of no effect, the original framing was steering the result, and the balanced answer is the one to act on. Note that this is not a hallucination fix; it is a framing fix, applied in the prompt, and it sits alongside broader prompt-level grounding techniques for getting more trustworthy answers.
Common misreadings to avoid
Misconception
If Claude agrees with my hypothesis, that agreement supports the hypothesis.
What's actually true
Misconception
Bias in AI output comes only from training data, not from my prompt.
What's actually true
How this shows up on the exam
Domain 2 questions on this knowledge point present a leading prompt and an agreeable, factually clean answer, then ask for the risk or the fix. The reliable reading is that the framing, not a fabricated fact, is the problem, that too-ready agreement on a contestable question is the signature, and that a neutral rephrase is the test.
Confirmation bias in framing is one of the failure patterns catalogued in the hallucination pattern taxonomy, distinguished from it by living in the prompt rather than the facts. The neutral-rephrase test connects to comparing multiple drafts, and reducing framing pressure at the source is part of prompt-level grounding techniques.
A prompt asks Claude to 'explain why our new process is clearly working,' and the response agrees fluently with no factual errors. What is the most likely risk and the best test?
People also ask
Can how you word a prompt bias the answer?
What is confirmation bias in AI output?
How do you test whether a prompt is leading?
Watch and learn
Official Anthropic Academy lessons first, then hand-picked walkthroughs. Videos load only when you press play.
No videos curated for this concept yet
We are still curating the best official and community videos for this topic.
Official prep for this domain
Anthropic's own free prep module for this part of the syllabus, on the official prep course. Free with an Anthropic Academy sign-in.
References & primary sources
Master this concept with Archie
Practice it inside an adaptive study session. Archie, your Socratic AI tutor, tracks your mastery with Bayesian Knowledge Tracing and schedules the perfect next review.