- In short
- The two-reader test asks whether two different people would interpret a standing instruction identically. An instruction passes if two independent readers would apply it the same way; if it could reasonably be read two different ways, it needs concrete criteria added. The test is a practical check to run before adding an instruction to a Project, and it catches subtle vagueness that survives a first read, since a word like "professional" sounds specific but is not.
A concrete test for an abstract quality
Knowing that instructions should be precise rather than vague raises a practical question: how do you tell, for a specific instruction you have just written, whether it is precise enough? The CCAO-F exam treats the answer as an apply-level skill, and the answer is a simple, portable test.
The test is this: would two different people, reading the instruction independently, interpret it the same way and apply it the same way? If yes, the instruction is precise enough to be reliable. If two reasonable readers could apply it differently, it is not, and it needs concrete criteria added. The two-reader test turns the abstract idea of "precise" into something you can actually run.
- The two-reader test
- A clarity check for a standing instruction: would two different, independent readers interpret and apply it the same way? If they would, the instruction is precise enough. If it could reasonably be read two different ways, it needs concrete criteria. The test is run before adding an instruction to a Project and repeated as instructions are written or revised, and it catches subtle vagueness that a single read misses.
Why a single read is not enough
The reason the test names two readers is that a single reader is a poor judge of clarity. When you read your own instruction, you already know what you meant, so it reads as clear to you even when the words do not pin the meaning down. The word "professional" is the standard example: it sounds specific, and on a first read it feels like a real directive, but two people asked to apply it will produce different things.
The two-reader framing exposes exactly that gap. By imagining a second, independent reader who does not share your intent, you test whether the words themselves carry the meaning or whether they were leaning on knowledge only you had. Subtle vagueness that survives your own first read, which is the dangerous kind because it feels fine, is what the second reader catches.
Running the test and fixing failures
In practice, running the test means pausing on an instruction and asking honestly whether it admits more than one reasonable application. "Be concise" fails: one reader trims to a page, another to a paragraph. "Keep the executive summary to five bullet points or fewer" passes: both readers do the same thing. When an instruction fails, the fix is not to argue it is clear but to add the concrete criteria that collapse the interpretations into one, exactly the measurable rules that make an instruction precise.
Crucially, this is not a one-time exercise. The test is a habit to run every time you write or revise a standing instruction, before you add it to the Project. Building it into the authoring step is what keeps vague instructions from getting in, and it is far cheaper than diagnosing inconsistent output later and tracing it back to wording.
What the CCAO-F exam trips candidates on
The exam sets two traps. The first is assuming an instruction is precise simply because it uses formal-sounding language rather than because it specifies checkable criteria. Formal-sounding words like "professional" or "rigorous" can still admit multiple interpretations; the two-reader test cuts through the sound to the substance. The credited answer judges by whether two readers would agree, not by how authoritative the wording feels.
The second is skipping a clarity check on instructions that sound concrete but leave room for differing interpretations. An instruction that reads fine on a first pass may still fail the two-reader test, and not running the check is how subtle vagueness slips through. The credited answer runs the test as a routine authoring step rather than assuming a confident-sounding instruction is clear. Both traps reward substituting the two-reader check for a first-impression judgement.
Worked example
A team is about to add the standing instruction 'Keep all client communications professional and appropriately detailed.' It reads well to the author. Apply the two-reader test and decide whether it is ready, then fix it if needed.
Run the test and the instruction fails, despite reading well to its author.
Imagine two independent readers applying "professional and appropriately detailed." One might write terse, formal three-line replies; another might write warm, thorough half-page responses. Both could honestly claim to be following the instruction, which means it admits two reasonable interpretations and is not precise enough. The reason it reads fine to the author is the classic trap: "professional" and "appropriately detailed" sound specific but carry no checkable criteria, and the author already knows what they meant, so their single read misses the gap.
The fix is to add concrete criteria that collapse the interpretations into one, for example: "Use a formal register. Open with a one-line acknowledgment of the client's request. Address every question the client raised. Keep the reply under 200 words unless the client asked for detail." Now two readers would produce closely matching output, so the instruction passes the two-reader test and is ready to add. The point is that the check, not the author's first impression, is what decided it.
Common misreadings to avoid
Misconception
If an instruction uses formal, authoritative-sounding language, it is precise enough.
What's actually true
Misconception
Running a clarity check is a one-time step you can skip once you have the hang of writing instructions.
What's actually true
How this shows up on the exam
Domain 5 questions present an instruction that sounds fine and ask whether it is ready. Apply the two-reader test: if two reasonable readers could apply it differently, it needs concrete criteria, regardless of how authoritative it sounds. Treat the check as a routine authoring gate.
This test operationalises vague versus precise instructions and is the front-line defence against silent instruction failure, where unchecked vagueness degrades output without warning. It supports anticipating use cases, since anticipated instructions still have to pass the clarity check.
A team is about to add the standing instruction 'Keep all client communications professional and appropriately detailed,' which reads well to its author. What does the two-reader test indicate?
People also ask
What is the two-reader test for instructions?
Why does an instruction that sounds specific still fail?
When should I run a clarity check on an instruction?
Watch and learn
Official Anthropic Academy lessons first, then hand-picked walkthroughs. Videos load only when you press play.
No videos curated for this concept yet
We are still curating the best official and community videos for this topic.
Official prep for this domain
Anthropic's own free prep module for this part of the syllabus, on the official prep course. Free with an Anthropic Academy sign-in.
References & primary sources
Master this concept with Archie
Practice it inside an adaptive study session. Archie, your Socratic AI tutor, tracks your mastery with Bayesian Knowledge Tracing and schedules the perfect next review.