- In short
- Context degradation shows up through observable symptoms: Claude no longer following an instruction it was correctly following earlier in the same session, responses that address only the most recent message and ignore earlier decisions, and accuracy drops that track with content from early in a long session. These symptoms point to a context-management problem -- earlier content being compressed as the window fills -- not to a change in the model's capability mid-session.
Reading the symptoms
Knowing that the context window is a finite budget is only useful if you can spot when the budget has filled and started to bite. The CCAO-F exam tests this at the understand level: not the mechanism itself, but the observable symptoms that tell you a session's context has degraded. Getting the diagnosis right matters because the wrong diagnosis leads to the wrong fix.
There are three characteristic signs, and they share a source: earlier content being compressed as the window fills. The critical interpretive move is that all three point to a context-management problem, not to the model getting worse partway through. A model does not lose capability mid-session; a session loses early context. Reading the symptoms correctly is what routes you to the restart, summarise, or persist responses.
- Signs of context degradation
- Observable symptoms that a session's context has degraded: Claude no longer following an instruction it followed correctly earlier in the same session; responses addressing only the most recent message while ignoring earlier decisions; and accuracy drops that track with early-session content. All three indicate a context-management problem -- earlier content compressed as the window filled -- not a mid-session change in model capability.
The three signals
The first signal is dropped instructions: Claude no longer follows an instruction it was correctly following earlier in the same session. The key qualifier is "earlier in the same session." Because Claude was following the instruction before and has stopped, the instruction itself was understood; what changed is that its detail has been compressed out of the active context. This is a strong, specific tell.
The second signal is recency-anchoring: responses address only the most recent message and ignore earlier decisions or context. Instead of integrating the whole thread, the reply behaves as if only the latest exchange exists. The third is a pattern in accuracy: accuracy drops in ways consistent with missing early-session content. When errors line up with material from the start of a long session -- exactly the content most compressed -- the drop is a degradation pattern, not random noise. Together these three form a recognisable profile.
Symptom, not model failure
The interpretation the exam most wants is that these symptoms are a context-management problem, not a change in model capability. It is natural, when Claude stops following an instruction, to conclude that the model has "forgotten how to follow instructions" or become unreliable. That reading is wrong and leads nowhere useful, because the model's capability has not changed within the session.
What has changed is the state of the context window. The early instruction, the earlier decisions, the start-of-session content -- all have been compressed as the budget filled, so the model is acting on a thinner version of the context than it had at the start. Attributing the drop to model unreliability rather than to the context budget misdirects the fix toward the model when it belongs on the context. Recognising the symptoms as context state, not model state, is the whole diagnostic skill, and it is what the long-session diagnosis applies in depth.
What the CCAO-F exam trips candidates on
Two misdiagnoses are tested. The first is diagnosing dropped early instructions as the model forgetting how to follow instructions, rather than as context compression. The credited reading is that the instruction was followed earlier and the model still can follow instructions; the early instruction's detail has simply been compressed out of active context.
The second is attributing degraded accuracy late in a session to model unreliability instead of to the context budget. When the accuracy loss tracks with early-session content in a long session, the cause is the context budget, not a flaky model. Both traps share the error of blaming the model for what is a context-state symptom.
Worked example
Ninety minutes into a long working session, Claude begins ignoring a formatting rule it applied perfectly at the start, its replies stop referencing decisions made earlier in the thread, and it gets a fact wrong that was stated clearly in the first few messages. A teammate says 'the model has become unreliable today.' What is the correct diagnosis?
The teammate is blaming the model, which is the misdiagnosis the exam targets. Look at the three symptoms together, and they form the classic context-degradation profile, all pointing at the context budget rather than the model.
The formatting rule was applied perfectly at the start and is now ignored -- dropped instruction. Because it was followed earlier in the same session, the model plainly understood it and can follow instructions; what changed is that the rule's detail has been compressed as the window filled over ninety minutes. The replies no longer referencing earlier decisions is recency-anchoring: the response is behaving as if only the latest messages exist. And the fact stated clearly in the first few messages now coming out wrong is accuracy loss tracking early-session content -- precisely the material most compressed.
Three symptoms, one cause: a filled context window compressing the earliest content. Nothing about the model's capability changed within the session; the session's context state did. "The model has become unreliable today" points the fix at the wrong place. The correct diagnosis is a context-management problem, which then routes to the right response -- restart, summarise, or persist -- rather than to distrusting the model.
Common misreadings to avoid
Misconception
When Claude stops following an early instruction, the model has forgotten how to follow instructions.
What's actually true
Misconception
Accuracy dropping late in a session means the model has become unreliable.
What's actually true
How this shows up on the exam
Questions describe late-session symptoms -- dropped instructions, recency-anchored replies, patterned accuracy loss -- and ask for the diagnosis, with a "the model became unreliable" distractor. Read the symptoms as context state: the budget filled and early content was compressed. The reliable answer names a context-management problem and never blames the model's capability.
This knowledge point builds on the finite context budget and unlocks restart, summarise, or persist and the long-session diagnosis. It pairs with choosing between restart and summarise, which acts on the diagnosis.
Late in a long session, Claude ignores a formatting rule it followed at the start, references only the latest message, and misstates a fact given clearly at the beginning. Which diagnosis is correct?
People also ask
How do I know a Claude conversation has degraded?
Why does Claude stop following earlier instructions?
Is degraded accuracy a model problem or a context problem?
Watch and learn
Official Anthropic Academy lessons first, then hand-picked walkthroughs. Videos load only when you press play.
No videos curated for this concept yet
We are still curating the best official and community videos for this topic.
Official prep for this domain
Anthropic's own free prep module for this part of the syllabus, on the official prep course. Free with an Anthropic Academy sign-in.
References & primary sources
Master this concept with Archie
Practice it inside an adaptive study session. Archie, your Socratic AI tutor, tracks your mastery with Bayesian Knowledge Tracing and schedules the perfect next review.