- In short
- Diagnosing a flawed A/B comparison means recognising the compound failure of an undersized sample, an uncontrolled input distribution, and post-hoc metric selection. A small sample such as 50 sessions per arm can show a gap within the noise floor that disappears at scale; an uncontrolled distribution where one arm sees easier inputs can produce an apparent win that is really a sampling artifact; and choosing the primary metric after seeing which one moved converts the experiment into outcome-shopping. Any one of the three is sufficient to invalidate a result.
Three ways a convincing result can be worthless
The most dangerous experiment result is the one that looks like a clear win and is not. This evaluate-level knowledge point is about recognizing the compound failure that produces exactly that: an undersized sample, an uncontrolled input distribution between the arms, and a metric chosen after the fact. Each is independently fatal, and they often appear together, reinforcing one another into a result that feels solid and means nothing. The skill is diagnosing all three and understanding that fixing only one still leaves the result invalid.
- Underpowered, outcome-shopped experiment
- A flawed A/B comparison undermined by up to three independent failures: an undersized sample too small to distinguish a real effect from noise; an uncontrolled input distribution where one arm receives easier or harder inputs, making any gap a sampling artifact; and post-hoc metric selection, choosing the primary metric after seeing which one moved favourably. Any one alone invalidates the result; together they manufacture a convincing but meaningless win.
Failure one: the sample is too small
The first failure is a sample too small for the effect being claimed. A comparison of 50 sessions per arm on a noisy LLM metric cannot distinguish a real effect from random variation, because a several-point gap is well inside the noise floor at that size. The gap looks like signal, but it is the kind of fluctuation that appears and vanishes as the sample grows. Its signature is that the apparent effect disappears when the experiment is rerun at an adequate scale, which is exactly what happens when the small win settles back to nothing weeks later.
Failure two: the arms saw different inputs
The second failure is an uncontrolled input distribution. If assignment is not properly randomized and held consistent, one arm can happen to receive an easier or harder mix of inputs than the other. When that happens, the metric gap between the arms reflects which inputs each arm saw, not the change under test. The treatment arm might look better simply because it drew fewer edge cases, a sampling artifact wearing the costume of a real improvement. This failure is independent of sample size: even a large experiment produces a biased result if the arms are not comparable.
Failure three: the metric was chosen afterward
The third failure is post-hoc metric selection, outcome-shopping. If the primary metric is picked after seeing which metric moved favourably, the experiment stops being a test. With several metrics on the dashboard and noisy outputs, one is likely to have drifted in a flattering direction by chance, and selecting it retroactively guarantees a "win" that carries no evidential weight. The tell is that had a different metric moved, the team would have reported that one instead, meaning no possible result could have counted as a failure.
Any one is fatal, and they compound
The crucial evaluative point is that these three are independent and each is individually sufficient to invalidate the result. A perfectly sized, perfectly balanced experiment is still worthless if the metric was chosen after the fact. A perfectly pre-specified metric on a huge sample is still worthless if one arm saw easier inputs. You do not get to fix the most obvious problem and salvage the conclusion; a sound result requires all three to be right at once. And in practice they cluster, the same informal, rushed process that undersizes a sample also tends to skip randomization and shop for a metric, so a flawed experiment usually carries more than one of these defects.
What the exam trips candidates on
The first trap is deploying a change because a 50-session comparison showed a 6-point improvement, without checking sample size, input-distribution balance, or metric pre-specification. A scenario will present exactly that small comparison as grounds to ship; the credited reading refuses, and names which of the three failures apply.
The second trap is attributing an early apparent win entirely to the change when the treatment arm happened to receive an easier input mix. A scenario will show an apparent gain that is really an input-distribution artifact; the correct answer separates the change's effect from the sampling artifact and declines to credit the change. Both traps reward diagnosing the specific failures rather than accepting a convincing-looking number.
Worked example
A team ran a new customer-service prompt against 50 sessions and the old prompt against another 50, saw task success at 68% versus 62%, declared a winner, and deployed. Two weeks later the new prompt's task success settled at 61%. Diagnose everything that went wrong.
The two-week collapse from 68% to 61% is the signature of a result that was never real, and all three failure modes are in play.
Undersized sample: 50 sessions per arm is far too small for a noisy LLM metric like task success. A 6-point gap at that size is within the noise floor, the range random variation produces on its own. The correct sample, computed from the effect, baseline, and confidence with LLM variance accounted for, would run into the hundreds per arm. That the gap evaporated at scale is exactly what an underpowered result does.
Uncontrolled input distribution: the 50 treatment sessions happened to contain fewer edge-case inputs than the 50 control sessions. Part of the apparent 6-point gain was therefore an artifact of which inputs landed in which arm, not an effect of the prompt. Even a larger sample would carry this bias if assignment was not randomized and held consistent.
Post-hoc metric selection: the team reported task success because it moved in the right direction. Had a different metric moved instead, they would have reported that one. Choosing the metric after seeing the results means no result could have counted as a failure, which turns the exercise into a search rather than a test.
Any one of these invalidates the deployment decision, and all three are present. The correct process would have pre-specified a single primary metric with a threshold, randomized assignment so the arms were comparable, and run to a properly powered sample before concluding anything. On this evidence, the honest verdict is that the new prompt was never shown to be better, and the 61% settling point suggests it was not.
Common misreadings to avoid
Misconception
A 6-point improvement on a 50-session-per-arm comparison is enough to justify deploying the change.
What's actually true
Misconception
If the treatment arm scored higher, the change caused the improvement.
What's actually true
How this shows up on the exam
Domain 4 questions on this knowledge point present a small, informal comparison declared a winner and ask you to diagnose it. The reliable approach checks all three independent failures, sample size, input-distribution balance, and metric pre-specification, recognizes that any one is fatal, and declines to credit the change until a properly designed experiment confirms it.
This is the evaluate-level capstone of the experimentation task statement, drawing the undersizing failure from sample size and statistical power and the missing-metric failure from a falsifiable hypothesis, both grounded in the five components of an A/B test. It rhymes structurally with diagnosing a stale eval suite and change attribution, two other places where a confident-looking signal hides a broken measurement.
A team deploys a prompt after a 50-session-per-arm comparison showed task success at 68% vs 62%; two weeks later it settles at 61%. Which diagnosis is most complete?
People also ask
What makes an A/B test result invalid?
What is outcome-shopping in an experiment?
How can input distribution bias an A/B test?
Watch and learn
Official Anthropic Academy lessons first, then hand-picked walkthroughs. Videos load only when you press play.
No videos curated for this concept yet
We are still curating the best official and community videos for this topic.
Official prep for this domain
Anthropic's own free prep module for this part of the syllabus, on the official prep course. Free with an Anthropic Academy sign-in.
References & primary sources
Master this concept with Archie
Practice it inside an adaptive study session. Archie, your Socratic AI tutor, tracks your mastery with Bayesian Knowledge Tracing and schedules the perfect next review.