A center back retests his countermovement jump six weeks into a power block. Pre-test: 38.6 cm. Post-test: 40.0 cm. That is a 1.4 cm gain, and it is the kind of number that ends up in a PDF report with a green up-arrow next to it. The problem is that a CMJ, tested with a jump mat or force plate under normal field conditions, has a typical error of roughly 1.0 to 1.5 cm session to session in a trained athlete who has done nothing at all between tests. So that 1.4 cm "improvement" sits almost entirely inside the test's own noise band. It might be real. It might be a slightly different warm-up, a slightly different night of sleep, or the athlete standing a centimeter differently on the mat.
This is the single most common statistical mistake in athlete testing: treating every number that goes up as an improvement. The fix does not require a statistics degree. It requires two numbers calculated once for your test and your population — typical error (TE) and smallest worthwhile change (SWC) — and a simple rule for comparing an observed change against both.
The Problem: A Gain That Might Not Be a Gain
The Problem: A Gain That Might Not Be a Gain
Every performance test — CMJ, 1RM, 10m sprint, isometric mid-thigh pull — produces a different number every time you run it, even with zero true change in the athlete. Warm-up quality, time of day, motivation, mat calibration, and simple biological noise all add variance. Sports scientists call this measurement noise, and it never goes to zero no matter how good the equipment is.
The practical question a coach actually needs answered is not "did the number go up?" It is "did the number go up by more than the test's own noise, and if so, is that amount of change big enough to matter for performance?" Those are two separate questions, and they require two separate numbers. Confusing them is why so many training logs are full of athletes flagged as "improving" or "declining" based on changes that a repeat test on the same day, with no intervention at all, would have produced just as easily.
Two Numbers You Need Before You Trust Any Retest
Two Numbers You Need Before You Trust Any Retest
Typical error (TE) is the noise floor of your test. It is derived from repeated trials on the same athletes under stable conditions (no training effect expected between sessions), and it answers: how much does this number bounce around on its own? TE is calculated from the standard deviation of the difference between two repeat tests, divided by the square root of 2: TE = SD(diff) / √2. Expressed as a percentage of the mean score, it becomes the coefficient of variation (CV%) — the version most coaches actually quote.
Smallest worthwhile change (SWC) is a completely different question: even assuming the change is real, how big does it need to be to matter for performance? Hopkins' widely used convention sets SWC at 0.2 times the between-athlete standard deviation for the group being tested — borrowed from Cohen's "small effect size" threshold. It is a statistical proxy for practical importance, not a measurement of noise.
The relationship between the two numbers is what actually tells you whether a test result is trustworthy. If TE is smaller than SWC, your test is precise enough that a change of practical importance will usually be visible above the noise. If TE is larger than SWC, a change that matters can easily hide inside measurement error — and a change that shows up clearly might still be too small to matter competitively.
Where TE and SWC Come From: The Research
Where TE and SWC Come From: The Research
The framework did not originate in a coaching manual — it came out of sports statistics research trying to solve exactly this problem for exercise scientists running small-sample studies.
| Study | Contribution | Key Figure / Effect Size | Limitation |
|---|---|---|---|
| Hopkins WG (2000), Sports Medicine, 30(1):1-15 | Formalized typical error as SD(diff)/√2 and proposed SWC = 0.2 × between-subject SD, adapting Cohen's small-effect convention to athletic performance | SWC threshold set at 0.2 SD ("small" effect); 0.5 SD reserved for "moderate" changes worth a bigger reaction | The 0.2 SD threshold is a statistical convention, not derived from what actually predicts winning in any given sport — it needs local calibration against competition outcomes where possible |
| Hopkins WG (2004), Sportscience, 8:1-7, "How to Interpret Changes in an Athletic Performance Test" | Introduced the SWC:TE ratio as a test-quality rating (good, OK, marginal) and showed how to combine both numbers into a single decision rule for retest interpretation | Ratio thresholds: good >1.5, OK 0.5-1.5, marginal <0.5 (SWC divided by TE) | Ratings depend entirely on the quality of the underlying TE estimate; a TE calculated from too few reliability trials (fewer than ~20 paired tests) is itself noisy and can mislabel a test |
| Turner AN, Brazier J, Bishop C, et al. (2015), Strength and Conditioning Journal, 37(1):76-83 | Published a step-by-step spreadsheet method for calculating TE, CV%, and SWC from raw session data using log-transformed scores | Demonstrated with CMJ data showing CV% in the 3-5% range for field-based teams — enough to swallow small individual gains | Log-transformation assumes the underlying error is proportional (multiplicative) to score magnitude, which fits jump and strength tests reasonably well but can misrepresent tests with a hard floor or ceiling |
The consistent finding across this line of research: for most field tests used in team sport, typical error and smallest worthwhile change sit close enough together that the difference between a "good" and a "marginal" test often comes down to testing discipline — standardized warm-up, fixed time of day, and enough familiarization trials — rather than the equipment itself.
Calculating TE and SWC From Your Own Squad
Calculating TE and SWC From Your Own Squad
Published TE and SWC values are a starting point, not a substitute for your own numbers — your athletes, equipment, and protocol all shift the estimate. The process takes one extra testing session and about twenty minutes of spreadsheet work.
- Run a reliability pair. Test the full squad twice, 24-72 hours apart, zero training load between sessions, same time of day, warm-up, equipment, and tester. Fifteen to twenty athletes is the practical minimum; fewer than ten produces a TE that swings wildly with each new data point.
- Take the best trial from each session if that is also your normal testing protocol — TE has to match how the test is actually used day to day.
- Calculate each athlete's difference score (session 2 minus session 1), then the standard deviation of those difference scores across the squad.
- Divide by √2 (1.414) to get TE in the test's native units; divide TE by the group mean and multiply by 100 for CV%.
- Take the between-athlete SD from either session's raw scores (a single session's SD is normally an adequate estimate).
- Multiply that SD by 0.2 for SWC, and by 0.5 for a moderate-change threshold worth a stronger reaction.
Re-run this check roughly twice a year, and whenever equipment, tester, or protocol changes — TE is a property of your specific setup, not a fixed property of the test itself.
Worked Example: Three Athletes, Same Retest, Three Different Verdicts
Worked Example: Three Athletes, Same Retest, Three Different Verdicts
Take a 22-player academy squad. A reliability pair on CMJ height (best of three trials, arms-akimbo, jump mat) produces a between-session SD of difference scores of 1.55 cm, giving TE = 1.55 / 1.414 = 1.10 cm (CV ≈ 3.1% against a squad mean of 35.4 cm). The squad's between-athlete SD on a single session is 4.0 cm, so SWC (0.2 × SD) = 0.80 cm and the moderate-change threshold (0.5 × SD) = 2.00 cm. Three athletes retest six weeks later after an identical power block:
| Athlete | Pre (cm) | Post (cm) | Observed change | vs. TE (1.10 cm) | vs. SWC (0.80 cm) | Verdict |
|---|---|---|---|---|---|---|
| A | 34.0 | 34.3 | +0.3 cm | Below TE | Below SWC | Noise. No usable evidence of change either way. |
| B | 36.0 | 37.0 | +1.0 cm | Below TE | Above SWC | Ambiguous. Large enough to matter if real, but too small to separate confidently from measurement error — this is the exact zone this test cannot resolve well. |
| C | 33.5 | 34.9 | +1.4 cm | Above TE | Above SWC | Real and worth reacting to. The change exceeds both the noise floor and the practical-importance threshold. |
Athlete B is the case that trips up most testing programs. A gain of 1.0 cm clears the bar for "would matter if true" but sits inside a test whose own noise band is 1.10 cm — meaning a repeat test the same week, with no training at all, could easily produce a swing of that size in either direction. The honest read for Athlete B is not "improved" or "unchanged" — it is "retest before deciding anything," ideally on a day with matched conditions to the original test.
Rating Whether a Test Is Even Worth Trusting
Rating Whether a Test Is Even Worth Trusting
Before interpreting individual results, it is worth asking a blunter question: is this particular test, on this particular squad, precise enough to ever detect a worthwhile change at the individual level? Divide SWC by TE and use Hopkins' rating scale.
| SWC ÷ TE Ratio | Rating | What It Means in Practice |
|---|---|---|
| > 1.5 | Good | A practically important change will usually clear the noise floor; individual-level retests can be trusted with reasonable confidence. |
| 0.5 - 1.5 | OK | Individual changes near the SWC threshold (like Athlete B above) will often be ambiguous; lean on squad-average trends and repeated testing rather than single retests. |
| < 0.5 | Marginal | The test's noise is large relative to what matters; single retests at the individual level are close to meaningless — use it only for group-level tracking or replace it with a lower-noise metric. |
The CMJ example above rates 0.80 / 1.10 = 0.73 — solidly "OK." That is a realistic result for a field-tested jump metric and matches the CV ranges reported by Turner et al. (2015): usable, but not precise enough to hang individual programming decisions on a single retest without corroborating data.
Where Coaches Get This Wrong
Where Coaches Get This Wrong
Three mistakes account for most of the bad calls in retest interpretation. The first is using standard deviation where typical error belongs — SD describes how spread out a squad's scores are, not how much a single athlete's score bounces between sessions. Plugging squad SD into a noise-floor decision answers the wrong question entirely, even though the resulting number looks plausible on a report.
The second is borrowing a published SWC or TE from a different population and assuming it transfers. A CMJ TE from elite sprinters on a lab force plate does not describe a 16-year-old academy squad tested on a jump mat in a gym hallway — equipment, training age, and testing environment can move the number by a factor of two.
The third is applying SWC without ever calculating TE — flagging any change larger than 0.2 SD as "real" with no check on whether that size of change is even distinguishable from noise. This is exactly how Athlete B in the worked example above gets misfiled as a confirmed improvement when the honest answer is "we don't know yet."
Frequently asked questions
01What is the difference between typical error and smallest worthwhile change?+
02How many athletes do I need to calculate a reliable TE?+
03My athlete's number went up by more than the SWC but I'm still not confident it's real. Why?+
04Does a higher SWC:TE ratio mean I should test more often?+
05Can I use published TE and SWC values instead of calculating my own?+
Related Articles
Inter-Individual Response Variability: Why Same Program Produces Different Results
481 people did the identical 20-week program; VO2max changes ranged from -6% to +100%. The genetic and lifestyle factors behind high vs low responders.
CMJ as a Monitoring Tool: Research Review
Not every CMJ metric tracks fatigue equally well. This review ranks jump height, RSI, and flight-time ratio by reliability, with thresholds coaches use.
Load-Velocity Profiling for 1RM Prediction: Accuracy Review
Can load-velocity profiling replace a maximal 1RM test? This review breaks down error rates by method and which protocols hold up in practice.
Velocity-Based 1RM Prediction Accuracy: How Error Varies by Exercise
The back squat's minimal velocity threshold sits near 0.30 m/s, but error swings widely — free-weight hip lifts are less predictable than machines.
Monitoring Training Load: Research on Best Practices
An ACWR above 1.5 links to 2-4x higher injury risk, yet explains only 10-15% of variance. Compare ACWR, RPE, and velocity monitoring on sensitivity.
CMJ Monitoring for Athlete Readiness: Research on Countermovement Jump as a Fatigue and...
A lower jump isn't always fatigue, but research points to thresholds that matter. Here's how CMJ flags overreaching and guides same-day training calls.
Hamstring-to-Quadriceps Isokinetic Ratio Test: Reading Injury Risk Beyond the Sprint Score
A H:Q ratio under 0.47 or a 15% bilateral gap flags real hamstring injury risk on isokinetic testing. Here is the full protocol, norms, and interpretation.
Tethered Swim Force Test: Validity, Setup, and Interpretation
A fixed rope and load cell measure real propulsive force in the water, not a land proxy. The rig, protocol, and force ranges that actually hold up.
Measure performance with lab-grade accuracy