A cornerback jogs off after a hard hit, one hand pressed to the side of his helmet. The athletic trainer runs him through the sideline checklist, then reaches for the Balance Error Scoring System sheet and counts six errors on tandem stance with his eyes closed. Six errors sounds fine against a published norm table built from a few hundred college athletes. It sounds a lot less fine once someone remembers that this particular kid tested at two errors during preseason, on the same foam pad, with the same tester, at the same time of day. The number that decides whether he practices Wednesday is not the population average. It is the gap between today and his own baseline, and that comparison only works if the baseline was collected under conditions rigid enough to trust.
Most programs run a baseline balance test because a checklist somewhere says to. Far fewer run it the same way twice: same surface, same footwear, same tester, same number of trials, banked before the season rather than during a rushed physical day. This protocol standardizes preseason balance baselining as an individual control a trainer can actually stand behind on a return-to-play decision, not a form that gets filed and forgotten.
Why a Team Average Cannot Answer a Return-to-Play Question
Why a Team Average Cannot Answer a Return-to-Play Question
Postural control varies enormously between healthy, uninjured athletes for reasons that have nothing to do with brain injury. A gymnast or a wrestler with years of proprioceptive training routinely posts single-digit BESS totals. A 200kg offensive lineman with two prior ankle sprains and stiff ankles can post totals in the high teens on a good day, foam pad included, and still be completely uninjured. Age matters too: adolescent athletes, particularly younger high schoolers, tend to score worse than college-age athletes on foam-surface stances simply because the vestibular and somatosensory systems are still maturing.
Compare either athlete to a single published norm table and one of two things happens. The lineman gets flagged as impaired when he is not, which either benches him unnecessarily or teaches the staff to stop trusting the test. Or the gymnast clears a real deficit because her post-injury score of five errors still beats a team-wide cutoff of eight, even though five errors is nearly triple her own resting number. A baseline collected on that specific athlete, under the same conditions used to retest her after a hit, removes both failure modes: it turns the balance test into a comparison against the one dataset that matters, this athlete on a day nobody suspected anything was wrong.
Equipment and Test Environment
Equipment and Test Environment
The Balance Error Scoring System (BESS) needs almost nothing exotic. What it needs is consistency, because the score is sensitive to surface, footwear, and even room noise.
| Item | Specification | Why it matters |
|---|---|---|
| Firm surface | Level, non-slip floor | Baseline for the three firm-surface stances |
| Foam pad | Medium-density foam, roughly 6cm thick (an Airex-style balance pad) | Density and thickness change task difficulty; swapping brands mid-season shifts scores independent of ability |
| Stopwatch or timer app | Audible cue at 20 seconds | Each trial runs 20 seconds; testers need a consistent signal to start and stop counting |
| Quiet, distraction-limited room | No open windows onto a practice field, no phone alerts | Visual or auditory distraction increases errors independent of true balance ability |
| Standardized footwear | Barefoot or thin socks, same choice every session | Shoes change proprioceptive feedback at the foot; a shod baseline compared to a barefoot retest is not a valid comparison |
Test in a rested state, not right after a training session or on a day the athlete reports poor sleep or an unrelated lower-body ache. Fatigue and soreness from a hard leg day both elevate error counts in athletes with normal vestibular function, and a baseline collected on a bad day becomes a bad reference point for the rest of the season.
Step-by-Step Preseason Baseline Protocol
Step-by-Step Preseason Baseline Protocol
The athlete stands with hands on the iliac crests, eyes closed, holding each of six stance-and-surface combinations for 20 seconds while a tester counts errors. Total possible score is 60, with a 10-error maximum per 20-second trial; lower totals mean better postural control.
- Double-leg stance, firm surface: Feet together, eyes closed, 20 seconds.
- Single-leg stance, firm surface: Standing on the non-dominant leg, opposite hip flexed to roughly 20-30 degrees and knee flexed to roughly 45 degrees, eyes closed, 20 seconds.
- Tandem stance, firm surface: Non-dominant foot directly behind the dominant foot, heel to toe, eyes closed, 20 seconds.
- Double-leg stance, foam pad: Same as step 1, standing on the foam pad.
- Single-leg stance, foam pad: Same as step 2, standing on the foam pad.
- Tandem stance, foam pad: Same as step 3, standing on the foam pad.
Six error types get tallied per trial, capped at 10: lifting the hands off the iliac crests, opening the eyes, a step, stumble, or fall, moving the hip past 30 degrees of flexion or abduction, lifting the forefoot or heel off the surface, and staying out of position for more than 5 seconds. Sum errors across all six trials for the total score.
Run two full baseline sessions, ideally 24-72 hours apart during preseason, and record both totals rather than only the second, better-looking one.
Why One Baseline Trial Is Not Enough
Why One Baseline Trial Is Not Enough
Valovich, Perrin, and Gansneder (2003), publishing in the Journal of Athletic Training, tested healthy, non-concussed high school athletes on the BESS multiple times across a several-day window and found a statistically significant practice effect: total error counts dropped meaningfully from the first session to later sessions purely from familiarity with the task, even though nothing about the athletes' actual balance had changed. The same study found no comparable practice effect on the Standardized Assessment of Concussion, helping confirm the BESS finding was not just general test-retest noise.
The practical problem this creates is specific: an athlete's first exposure to the test, run once during a rushed physical day, tends to produce a worse score than his true resting ability, because he has not yet learned how to hold tandem stance with his eyes closed on a foam pad. Bank that inflated single trial as the season's baseline, and a genuinely concerning post-concussion score can look like a small, unremarkable increase rather than the real deficit it is. Running two sessions during preseason and averaging them, or treating the first explicitly as familiarization and banking the second as the baseline of record, closes most of that gap.
What the Concussion Literature Shows
What the Concussion Literature Shows
Riemann and Guskiewicz (2000), in the Journal of Athletic Training, compared BESS scores from concussed college athletes against their own preseason baselines and against a matched control group. Concussed athletes showed a clear, large elevation in total errors within the first 24 to 72 hours post-injury relative to their individual baseline, consistent with a large effect on the total score, before performance trended back down toward baseline over the following three to five days for most athletes. The finding that matters most for building a protocol is not simply that concussion worsens balance; it is that the deficit is largest and most detectable in the first day or two and fades quickly, so a comparison run a week after injury is far less useful than one run at 24 and 72 hours.
Finnoff, Peterson, Hollman, and Smith (2009), in PM&R, examined how consistent BESS scores are when the same tester scores an athlete twice (intrarater reliability) versus when two different testers score the same trial (interrater reliability). Intrarater agreement came out meaningfully stronger than interrater agreement across most stances, pointing to a limitation of manual scoring: because errors are counted by eye in real time, switching which staff member runs the retest introduces noise unrelated to the athlete's actual balance. Training one dedicated tester, or standardizing scoring criteria tightly across staff, protects the comparison; letting whoever is free that day run the test does not.
Turning a Post-Injury Score Into a Decision
Turning a Post-Injury Score Into a Decision
Because BESS scores carry measurement error even in a healthy athlete retested on an ordinary day, not every increase means something happened. Published test-retest data put the standard error of measurement for the total score at roughly 3-4 points, a rough floor for what counts as real change rather than day-to-day noise.
| Change from own baseline | Interpretation |
|---|---|
| 0 to 3 points higher | Within typical measurement noise; not evidence of a balance deficit on its own |
| 4 to 6 points higher | Borderline; retest within 24 hours and weigh alongside symptom and cognitive findings rather than treating balance alone as decisive |
| 7 or more points higher | Consistent with a meaningful postural control deficit; hold from balance-dependent activity and retest before progressing return-to-play stages |
These bands describe the total score across all six trials, not a single stance. A jump concentrated in the two foam-pad stances, with firm-surface trials unchanged, is itself informative: it suggests the deficit is specific to the harder sensory-conflict conditions rather than a global collapse in postural control.
Mistakes That Undermine the Baseline
Mistakes That Undermine the Baseline
| Mistake | Effect | Fix |
|---|---|---|
| Recording only one preseason trial | Bakes the practice effect into the season's reference number | Run two sessions 24-72 hours apart and average, or bank session two as the baseline of record |
| Testing right after a hard conditioning session | Fatigue inflates error counts independent of true balance | Test in a rested state, ideally at the same time of day used for future retests |
| Swapping foam pad brand or density mid-season | Shifts scores on the three foam-pad-dependent trials | Keep the same physical pad, or at minimum the same specification, for baseline and every retest |
| Letting whichever staff member is free run the test | Interrater scoring differences add noise the athlete never produced | Assign one primary tester per athlete, or train all testers against identical scoring criteria |
| Comparing a post-injury score to a team average instead of the athlete's own baseline | Misses real deficits in athletes with strong baseline balance and flags healthy athletes with naturally higher baseline totals | Always compare to the individual's stored baseline first |
Where Balance Testing Fits in the Return-to-Play Decision
Where Balance Testing Fits in the Return-to-Play Decision
A clean balance score never clears an athlete on its own, and a single elevated score never holds one out on its own either. Balance testing is one leg of a multidimensional battery that also includes symptom checklists, cognitive screening, and a graded exertion protocol, and its real value is catching athletes whose symptoms have resolved on paper but whose postural control has not caught up yet. Retest at 24 and 72 hours post-injury while the deficit is most detectable, then again before each step-up in the return-to-play progression, since exertion can transiently worsen balance in an athlete who otherwise looks recovered at rest.
Re-baseline every preseason rather than carrying a number forward for years, particularly for adolescent athletes whose vestibular systems continue maturing, and for any athlete with a significant new lower-limb injury, since ankle or knee mechanics feed directly into the somatosensory input the test measures. A three-year-old baseline on a 14-year-old who is now 17 and coming off an ACL reconstruction is a number from a different athlete.
Frequently asked questions
01How many trials make up a proper baseline balance test?+
02My score this week is a few points worse than my preseason number, but I mostly feel fine. Does that mean something is wrong?+
03What happens if the foam pad brand changes between the baseline test and a later retest?+
04How soon after a suspected concussion should the balance retest happen?+
05Can a team-wide balance norm substitute for individual baseline testing?+
Related Articles
How to Perform the Y-Balance Test: Dynamic Balance Screening
A single asymmetry score doesn't reveal injury risk alone. Y-Balance Test protocol: asymmetry thresholds, sport norms, and injury-risk scoring rules.
Neck Strength and Concussion Risk: What Studies Show
Coaches ask if neck exercises stop concussions. A 6,704-athlete study found 5% lower odds per pound of strength - here's what it does and doesn't prove.
How to Test Single-Leg Asymmetry: A Complete 800Hz IMU Protocol
A 10% gap between legs raises injury risk. This 5-step 800Hz IMU protocol finds single-leg asymmetry and shows how LSI data should change your programming.
How to Assess Landing Mechanics for ACL Prevention
One landing pattern predicts most non-contact ACL injuries. Score it with the drop-landing protocol and LESS criteria, then apply corrective progressions.
Ankle Sprain Return-to-Play: Hop Test and Balance Cutoffs Before Cutting Resumes
Pain-free jogging isn't clearance to cut. A hop-and-balance protocol with the LSI and reach cutoffs research actually supports before cutting resumes.
Concussion Collision Readiness: Dual-Task Reactive Balance and Neck Stability
Symptom-free and a clean BESS score are not collision-ready. A dual-task reactive-balance and neck-stability protocol for the final call before full contact.
Star Excursion Balance Test: Protocol and Normalizing Reach Distance to Limb Length
A taller athlete's raw SEBT reach looks better on paper, until you normalize to limb length. Full 8-direction protocol, the math, and injury-risk cutoffs.
Curling Delivery Slide Stability Test: Measuring Lunge-Hold Sway
A stone drifting wide despite consistent weight often traces to slide-hold sway. An IMU protocol for measuring mediolateral wobble in the curling delivery.
Measure performance with lab-grade accuracy