PoinT GOResearch
how to·how to

Concussion Baseline Balance Testing Protocol: Equipment, Steps, and Return-to-Play Cutoffs

Team norms miss individual deficits. A preseason BESS baseline protocol with scoring steps, the practice-effect fix, and return-to-play interpretation.

PoinT GO Research Team··9 min read
Concussion Baseline Balance Testing Protocol: Equipment, Steps, and Return-to-Play Cutoffs

A cornerback jogs off after a hard hit, one hand pressed to the side of his helmet. The athletic trainer runs him through the sideline checklist, then reaches for the Balance Error Scoring System sheet and counts six errors on tandem stance with his eyes closed. Six errors sounds fine against a published norm table built from a few hundred college athletes. It sounds a lot less fine once someone remembers that this particular kid tested at two errors during preseason, on the same foam pad, with the same tester, at the same time of day. The number that decides whether he practices Wednesday is not the population average. It is the gap between today and his own baseline, and that comparison only works if the baseline was collected under conditions rigid enough to trust.

Most programs run a baseline balance test because a checklist somewhere says to. Far fewer run it the same way twice: same surface, same footwear, same tester, same number of trials, banked before the season rather than during a rushed physical day. This protocol standardizes preseason balance baselining as an individual control a trainer can actually stand behind on a return-to-play decision, not a form that gets filed and forgotten.

Why a Team Average Cannot Answer a Return-to-Play Question

Why a Team Average Cannot Answer a Return-to-Play Question

Postural control varies enormously between healthy, uninjured athletes for reasons that have nothing to do with brain injury. A gymnast or a wrestler with years of proprioceptive training routinely posts single-digit BESS totals. A 200kg offensive lineman with two prior ankle sprains and stiff ankles can post totals in the high teens on a good day, foam pad included, and still be completely uninjured. Age matters too: adolescent athletes, particularly younger high schoolers, tend to score worse than college-age athletes on foam-surface stances simply because the vestibular and somatosensory systems are still maturing.

Compare either athlete to a single published norm table and one of two things happens. The lineman gets flagged as impaired when he is not, which either benches him unnecessarily or teaches the staff to stop trusting the test. Or the gymnast clears a real deficit because her post-injury score of five errors still beats a team-wide cutoff of eight, even though five errors is nearly triple her own resting number. A baseline collected on that specific athlete, under the same conditions used to retest her after a hit, removes both failure modes: it turns the balance test into a comparison against the one dataset that matters, this athlete on a day nobody suspected anything was wrong.

Equipment and Test Environment

Equipment and Test Environment

The Balance Error Scoring System (BESS) needs almost nothing exotic. What it needs is consistency, because the score is sensitive to surface, footwear, and even room noise.

ItemSpecificationWhy it matters
Firm surfaceLevel, non-slip floorBaseline for the three firm-surface stances
Foam padMedium-density foam, roughly 6cm thick (an Airex-style balance pad)Density and thickness change task difficulty; swapping brands mid-season shifts scores independent of ability
Stopwatch or timer appAudible cue at 20 secondsEach trial runs 20 seconds; testers need a consistent signal to start and stop counting
Quiet, distraction-limited roomNo open windows onto a practice field, no phone alertsVisual or auditory distraction increases errors independent of true balance ability
Standardized footwearBarefoot or thin socks, same choice every sessionShoes change proprioceptive feedback at the foot; a shod baseline compared to a barefoot retest is not a valid comparison

Test in a rested state, not right after a training session or on a day the athlete reports poor sleep or an unrelated lower-body ache. Fatigue and soreness from a hard leg day both elevate error counts in athletes with normal vestibular function, and a baseline collected on a bad day becomes a bad reference point for the rest of the season.

Step-by-Step Preseason Baseline Protocol

Step-by-Step Preseason Baseline Protocol

The athlete stands with hands on the iliac crests, eyes closed, holding each of six stance-and-surface combinations for 20 seconds while a tester counts errors. Total possible score is 60, with a 10-error maximum per 20-second trial; lower totals mean better postural control.

  1. Double-leg stance, firm surface: Feet together, eyes closed, 20 seconds.
  2. Single-leg stance, firm surface: Standing on the non-dominant leg, opposite hip flexed to roughly 20-30 degrees and knee flexed to roughly 45 degrees, eyes closed, 20 seconds.
  3. Tandem stance, firm surface: Non-dominant foot directly behind the dominant foot, heel to toe, eyes closed, 20 seconds.
  4. Double-leg stance, foam pad: Same as step 1, standing on the foam pad.
  5. Single-leg stance, foam pad: Same as step 2, standing on the foam pad.
  6. Tandem stance, foam pad: Same as step 3, standing on the foam pad.

Six error types get tallied per trial, capped at 10: lifting the hands off the iliac crests, opening the eyes, a step, stumble, or fall, moving the hip past 30 degrees of flexion or abduction, lifting the forefoot or heel off the surface, and staying out of position for more than 5 seconds. Sum errors across all six trials for the total score.

Run two full baseline sessions, ideally 24-72 hours apart during preseason, and record both totals rather than only the second, better-looking one.

Why One Baseline Trial Is Not Enough

Why One Baseline Trial Is Not Enough

Valovich, Perrin, and Gansneder (2003), publishing in the Journal of Athletic Training, tested healthy, non-concussed high school athletes on the BESS multiple times across a several-day window and found a statistically significant practice effect: total error counts dropped meaningfully from the first session to later sessions purely from familiarity with the task, even though nothing about the athletes' actual balance had changed. The same study found no comparable practice effect on the Standardized Assessment of Concussion, helping confirm the BESS finding was not just general test-retest noise.

The practical problem this creates is specific: an athlete's first exposure to the test, run once during a rushed physical day, tends to produce a worse score than his true resting ability, because he has not yet learned how to hold tandem stance with his eyes closed on a foam pad. Bank that inflated single trial as the season's baseline, and a genuinely concerning post-concussion score can look like a small, unremarkable increase rather than the real deficit it is. Running two sessions during preseason and averaging them, or treating the first explicitly as familiarization and banking the second as the baseline of record, closes most of that gap.

What the Concussion Literature Shows

What the Concussion Literature Shows

Riemann and Guskiewicz (2000), in the Journal of Athletic Training, compared BESS scores from concussed college athletes against their own preseason baselines and against a matched control group. Concussed athletes showed a clear, large elevation in total errors within the first 24 to 72 hours post-injury relative to their individual baseline, consistent with a large effect on the total score, before performance trended back down toward baseline over the following three to five days for most athletes. The finding that matters most for building a protocol is not simply that concussion worsens balance; it is that the deficit is largest and most detectable in the first day or two and fades quickly, so a comparison run a week after injury is far less useful than one run at 24 and 72 hours.

Finnoff, Peterson, Hollman, and Smith (2009), in PM&R, examined how consistent BESS scores are when the same tester scores an athlete twice (intrarater reliability) versus when two different testers score the same trial (interrater reliability). Intrarater agreement came out meaningfully stronger than interrater agreement across most stances, pointing to a limitation of manual scoring: because errors are counted by eye in real time, switching which staff member runs the retest introduces noise unrelated to the athlete's actual balance. Training one dedicated tester, or standardizing scoring criteria tightly across staff, protects the comparison; letting whoever is free that day run the test does not.

Turning a Post-Injury Score Into a Decision

Turning a Post-Injury Score Into a Decision

Because BESS scores carry measurement error even in a healthy athlete retested on an ordinary day, not every increase means something happened. Published test-retest data put the standard error of measurement for the total score at roughly 3-4 points, a rough floor for what counts as real change rather than day-to-day noise.

Change from own baselineInterpretation
0 to 3 points higherWithin typical measurement noise; not evidence of a balance deficit on its own
4 to 6 points higherBorderline; retest within 24 hours and weigh alongside symptom and cognitive findings rather than treating balance alone as decisive
7 or more points higherConsistent with a meaningful postural control deficit; hold from balance-dependent activity and retest before progressing return-to-play stages

These bands describe the total score across all six trials, not a single stance. A jump concentrated in the two foam-pad stances, with firm-surface trials unchanged, is itself informative: it suggests the deficit is specific to the harder sensory-conflict conditions rather than a global collapse in postural control.

Mistakes That Undermine the Baseline

Mistakes That Undermine the Baseline

MistakeEffectFix
Recording only one preseason trialBakes the practice effect into the season's reference numberRun two sessions 24-72 hours apart and average, or bank session two as the baseline of record
Testing right after a hard conditioning sessionFatigue inflates error counts independent of true balanceTest in a rested state, ideally at the same time of day used for future retests
Swapping foam pad brand or density mid-seasonShifts scores on the three foam-pad-dependent trialsKeep the same physical pad, or at minimum the same specification, for baseline and every retest
Letting whichever staff member is free run the testInterrater scoring differences add noise the athlete never producedAssign one primary tester per athlete, or train all testers against identical scoring criteria
Comparing a post-injury score to a team average instead of the athlete's own baselineMisses real deficits in athletes with strong baseline balance and flags healthy athletes with naturally higher baseline totalsAlways compare to the individual's stored baseline first

Where Balance Testing Fits in the Return-to-Play Decision

Where Balance Testing Fits in the Return-to-Play Decision

A clean balance score never clears an athlete on its own, and a single elevated score never holds one out on its own either. Balance testing is one leg of a multidimensional battery that also includes symptom checklists, cognitive screening, and a graded exertion protocol, and its real value is catching athletes whose symptoms have resolved on paper but whose postural control has not caught up yet. Retest at 24 and 72 hours post-injury while the deficit is most detectable, then again before each step-up in the return-to-play progression, since exertion can transiently worsen balance in an athlete who otherwise looks recovered at rest.

Re-baseline every preseason rather than carrying a number forward for years, particularly for adolescent athletes whose vestibular systems continue maturing, and for any athlete with a significant new lower-limb injury, since ankle or knee mechanics feed directly into the somatosensory input the test measures. A three-year-old baseline on a 14-year-old who is now 17 and coming off an ACL reconstruction is a number from a different athlete.

FAQ

Frequently asked questions

01How many trials make up a proper baseline balance test?
+
Run the full six-trial BESS battery twice, ideally 24-72 hours apart during preseason, and record both total scores. Averaging the two sessions, or treating the first as familiarization and banking the second, corrects for the practice effect documented when athletes take the test for the first time.
02My score this week is a few points worse than my preseason number, but I mostly feel fine. Does that mean something is wrong?
+
Not necessarily. Published test-retest data put the measurement error for the total BESS score at roughly 3-4 points, so small shifts happen in healthy athletes retested on an ordinary day. A change of 7 or more points from your own baseline is where it starts looking like a real deficit rather than normal noise.
03What happens if the foam pad brand changes between the baseline test and a later retest?
+
Density and thickness differences between foam pads change how hard the balance task is, which shifts scores on the three foam-pad trials independent of any real change in the athlete's balance. Keep the same physical pad, or at least the same specification, for every session across a season.
04How soon after a suspected concussion should the balance retest happen?
+
Research on postural control after concussion has found the deficit is largest and most detectable in the first 24 to 72 hours, with most athletes trending back toward their baseline over the following three to five days. Retesting inside that early window, and again before each stage of a graded return-to-play progression, catches the deficit when it is easiest to see.
05Can a team-wide balance norm substitute for individual baseline testing?
+
It substitutes poorly. Postural control varies widely between healthy athletes based on training background, ankle history, and age, so a single population cutoff both flags healthy athletes with naturally higher scores and clears athletes with strong baseline balance who have a real post-injury deficit. An individual baseline, collected under the same conditions used for every retest, avoids both errors.
Keep reading

Related Articles

how to

How to Perform the Y-Balance Test: Dynamic Balance Screening

A single asymmetry score doesn't reveal injury risk alone. Y-Balance Test protocol: asymmetry thresholds, sport norms, and injury-risk scoring rules.

research

Neck Strength and Concussion Risk: What Studies Show

Coaches ask if neck exercises stop concussions. A 6,704-athlete study found 5% lower odds per pound of strength - here's what it does and doesn't prove.

how to

How to Test Single-Leg Asymmetry: A Complete 800Hz IMU Protocol

A 10% gap between legs raises injury risk. This 5-step 800Hz IMU protocol finds single-leg asymmetry and shows how LSI data should change your programming.

how to

How to Assess Landing Mechanics for ACL Prevention

One landing pattern predicts most non-contact ACL injuries. Score it with the drop-landing protocol and LESS criteria, then apply corrective progressions.

how to

Ankle Sprain Return-to-Play: Hop Test and Balance Cutoffs Before Cutting Resumes

Pain-free jogging isn't clearance to cut. A hop-and-balance protocol with the LSI and reach cutoffs research actually supports before cutting resumes.

how to

Concussion Collision Readiness: Dual-Task Reactive Balance and Neck Stability

Symptom-free and a clean BESS score are not collision-ready. A dual-task reactive-balance and neck-stability protocol for the final call before full contact.

how to

Star Excursion Balance Test: Protocol and Normalizing Reach Distance to Limb Length

A taller athlete's raw SEBT reach looks better on paper, until you normalize to limb length. Full 8-direction protocol, the math, and injury-risk cutoffs.

how to

Curling Delivery Slide Stability Test: Measuring Lunge-Hold Sway

A stone drifting wide despite consistent weight often traces to slide-hold sway. An IMU protocol for measuring mediolateral wobble in the curling delivery.

Measure performance with lab-grade accuracy

Get PoinT GO