A physical therapist tests a post-ACL reconstruction patient's quad strength on a Tuesday and records a limb symmetry index of 83%. A different clinician retests the same knee on the same handheld dynamometer three days later and gets 94%. The graft hasn't changed. The rehab hasn't changed. What changed is who was holding the sensor.
This isn't the usual kind of measurement noise - a low battery, a slipped strap, a bad angle. It's the tester's own strength, leverage, and bracing technique leaking directly into the number on the screen. Wikholm and Bohannon flagged this in the Journal of Orthopaedic & Sports Physical Therapy back in 1991: when a tester tries to overpower a strong patient's isometric contraction, part of the resulting score reflects how strong the tester is, not just how strong the patient is. Three decades on, most clinics and performance labs still test the same way - one hand on the dynamometer, a forearm or a knee braced against the limb, hoping the tester's own arm holds long enough to get a real number.
What follows is the mechanical reason make and break testing fail in opposite directions, what the reliability literature says about which muscle groups get hit hardest, and a strap-fixation protocol that takes the tester's own strength almost entirely out of the equation.
Make Test vs. Break Test: Where Tester Strength Enters the Number
Handheld dynamometry (HHD) runs one of two protocols, and each one introduces bias through a different mechanical pathway.
In a break test, the patient contracts maximally and the tester pushes against the limb until the contraction breaks - the limb moves past the sensor. The peak force recorded is capped at whatever the tester can generate before losing position. Test a 220-pound collegiate lineman's knee extensors with a 140-pound clinician holding the dynamometer, and the trial frequently ends because the tester's grip, shoulder, or base of support gave out - not because the athlete's quad reached true failure. The screen reads a number, but that number partly describes the tester, not the athlete.
In a make test, the tester anchors the dynamometer against a fixed point (or braces hard enough to approximate one) and the patient pushes into it without the limb moving. This removes the arm-wrestle problem but introduces a quieter one: if the tester's bracing arm has any give at all, some of the patient's true force gets absorbed into that compliance instead of registering on the load cell. A tester with a soft grip or unstable trunk position will under-read a strong patient on a make test through essentially the same mechanism that lets a weak tester over-cap a strong patient on a break test - the human holding the sensor isn't rigid, and the dynamometer can only be as fixed as they are.
Both failure modes trace back to one root problem: the sensor is only as stationary as the person holding it, and that person varies session to session, clinic to clinic, shift to shift.
| Factor | Break Test | Make Test |
|---|---|---|
| Test ends when | Tester overcomes the patient's contraction | Patient reaches max effort against fixed resistance |
| Tester's role | Active force generator | Passive but rigid anchor |
| Bias direction | Score capped by tester's own strength - ceiling effect on strong patients | Score reduced by tester compliance - floor effect from arm or grip give |
| Groups most affected | Knee extensors, hip abductors, plantarflexors | Any group once the tester fatigues across repeated trials |
| Typical unfixed ICC (knee extension) | 0.60-0.85 | 0.75-0.90 |
| Structural fix | External strap or frame fixation | External strap or frame fixation |
What the Reliability Literature Actually Shows
Wikholm and Bohannon (1991) tested this directly by having testers of varying strength measure the same subjects, then correlating tester strength with the recorded scores. For elbow flexion - a smaller muscle group nearly any adult tester can match or overcome - tester strength barely moved the number. For knee extension and other large lower-body groups, the relationship between tester strength and recorded scores was strong enough that the authors stated it plainly: tester strength changes the reading, independent of how strong the patient actually is.
Stark and colleagues' 2011 systematic review in PM&R found the same pattern from the opposite angle, pooling HHD-versus-isokinetic-dynamometer comparisons across dozens of studies. Correlations between handheld readings and the isokinetic gold standard stayed strong for smaller, weaker muscle groups - elbow and wrist testing often exceeded r = 0.85 - but dropped and grew inconsistent for knee extensors and hip muscles, exactly the groups most likely to overpower an average-build tester in an unfixed break test. The review flagged external fixation - strapping the dynamometer to a frame, wall, or plinth leg rather than the tester's own hand - as the single variable that most consistently closed that gap.
Katoh and colleagues put a number on the fix. Comparing hand-held-only measurement against belt-stabilized measurement for knee extension strength, they reported test-retest ICCs of roughly 0.75-0.88 for hand-held-only testing, depending on how mismatched the tester's and patient's strength were. Adding a stabilization belt - anchoring the dynamometer to a fixed point rather than the tester's arm - pushed ICCs to 0.93-0.98, and, more tellingly, collapsed most of the gap between strong and weak testers. Put a belt in the chain and it stops mattering much whether the person running the test benches 200 pounds or 95.
None of this means handheld dynamometry is unreliable in general - grip strength testing, for instance, is largely self-limited by the patient's own hand and shows minimal tester-strength dependence. The bias is specific to tests where the tester has to generate or resist force against a large muscle group without external help.
A Fixation Protocol That Removes the Tester Variable
Fixing this doesn't require replacing a handheld dynamometer with an isokinetic dynamometer most clinics can't budget for. It requires anchoring the sensor to something other than a person.
Setup. Loop a strap around a stable structure - a plinth leg, a squat rack upright, or a purpose-built stabilization frame - and attach the dynamometer so the strap runs perpendicular to the limb at the contact point. For knee extension, that means securing the strap just proximal to the malleoli with the patient seated, hip and knee at 90°, and the strap pulled to zero slack before the contraction starts.
Familiarization. Run two submaximal contractions at roughly 50% effort before recording anything. This matters more with fixation than without it - patients unconsciously modulate their output against the compliance they feel in a human tester's arm, and a first true maximal effort against a rigid strap often surprises them enough to skew the very first trial.
Recording. Take three trials of 3-5 second maximal isometric contractions with 30-60 seconds of rest between each. Discard any trial whose force-time curve shows a dip-and-recover shape - the signature of a patient bracing, easing off, then re-pushing - rather than a clean rise to plateau. If the two closest trials fall within 10% of each other, average them; if they don't, run a fourth trial before deciding on a value.
Normal ranges and interpretation. On a belt-stabilized HHD, healthy recreationally active adults typically produce roughly 3.0-4.5 N/kg of body mass in knee extension force per limb; trained field-sport athletes often exceed 5.0 N/kg. For return-to-sport decisions, the number that matters most is the limb symmetry index (LSI) - involved-limb force divided by uninvolved-limb force, times 100. An LSI at or above 90% is the threshold most commonly cited for clearing lower-body strength criteria, though it should sit alongside hop testing and readiness screening rather than stand alone.
Here's where tester bias connects directly back to a real decision: an unfixed LSI of 86-89% is genuinely ambiguous. It could reflect a real 11-14% deficit, or it could reflect a smaller tester failing to fully break the involved limb's contraction while comfortably breaking the stronger, uninvolved side. Any borderline result in that range deserves a retest with a fixed strap before it drives a return-to-play call. We've seen LSI shift 6-9 percentage points on retest from fixation alone - enough to flip a clearance decision in either direction. Reading the result alongside asymmetry noise thresholds keeps a single ambiguous number from driving a clearance decision on its own.
Case: Two Testers, One Knee, Two Verdicts
A Division I athletic training staff we worked with ran into this almost by accident. Two staff members - a 145-pound graduate assistant and a 210-pound former offensive lineman now on staff - independently tested the same post-ACLR soccer player's knee extension strength within the same week, both using an unfixed break test.
The graduate assistant recorded an LSI of 83%. The former lineman, testing three days later, recorded 94% on the same knee. The player hadn't trained differently in three days, and nothing about the graft had changed. What differed was which tester could physically overpower the athlete's quad contraction - the smaller tester's break point arrived earlier on the involved limb specifically, since the uninvolved limb was strong enough that neither tester could fully overpower it, which masked most of the discrepancy on that side.
The staff re-tested both limbs with a belt-stabilized setup, three trials per limb, averaging the closest two.
| Tester | Method | Involved Limb (N) | Uninvolved Limb (N) | LSI |
|---|---|---|---|---|
| Grad assistant (145 lb) | Unfixed break test | 287 | 346 | 83% |
| Former lineman (210 lb) | Unfixed break test | 324 | 345 | 94% |
| Grad assistant (145 lb) | Belt-fixed | 318 | 352 | 90% |
| Former lineman (210 lb) | Belt-fixed | 320 | 349 | 92% |
Fixation brought both testers within two percentage points of each other on the same knee. The gap that had looked like a rehab problem was actually a testing-method problem the whole time. The program's return-to-sport committee now requires belt fixation for any strength testing used in a clearance decision, regardless of who's running the session that day.
Frequently asked questions
01Does a stronger tester always produce a more accurate reading?+
02Should clinics stop using break tests entirely?+
03How much can fixation change a single knee extension score?+
04What ICC should I expect from handheld dynamometry without any fixation?+
05Is grip strength testing affected by this same bias?+
Related Articles
Grip Strength Asymmetry and Injury Monitoring: Testing Protocol, Thresholds, and Correction
A 10%+ grip gap between hands can predate elbow and wrist injuries by weeks. See the dynamometer protocol, threshold table, and retest schedule.
Asymmetry Percentage: Noise vs. Real Difference in Strength Testing
A 14% strength gap can flip to 4% on next-day retest with nothing changed. See the typical error formula and 3-session protocol that separate noise from real.
Isometric Mid-Thigh Pull (IMTP): Testing Protocol, Norms & Applications
How to run the isometric mid-thigh pull correctly: bar height, ramp instructions, force-time variables to track, and normative data for comparison.
Return to Sport Protocol After Injury
Cleared to return does not mean ready to perform. This protocol covers a 3-stage framework, clearance criteria, load progression, and asymmetry tests.
Limb Symmetry Index Cutoffs After Injury: Why a 90% Pass Can Still Mean the Athlete Isn't Ready
A 90% limb symmetry index can pass even when the healthy leg got weaker too. See the absolute-value check, real cutoff table, and testing protocol.
IMU Jump Height Accuracy vs Force Plate: Research Review
Can wearable IMU sensors replace a force plate for measuring jump height? This review breaks down validity and reliability data across lab and field.
Mean vs Peak Velocity in VBT: Which Metric Should You Use?
Mean velocity, mean propulsive velocity, or peak velocity — they do not track load the same way. This review shows which metric fits which exercise.
Curved Sprint Asymmetry: Left vs Right-Turn Gaps as an Injury Flag
A hamstring re-tears three weeks after clearing straight-line sprint tests -- it happened on a curve. Here's how to test left vs right bend sprint gaps first.
Measure performance with lab-grade accuracy