How to Do Muscle Testing: A Step-by-Step Guide

Manual muscle testing is a hands-on clinical technique in which an examiner applies force to a specific body part while the person being tested resists, allowing the examiner to gauge muscle strength on a standardized scale. The procedure has been a staple of physical therapy, neurology, and orthopedic assessment for over a century. But “muscle testing” also refers to something quite different in alternative health circles, where practitioners use perceived changes in muscle resistance to diagnose food sensitivities, emotional blockages, or nutritional deficiencies. These two practices share a name and a superficial resemblance, but the evidence behind them diverges sharply.

Two Very Different Practices Under One Name

The confusion starts with terminology. Standard manual muscle testing (MMT) is a physical examination technique. A clinician isolates a single muscle or muscle group, places it in a specific position, and applies resistance to determine how strong it is. The goal is straightforward: find out whether a muscle is working at full capacity, partially weakened, or paralyzed. Neurologists, physical therapists, and orthopedic surgeons use this information to track recovery from injuries, monitor neuromuscular diseases, and plan rehabilitation programs.

Applied Kinesiology (AK), developed in the 1960s by a chiropractor named George Goodheart, borrowed the format of MMT but applied it to entirely different questions. In AK and its offshoots, a practitioner might press down on your outstretched arm while you hold a food item, and interpret any perceived weakness as a sign of intolerance. The assumption is that the body’s muscular response reflects its overall state of health, including things like organ function, allergies, or emotional stress. A critical review of the literature found that when AK’s unique diagnostic procedures are separated from standard orthopedic muscle testing, the studies either refute or fail to support the validity of those procedures as diagnostic tests.1PubMed Central. Disentangling manual muscle testing and Applied Kinesiology: critique and reinterpretation of a literature review The rest of this article covers both practices, but treats them as the distinct methods they are.

How Standard Manual Muscle Testing Works

Clinical MMT follows a structured sequence. Though specifics vary by muscle group, the overall process is the same whether you are testing a shoulder, hip, or ankle. Here is the general procedure a clinician follows:

  • Isolate the muscle: The person being tested is positioned so that the target muscle or muscle group is the primary mover for the action. Other muscles that could compensate are minimized through careful body positioning. For example, testing the deltoid for shoulder abduction means stabilizing the trunk so the person cannot lean to generate force.
  • Choose the test position: The joint is placed at a specific angle. Research using anatomical models of muscle force-length properties has identified optimal joint positions for major muscle groups of the shoulder, elbow, wrist, hip, knee, and ankle, where net muscle force production is predicted to be highest.2Journal of Sport Rehabilitation. Optimal joint positions for manual isometric muscle testing Testing at these positions gives the clearest picture of a muscle’s true capacity.
  • Apply resistance: The examiner pushes against the limb while the person tries to hold their position or push back. The examiner uses the response to assign a strength grade.
  • Grade the result: Muscle strength is recorded on a scale, usually from 0 (no detectable contraction) to 5 (normal strength against full resistance).

The entire test for a single muscle group takes only a few seconds of active effort. A full examination might cover dozens of muscles bilaterally, comparing the left side to the right for asymmetries that can point to nerve damage, muscle disease, or incomplete recovery from surgery.

Make Tests and Break Tests

There are two main ways to apply resistance during a manual muscle test, and the choice between them affects both the numbers you get and how reliable those numbers are.

In a “make” test, the body part starts at the beginning of its range of motion, and the person pushes as hard as they can against the examiner’s fixed hand. The examiner does not try to overpower the person; they simply hold steady and measure how much force the person generates. In a “break” test, the person first moves the body part to the end of its range, and the examiner then gradually increases pressure until the joint gives way and begins to move.3Europe PMC. A comparison of the reliability of make versus break testing in measuring palmar abduction strength of the thumb

Break tests consistently produce higher force readings than make tests, because the examiner is essentially finding the point of failure rather than measuring voluntary effort. However, research comparing the two approaches found that make tests tend to be more reliable when a hand-held device is used, meaning repeated measurements are more consistent.4PubMed Central. A comparison of make and break tests using a hand-held dynamometer and the Kin-Com For this reason, many clinicians default to make tests for tracking a patient’s progress over time, where consistency between measurements matters most. Break tests are sometimes preferred when the examiner needs to detect subtle weakness that a patient might compensate for during a make test.

The 0 to 5 Grading Scale

Most clinical settings use a grading system descended from the one developed in the early twentieth century for evaluating polio patients.5BioMed Central. On the reliability and validity of manual muscle testing: a literature review The scale runs from 0 to 5:

  • Grade 0: No visible or palpable contraction at all.
  • Grade 1: A flicker of contraction is visible or can be felt, but the muscle cannot move the joint.
  • Grade 2: The muscle can move the joint through its full range, but only when gravity is eliminated (for example, sliding the limb sideways on a table rather than lifting it upward).
  • Grade 3: The muscle can move the joint through its full range against gravity, but cannot resist any additional force from the examiner.
  • Grade 4: The muscle can move through full range against gravity and can resist some additional force, but not full resistance.
  • Grade 5: Normal strength. The muscle resists the examiner’s full manual resistance without giving way.

Plus and minus modifiers are sometimes added (e.g., 4+ or 3−) for finer distinctions, though this expansion of the scale introduces more subjectivity and is used inconsistently across clinicians. The scale is most reliable at the extremes. Telling the difference between a grade 0 and a grade 3 is straightforward; distinguishing a grade 4 from a grade 4+ requires judgment calls that different examiners may make differently.

Where Manual Muscle Testing Is Reliable and Where It Struggles

The reliability of MMT depends heavily on what you are testing and how sick the person is. A study of people with inflammatory muscle disease found that the MMT8, a standardized battery that scores eight muscle groups, produced excellent agreement both when the same examiner tested someone twice and when different examiners tested the same person. However, reliability dropped for individual muscle groups tested in isolation, with some like wrist and knee extension showing only moderate consistency, and ankle extension performing poorly.6PLOS ONE. Manual muscle testing and hand-held dynamometry in people with inflammatory myopathy: An intra- and interrater reliability and validity study

This pattern makes clinical sense. Larger, more accessible muscle groups like the hip abductors and shoulder abductors are easier to isolate and stabilize, meaning different examiners are more likely to agree. Smaller, more distal muscles (those farther from the trunk) are harder to isolate cleanly, and slight variations in hand placement or stabilization can change the result. The composite score approach, testing many muscles and adding up the grades, smooths out these inconsistencies, which is why clinical trials of drugs for muscle diseases almost always use a composite score rather than relying on a single muscle group.

For strong individuals, the grading scale also runs into a ceiling problem. If someone’s quadriceps are strong enough that no examiner could break the contraction, the muscle earns a grade 5 regardless of whether it is at 80% or 100% of its true capacity. This is where instrumented measurement becomes useful.

Hand-Held Dynamometry and When You Need Numbers

A hand-held dynamometer is a small device the examiner holds between their hand and the person’s limb during a muscle test. It records the peak force in newtons or pounds, replacing the subjective 0-to-5 grade with an objective number. Research in patients with knee osteoarthritis found that dynamometry is less subjective than MMT, particularly at the stronger end of the scale where manual grading struggles to differentiate.7Journal of Orthopaedic & Sports Physical Therapy. Reliability of hand-held dynamometry and its relationship with manual muscle testing in patients with osteoarthritis in the knee For documenting a patient’s progress over weeks or months of rehabilitation, dynamometry provides the kind of quantitative data that can show whether a treatment is actually working or whether a decline has started.

Dynamometry has its own limitations. The examiner still needs to stabilize the device consistently, and if the patient is much stronger than the examiner, the device simply slides. Expensive isokinetic machines (large, motorized devices that control speed while measuring force) solve this problem but are impractical outside of research labs and specialized clinics. In routine clinical work, the combination of manual grading for weaker muscles and dynamometry for stronger ones covers most needs.

How to Test Yourself at Home

Clinicians spend years learning to standardize their hand placement, stabilization, and grading. Self-testing at home will never replicate that precision, but there are situations where a rough gauge of your own strength is useful, particularly if you are rehabbing an injury and want to track progress between appointments.

The simplest approach is a functional comparison between sides. If your right knee was injured, sit on the edge of a table and straighten each leg in turn, noting how easy or difficult the movement feels. Can you hold the leg straight while someone gently pushes down on it? Do both legs feel equally strong, or does the injured side give out sooner? These are crude versions of the clinical test, and they will not catch a 10% difference in strength, but they can flag gross asymmetries worth mentioning to your therapist.

Gravity-based testing is another option for muscles that are significantly weak. If you cannot lift your arm above shoulder height against gravity, that is roughly a grade 2, meaning the muscle can move the joint only when gravity is taken out of the equation. If you can lift the arm but cannot hold it there against any downward pressure, that is roughly a grade 3. These observations give you a vocabulary to describe your symptoms accurately when you see a clinician.

What you should not do is use home muscle testing to diagnose a condition. Weakness in a specific distribution, say the muscles controlled by a particular nerve, means something very different from general fatigue. Interpreting the pattern of weakness requires training that a step-by-step guide cannot replace.

Applied Kinesiology and “Muscle Response Testing”

The version of muscle testing most people encounter on social media looks nothing like the clinical procedure described above. In these demonstrations, a practitioner pushes down on a person’s outstretched arm while the person holds a supplement, touches a body part, or thinks about a stressful memory. If the arm “goes weak,” the interpretation is that the body is reacting negatively to whatever stimulus was introduced. This is the domain of Applied Kinesiology and its various offshoots, sometimes called muscle response testing (MRT) or energy testing.

The theoretical premise is that the nervous system is so finely tuned that it will reflexively weaken a muscle in the presence of a harmful substance or thought. Proponents argue this creates a real-time feedback system that can diagnose food allergies, detect nutritional deficiencies, and guide supplement choices. One small study did find that AK screening procedures matched serum immunoglobulin tests for food allergies in about 90% of cases.8PubMed Central. Correlation of applied kinesiology muscle testing findings with serum immunoglobulin levels for food allergies That finding sounds impressive in isolation, but it came from a study with only 21 suspected allergies, and the broader body of evidence does not support the premise.

A more recent study designed to test whether MRT practitioners could accurately distinguish between true and false stimuli found that their mean overall accuracy was about 62%, compared to roughly 50% for people simply making intuitive guesses, which is essentially coin-flip performance.9PLoS One. Exploring the variation in muscle response testing accuracy through repeatability and reproducibility The practitioners did slightly better than chance, but not by a margin that would make the technique useful for clinical decisions. A diagnostic test with 62% accuracy means it gives the wrong answer nearly two out of every five times.

Why AK-Style Testing Feels Convincing

If the evidence is this weak, why does muscle response testing seem to work so convincingly during a demonstration? Several well-understood factors explain the experience without requiring the theoretical framework AK proposes.

The most significant factor is subtle variation in how the examiner applies force. Even a slight change in the angle, speed, or amount of pressure will make a muscle appear to give way. The examiner is usually also the person interpreting the result, and they often know what the “expected” outcome should be. This is not necessarily conscious deception. Unconscious shifts in force application are nearly impossible to control without instrumented measurement, which is precisely why clinical researchers use dynamometers.

There is also the role of expectation on the part of the person being tested. If you believe a certain food is bad for you, you may unconsciously brace less firmly when that food is introduced. Similarly, a confident practitioner who says “this one should be strong” before pressing can prime you to resist harder. These psychological influences do not mean the muscle response is fake; they mean the response is driven by the testing context rather than by the substance or thought being evaluated.

The critical review that separated AK’s claims from standard orthopedic MMT noted that the original literature review on muscle testing had conflated the two, implying that the established reliability of clinical MMT lent credibility to AK’s unique diagnostic uses.1PubMed Central. Disentangling manual muscle testing and Applied Kinesiology: critique and reinterpretation of a literature review That conflation persists in popular understanding and is one reason AK retains a veneer of scientific respectability.

Getting the Most Accurate Clinical Results

If you are being tested by a physical therapist, neurologist, or other clinician, a few practical considerations influence how accurate the results will be.

Fatigue matters. If the examiner tests the same muscle repeatedly, or if you have just finished an exercise session, the muscle will perform below its actual capacity. Testing should happen when you are reasonably rested and the muscle has not been pre-fatigued by earlier tests in the same session. For this reason, many protocols test the weakest muscles first.

Pain changes everything. If contracting a muscle hurts, you will instinctively limit your effort whether you mean to or not. A clinician who recognizes pain inhibition will note it rather than recording the result as true weakness. The distinction between “weak because the nerve is damaged” and “weak because it hurts to push hard” is clinically important, and a skilled examiner watches your face and body language as carefully as they feel the muscle contraction.

Consistency of positioning is the single biggest factor in whether results are comparable over time. If your shoulder was tested at 90 degrees of abduction last visit and 70 degrees this visit, the force values are not directly comparable. Standardized testing protocols specify exact joint angles for this reason, and the research on optimal joint positions for isometric testing exists precisely to reduce this source of variability.2Journal of Sport Rehabilitation. Optimal joint positions for manual isometric muscle testing

Muscle Testing in Specific Conditions

MMT plays different roles depending on the clinical context. In neuromuscular diseases like inflammatory myopathies, dermatomyositis, or muscular dystrophy, serial muscle testing is one of the primary ways clinicians track disease progression and treatment response. The MMT8 composite score, which tests neck flexion, deltoid, biceps, wrist extensors, gluteus maximus and medius, quadriceps, and ankle dorsiflexors, has become a standard outcome measure in clinical trials for these conditions.6PLOS ONE. Manual muscle testing and hand-held dynamometry in people with inflammatory myopathy: An intra- and interrater reliability and validity study

In nerve injuries, the pattern of weakness is often more informative than the degree. A radial nerve injury produces a characteristic set of weak muscles (wrist extensors, finger extensors, triceps) while leaving others intact. A clinician who systematically tests muscles innervated by different nerves can often localize where the damage occurred before any imaging is done. This is one of the oldest and most elegant applications of MMT, dating back to its origins in evaluating polio patients whose anterior horn cells had been destroyed in specific spinal cord segments.5BioMed Central. On the reliability and validity of manual muscle testing: a literature review

In orthopedic rehabilitation after surgery or injury, MMT is used more as a progress marker. The question is less “which nerve is damaged” and more “is the repaired tendon strong enough to return to sport.” Here, dynamometry often supplements the manual grade, because the clinical question demands finer resolution than a five-point scale can provide.

Common Mistakes That Compromise Results

Whether you are a student learning to perform MMT or a patient trying to understand why your results seem inconsistent, a few errors come up repeatedly.

Substitution is the most common. When a target muscle is weak, neighboring muscles try to take over the movement. A patient with a weak gluteus medius will hike their hip to abduct the leg using their trunk muscles instead. If the examiner does not stabilize the pelvis, the test will look stronger than it actually is. Skilled clinicians actively block substitution by holding the body part adjacent to the one being tested.

Inconsistent verbal instructions also skew results. Telling someone “push as hard as you can” produces different effort levels than “resist my pressure.” The phrasing, tone, and encouragement the examiner uses should be standardized between sessions, particularly in research settings.

Testing too fast is another pitfall, especially with break tests. If the examiner ramps up pressure abruptly, the muscle may yield because the person was not ready, not because the muscle is weak. A smooth, gradual application of force over about two to three seconds gives the nervous system time to fully recruit the muscle before the break point is reached.

Finally, the examiner’s own strength matters. A small examiner testing a large patient’s quadriceps may simply not be strong enough to break the contraction, producing a grade 5 that reflects the examiner’s limitation rather than the patient’s true strength. This is one of the practical situations where a dynamometer strapped to a fixed surface becomes essential rather than optional.