Fitness is not one quantity, and the tests that claim to measure it probe different physiological systems. Disagreement between two assessments is usually a category difference rather than an error.

Aerobic capacity is an oxygen measurement

The reference measure of endurance is the maximum rate at which the body can take in, transport and use oxygen during hard exercise, determined by heart output and muscle metabolism.

Measuring it directly requires a mask, analyzed expired air and a graded test to exhaustion, usually in a lab or a clinic. Everything else is an estimate.

Field versions substitute a known workload and a heart rate response, then infer capacity. The inference holds reasonably across groups and loosely for any single person.

Wearables estimate rather than measure

A watch reporting a fitness score is modeling the relationship between pace, heart rate and known population patterns. It never observes oxygen use at all.

That model is sensitive to conditions the device cannot see: heat, hills, sleep loss, illness, caffeine and how tightly the strap sits on the wrist.

The number is therefore most useful as a trend within one person over weeks, and least useful as a comparison against a friend wearing a different brand.

Strength tests measure a skill as well as a tissue

A one-repetition maximum reflects muscle force, but also technique, familiarity with the movement and the nervous system's willingness to recruit fully under load.

Early gains on such a test often come from coordination rather than added muscle, which is why a novice can improve quickly without visible physical change.

Grip strength is used differently. It correlates with overall muscular condition and is easy to measure repeatedly, so it functions as a cheap general indicator rather than a training target.

Clinical exercise testing asks a different question

A cardiologist's treadmill test is not scoring fitness. It observes how the heart's electrical activity, rhythm and blood pressure respond to increasing demand under supervision.

The endpoint is diagnostic information, and the test is stopped on clinical signs rather than on exhaustion. The setting includes monitoring and staff for that reason.

This is why a gym assessment cannot substitute for one. Symptoms during exertion, such as chest discomfort, unusual breathlessness or faintness, belong with a physician rather than a trainer.

Why the results diverge

Each test is specific to what it loads. Someone with a strong cycling background may score poorly on a running protocol simply because the movement pattern is unfamiliar.

Day-to-day variation is substantial. Hydration, ambient temperature, recent training and sleep all move results enough to swamp real change measured over a short window.

Repeating one test under consistent conditions therefore tells you more than collecting several different ones, because the comparison you can trust is against yourself.