🩸 Metabolic Health · 11 min read · Subtopic 1 of 5

DEXA vs Bioimpedance vs Calipers

Three devices all claim to answer the same question — how much of you is fat — and they can disagree with each other by eight percentage points for the same person on the same day. This page explains what each method physically measures, how big its error bars really are, and how to get the cheapest number that still tells the truth.

🔎 Evidence Snapshot ★★★★☆ Good — large validation literature; absolute accuracy varies by device and operator

What the evidence supports

  • DXA (dual-energy X-ray absorptiometry) is well validated against the four-compartment model, with repeat-scan error around 1–2 percentage points of body fat (Shepherd et al., Bone, 2017).
  • Bioelectrical impedance tracks total body water well enough to follow trends, but absolute body-fat estimates carry roughly ±3–8% error against reference methods (Sun et al., Am J Clin Nutr, 2003).
  • Skinfold equations have been validated in large samples since the 1970s, with accuracy concentrated in trained hands (Durnin & Womersley, Br J Nutr, 1974).

What remains uncertain

  • No home device matches the precision of lab methods; consumer "muscle quality" and "metabolic age" readouts are mostly proprietary math on top of impedance.
  • Individual error can exceed the population average — the same smart scale may misread one person by double digits while tracking another reliably.

Evidence last reviewed: August 15, 2026. Conclusions may change as new research is published.

two beams, one weak current

1–2%
repeat-scan error of DXA for body-fat percentage
±3–8%
typical error of consumer bioimpedance against lab references
4
compartments in the reference model: fat, water, protein, mineral

What Each Method Physically Measures

The three tools are not three versions of the same measurement. They are three different measurements that all get converted — through equations built on other people's bodies — into the same familiar number. DXA passes two low-energy X-ray beams through you and reads how much of each beam is absorbed; because fat, bone, and lean tissue attenuate the beams differently, software can estimate bone mineral, fat, and lean mass for the whole body and for each region. A scan takes roughly ten to twenty minutes and uses a radiation dose far below a day's natural background exposure. It is the field test the other methods are judged against.

The shared limitation: every method measures a proxy and runs it through a model built on other people. None of them "sees" your fat directly.

The Error Bars, Honestly

What does "±3%" mean in practice? If a scale reads 25% with ±3% error, the lab value is plausibly anywhere from 22% to 28% — and for some individuals the miss is larger than the average. In validation studies, consumer bioimpedance devices show errors of roughly 3–8 percentage points against multicomponent reference methods (Sun et al., Am J Clin Nutr, 2003; NIH Technology Assessment, 1996). DXA repeats within 1–2 points (Shepherd et al., Bone, 2017), and calipers in trained hands land around 3–4 points (Durnin & Womersley, Br J Nutr, 1974). The chart below shows the typical bands side by side.

Typical Error by Method
Approximate error ranges in body-fat percentage against lab reference methods, from validation studies (NIH 1996; Sun et al. 2003; Shepherd et al. 2017). Ranges are typical, not universal — individual devices and operators vary.
Error vs lab reference (± % body fat), 0–10 scale Bioimpedance scales ±3–8% Calipers, trained hands ±3–4% DXA scan ±1–2% 0% 10%

Two practical consequences follow. First, a single reading is a guess with a range, not a verdict — a difference under three points between two measurements may be pure noise. Second, error is often not random for you personally: if your scale reads four points high, it tends to read high every morning. That consistency is exactly what makes trends usable even when absolute numbers are not.

Where Each Method Goes Wrong

How We Know Any of It Works

The validation chain runs from a lab method called the four-compartment model — which measures body density, total body water, and bone mineral separately and is treated as the reference — down to the field devices. DXA has been validated against it and serves as the reference-standard field test (Shepherd et al., Bone, 2017; Toomey et al., Topics in Clinical Nutrition, 2015). The bioimpedance equations inside most consumer devices were derived by fitting impedance data against multicomponent measurements in roughly 1,800 adults (Sun et al., Am J Clin Nutr, 2003), and a 1996 NIH technology assessment spelled out the conditions under which impedance is trustworthy: hydrated, rested, consistent. Skinfold equations were likewise built against density measurements in 481 men and women (Durnin & Womersley, Br J Nutr, 1974). The lesson: every field number is a shadow cast by a better lab number, and the shadow is sharpest when your measuring conditions match the ones used to build the equation.

Reading Your Number

Body-fat percentage is simply the share of your mass that is fat — nothing more. It is not a health grade by itself; a marathoner and a frail older adult can carry identical percentages. The number earns meaning three ways: against the healthy ranges for your age and sex, as a trend over months measured under identical conditions, and beside its partners — waist circumference and strength. Two rules keep readings honest. Never mix devices: a DXA result and a smart-scale result are different currencies, and comparing them is how people conclude they lost 5% body fat in a weekend. And never interpret a single reading: the actionable unit of body-composition data is the three-to-six-month slope, which is why the quarterly audit's body-metrics routine spreads measurements across the year rather than the week.

⚠️ When the number becomes the problem

Body-composition numbers are tools, not grades. If tracking produces dread, daily weigh-ins, or restrictive eating, step back — and if you have a history of disordered eating, consider skipping the numbers entirely and working with a clinician on non-scale goals. Population ranges describe populations; your targets are a clinical conversation, not a chart reading.

A Practical Stack

For most people the value-per-dollar ordering is clear: tape monthly, DXA annually if convenient, smart scale as a daily trend only. The table is the whole decision, honestly graded.

MethodWhat it measuresTypical errorVerdict for most people
📏 Tape measureWaist and waist-to-height — the fat distribution that predicts risk±1–2 cm if consistentStrong value — monthly, free
🔬 DXABone, lean, and fat mass by body region±1–2% body fatStrong — annually if convenient
⚡ Bioimpedance scaleImpedance → estimated water, lean, and fat±3–8% body fatVariable — trends only
🫰 Skinfold calipersSubcutaneous fat at 4–7 pinch sites±3–4% body fat, trained handsModerate — operator-dependent
📱 "Muscle quality" app scoresProprietary math on top of impedanceUnvalidatedWeak — entertainment

When It's Worth Paying

A sensible budget tiers by what you actually get. The tape measure costs a few dollars and gives you the trend that matters most. A decent impedance scale runs in the tens of dollars, and its honest value is the daily trend, not the percentage. Calipers are cheap but worthless without a trained operator — a few sessions with a coach who uses them well beats owning a pair you pinch badly. DXA is the premium tier: typically tens to low hundreds of dollars per scan depending on where you live, and worth it one to two times a year if you want a trustworthy absolute number to anchor the trend — especially around a deliberate body-composition phase (the recomp, cut, or bulk decision is where a real measurement earns its keep). Skip everything that claims to grade your "metabolic age" or "muscle quality" from a bathroom scale: those numbers are not validated against anything, and they add noise where you need signal.

Questions, Answered Briefly

The Bottom Line

  1. DXA is the reference-standard field test — roughly 1–2% repeat error; use it once or twice a year to anchor a trustworthy trend.
  2. Smart scales are trend machines, not lab results — ±3–8% error driven mostly by hydration; ignore the absolute percentage.
  3. Calipers are only as good as the hand on them — excellent with trained operators, noisy for everyone else.
  4. The tape measure remains the best value — free, repeatable, and it tracks the fat distribution that actually predicts risk.

Related Topics

Sources & further reading