Calcylator
Variability

Coefficient of variation:
comparing spread when the averages differ

A bigger standard deviation does not always mean a messier process. Dividing by the mean puts different scales on equal footing.

Calcylator Editorial Team

Updated · 6 min read

Why a standard deviation alone can mislead

A standard deviation of 50 grams sounds worse than one of 4 grams. Yet if the first belongs to 5 kg bags of rice and the second to 200 g snack packs, the bagging line is the more consistent of the two. The size of the wobble has to be judged against the size of the thing being measured.

The coefficient of variation, often shortened to CV, does that judging for you. It expresses the standard deviation as a share of the mean, so the result is a plain percentage with no units. That makes it possible to compare heights with weights, or monthly sales of a tea stall with those of a supermarket.

It is also called relative standard deviation, and it is the figure behind phrases like 'results varied by 8%'. Lab reports, investment summaries and quality control sheets all lean on it.

The formula and how to read it

Coefficient of variation =CV = (s ÷ x̄) × 100%
s:
standard deviation (use σ for a population)
x̄:
mean of the same data (μ for a population)

Use the same kind of standard deviation you would normally use for the data set: the sample version for a sample, the population version for a full population. Then divide by the mean and multiply by 100 to express it as a percent.

There is no universal threshold for 'good'. A CV of 2% would be sloppy for a pharmacy scale and unrealistically tight for daily footfall at a market stall. Judge it against what is normal for that kind of data and against the CV of the alternative you are comparing.

Worked comparison: which line is steadier?

A packing unit fills snack packs with a mean of 200 g and a standard deviation of 4 g. A second line fills rice bags with a mean of 5,000 g and a standard deviation of 50 g. On raw SD the rice line looks twelve times worse.

  • Snack line

    mean 200 g, SD 4 g

  • Rice line

    mean 5,000 g, SD 50 g

  • Snack CV

    4 ÷ 200 × 100 = 2%

  • Rice CV

    50 ÷ 5,000 × 100 = 1%

Steadier line

Rice bagging (1% vs 2%)

The rice line has the larger SD but half the relative variation.

Another way to say it: the typical snack pack is off by 2% of its weight, whereas the typical rice bag is off by 1% of its weight. For the customer opening the pack, the percentage is what they feel.

A second worked case from small samples

Two market stalls record sales, in thousands of rupees, over five weeks. Stall A: 40, 44, 38, 46, 42. Stall B: 120, 140, 100, 150, 90. Because these are five weeks drawn from a longer trading history, the sample standard deviation is the right one.

  • Stall A

    mean 42, SD 3.16 (sum of squares 40 ÷ 4 = 10)

  • Stall B

    mean 120, SD 25.50 (sum of squares 2,600 ÷ 4 = 650)

  • Stall A CV

    3.16 ÷ 42 × 100 = 7.5%

  • Stall B CV

    25.50 ÷ 120 × 100 = 21.2%

More predictable stall

Stall A, by about three times

B earns nearly three times as much, but swings far more week to week

The absolute SDs are 3.16 and 25.50, a gap of eight times, which exaggerates the difference. The CVs show the proportional picture: B's weekly sales move by about a fifth of their average, A's by under a tenth.

When the coefficient of variation breaks down

  • The mean is close to zero. Dividing by a tiny number blows the percentage up, so a data set of profits and losses averaging near zero gives a meaningless CV.
  • The scale has no true zero. Temperatures in °C are the classic trap: readings of 20 °C with an SD of 2 °C give a CV of 10%, but the same readings in kelvin (293.15 K) give about 0.68%. Same data, very different CV.
  • Values can be negative. A CV only makes sense for data that is always positive, such as weights, prices, durations.
  • The distribution is badly skewed. Mean and SD are both pulled by extremes, so the CV inherits the problem.

Where people use it

FieldWhat the CV comparesTypical use
ManufacturingFill weights, dimensions across different productsIs the process consistent?
LaboratoriesReplicate readings of an assayPrecision of a method
FinanceSpread of returns relative to the average returnRisk per unit of return
AgricultureYield across plots or seasonsStability of a crop variety
BusinessWeekly sales of stores of different sizeWhich store is more predictable?

In investing, a lower CV means less risk per unit of expected return. It is a descriptive comparison, not a prediction, and it relies on past data that may not repeat.

Reporting the figure sensibly

Quote the CV with the mean and the number of observations, since a CV from five readings is much less certain than one from five hundred. A line such as 'CV 7.5% (n = 5)' is more honest than a bare percentage.

  • Round to one or two significant figures; a CV of 7.529% implies a precision that five data points cannot support.
  • State whether the SD behind it is the sample or the population version.
  • Do not compare CVs from different measurement methods without noting it; a different instrument can change both mean and SD.
  • If a data set is made of ratios or percentages already, a CV on top can confuse; consider the SD of the raw quantities instead.

Within any one process, tracking the CV over time is a simple health check. A rising CV on the same product means the process is drifting, even if the mean stays on target.

A quick way to compute it

  1. Find the mean of the data.
  2. Find the standard deviation, choosing sample or population as appropriate.
  3. Divide the standard deviation by the mean.
  4. Multiply by 100 and attach a % sign.
  5. Compare it with another data set's CV, not with its raw SD.

A calculator or spreadsheet makes this immediate; the main thing to check by eye is that the mean is not close to zero, and that the units of the standard deviation and the mean match. A CV of 100% or more is a sign that the data is very spread out or has a mean too small to divide by safely.

Common questions

What is the coefficient of variation in simple words?

It is the standard deviation expressed as a percentage of the mean. A CV of 10% means the typical value sits about one-tenth of the average away from that average, which makes spread comparable across different scales.

How do you calculate the coefficient of variation?

Divide the standard deviation by the mean and multiply by 100. For a mean of 200 g and an SD of 4 g, 4 ÷ 200 = 0.02, so the CV is 2%. Use consistent units for both.

Is a high coefficient of variation bad?

Not by itself. It only says the data varies a lot relative to its average. That may be fine for naturally volatile things like daily sales, but it signals a problem for a filling machine that should be precise.

Why is the coefficient of variation unreliable near zero?

Because the mean is the denominator. When the mean is close to zero, even a small standard deviation gives an enormous percentage, and data containing negatives can have a mean near zero while being widely spread.

Can I use CV on temperatures in Celsius?

It is not recommended. Celsius has an arbitrary zero, so the CV changes if you switch to kelvin: 20 °C with SD 2 gives 10%, but 293.15 K with the same SD gives about 0.68%.

Was this guide helpful?

Continue reading

View all blogs