Calcylator
T Test Statistic

The t statistic:
how far a sample mean sits from a claim

See how the t statistic scales a gap by its standard error, why degrees of freedom matter and what t = 2.0 does and does not prove.

Calcylator Editorial Team

Updated · 5 min read

What the t statistic actually measures

Suppose a supplier says their tea packets average 50 g, but your sample of 25 packets averages 54. Is that gap real, or the kind of wobble you would expect from picking 25 packets at random? The t statistic puts a number on that question by comparing the gap with the amount of wobble the sample itself suggests.

The calculation turns the gap into units of standard error. A t of 0 means the sample mean is exactly the claimed value. A t of 2 means it is two standard errors away. The farther from zero the t, the harder it is to blame luck.

The same idea sits behind many tests: a difference, divided by a measure of how much that difference would vary by chance. The one-sample t test is the simplest version, and it is the one this page works through.

Read the sign as well as the size. A positive t says the sample mean is above the claimed value and a negative t says it is below. In a two-sided test the sign does not affect significance, but it tells you which way the evidence points, and that is often the part a reader cares about.

The formula and its parts

One-sample t statistic =t = (x̄ − μ₀) ÷ (s ÷ √n)
x̄:
Sample mean
μ₀:
The value claimed or hypothesised for the population mean
s:
Sample standard deviation
n:
Number of observations
Degrees of freedom are n − 1.

The denominator, s ÷ √n, is the standard error of the mean. It shrinks as the sample gets bigger, which is why the same gap is more convincing in a large sample than a small one.

Use the sample standard deviation with n − 1 in its own calculation, as spreadsheets and statistical software do. If you were given the true population standard deviation, a z test would be the right tool instead.

Worked example: mean 54 against a claim of 50

  • Sample mean

    54

  • Claimed mean

    50

  • Sample SD

    10

  • Sample size

    25

  • Standard error

    10 ÷ √25 = 10 ÷ 5 = 2

  • Gap

    54 − 50 = 4

t statistic

t = 4 ÷ 2 = 2.0, with 24 degrees of freedom

A two-tailed p-value for t = 2.0 on 24 df is about 0.057.

The statistic alone is not the verdict. To decide, compare it with a critical value for 24 degrees of freedom, or convert it into a p-value using a table or software. With a two-sided test at the 5 per cent level, the critical value is 2.064, so t = 2.0 just misses it. With a one-sided test at 5 per cent the critical value is 1.711, and the result would count as significant.

That sensitivity to the choice of test is a reason to decide the hypothesis and the tails before looking at the data, not after.

It also helps to see what changes if only the sample size does. Keep the mean at 54, the claim at 50 and the SD at 10, but imagine 100 observations. The standard error falls to 10 ÷ 10 = 1, and t becomes 4.0, which is far past any usual critical value. The gap of 4 units is identical in both cases. What changed is how precisely the mean was pinned down.

Degrees of freedom and critical values

The t distribution looks like a normal curve with heavier tails, and the difference is largest in small samples. As n grows, the distribution approaches the normal curve, and the critical value approaches 1.96 for a 5 per cent two-sided test.

Selected critical values from a standard t table
Degrees of freedomTwo-sided 5% critical tOne-sided 5% critical t
33.1822.353
102.2281.812
242.0641.711
602.0001.671
Very large1.9601.645

The pattern shows why small samples demand a bigger gap. With only 4 observations, a t of 3.2 is needed to clear the same bar that 1.96 clears in a huge sample.

What to report next to the t value

A t statistic and a p-value say whether a gap is distinguishable from noise. They do not say how large the gap is, or whether it matters. Report an estimate and an interval alongside them.

  • Sample mean

    54

  • Standard error

    2

  • Critical t (24 df, 95%)

    2.064

  • Margin

    2.064 × 2 = 4.128

95% confidence interval

49.87 to 58.13

The interval includes 50, which matches the borderline test result.

A standardised effect size such as the gap divided by the SD, here 4 ÷ 10 = 0.4, also helps readers see that the difference is moderate in size. A larger sample could make a modest gap highly significant while leaving the practical meaning unchanged.

Reading the result in plain words

A good write-up avoids both overclaiming and mumbling. A sentence such as "The sample mean of 54 was above the claimed 50 (t(24) = 2.00, two-sided p ≈ 0.057, 95% CI 49.9 to 58.1)" carries the numbers a reader needs. It does not say the claim is false, and it does not say the sample proves it true.

Think about the decision behind the test. If the cost of wrongly rejecting a claim is high, a stricter threshold such as 1 per cent makes sense. If it is a cheap early look and you plan to follow up, 10 per cent may be fine. The data cannot choose that threshold for you, which is why the level should be agreed first.

Finally, consider the context. A mean of 54 against a claimed 50 may be important in a medicine dose and trivial in the weight of a bag of sugar. Statistical significance and practical importance are different questions, and the t statistic answers only the first. For anything with health or safety consequences, the interpretation belongs with a qualified specialist who knows the field.

Assumptions, and where the test goes wrong

  • Independence: observations must not influence each other. Repeated measurements on the same person break this.
  • Roughly normal data, or a reasonably large sample. For very small, strongly skewed samples a rank-based test may be safer.
  • No wild outliers: one extreme value changes both the mean and the SD.
  • One question: running the test on many outcomes and reporting the one that passes inflates the chance of a false positive.
  • Not the same as proof: a non-significant result means the data could not separate the claim from chance, not that the claim is true.

If you are comparing two groups rather than one group against a value, the structure is similar but uses the difference between the means and a combined standard error. Paired data, such as before and after on the same people, reduce to a one-sample test on the differences.

Common questions

What is the formula for the one-sample t statistic?

It is t = (x̄ − μ₀) ÷ (s ÷ √n): the sample mean minus the hypothesised mean, divided by the sample standard deviation over the square root of the sample size. Degrees of freedom are n − 1.

How many degrees of freedom does a one-sample t test have?

It has n − 1. A sample of 25 observations gives 24 degrees of freedom. One degree is lost because the sample mean is estimated from the same data, and the distribution is read at that row of the table.

Is t = 2.0 significant?

It depends on the degrees of freedom and the test type. With 24 df, the two-sided 5% critical value is 2.064, so t = 2.0 narrowly misses. A one-sided test at 5% uses 1.711, which it would pass.

What is the difference between a t statistic and a p-value?

The t statistic is the standardised gap itself. The p-value is the probability of seeing a gap at least that large if the claimed mean were true. You obtain the p-value by looking up t on a t distribution with the right degrees of freedom.

When should I use a t test instead of a z test?

Use a t test when the population standard deviation is unknown and you estimate it from the sample, which is the usual case. A z test applies when the population standard deviation is known, or in very large samples where the two give nearly the same answer.

Was this guide helpful?

Continue reading

View all blogs