IQTest EN

Methods

The error band on your IQ score

Every score is an estimate with a range around it. Here is how wide that range is, where it comes from, and what follows for comparing scores.

Published 19 September 2026 7 min read By the IQTest EN editorial team
The error band on your IQ score

A score is an estimate, not a reading

A tape measure gives you a reading. An IQ test gives you an estimate, and the difference is not pedantry.

Your performance on any given day depends on the specific items you drew, how well you slept, whether the room was quiet, whether you happened to spot one rule quickly, and dozens of other things that have nothing to do with the ability being measured. Test theory treats your observed score as your true score plus error.

The size of that error is quantifiable, and it is called the standard error of measurement.

How wide the band is

The standard error of measurement depends on the test's reliability and the spread of scores. For a full-scale IQ on a modern clinical battery with a reliability around 0.97, it works out at roughly two to three points.

A 95 percent confidence interval is about twice that in each direction, so a clinical full-scale score carries an interval of roughly plus or minus five points.

Subtest and index scores are shorter and therefore less reliable, with correspondingly wider bands — often eight to ten points. This is why clinicians interpret the full-scale score with more confidence than the individual indices, and why a pattern of highs and lows across indices should not be over-read.

The practical rule

If two scores are within about six points of each other, treat them as the same score. This holds whether you are comparing two people or the same person on two occasions.

What this means for online tests

Everything above assumes a long, well-constructed test administered under controlled conditions. An unsupervised web test has a wider band for several compounding reasons.

It is shorter, and the standard error falls roughly with the square root of test length. Conditions are uncontrolled, which adds variance that a clinical setting removes. And the norm group is self-selected, which adds a systematic offset on top of the random error — a different kind of problem that a confidence interval does not capture at all.

A ten-item test can reliably place you in a quarter of the distribution. That is a genuine piece of information. What it cannot do is distinguish 112 from 118, which is why this site reports a band.

Why your score changes when you retake

People are often unsettled to score 112 one week and 121 the next, and conclude that one of the tests must be broken. Usually neither is.

Three things drive the difference. Measurement error alone accounts for a few points in either direction. Practice effects account for several more if the format is the same, concentrated in the first repeat. And state factors — sleep, caffeine, time of day, mood — account for the rest.

Clinical practice deals with this by using alternate forms and observing minimum retest intervals, often six to twelve months. Online, none of that applies, which is why a personal-best score collected across five attempts is a measurement of your persistence.

Where the band has real consequences

Confidence intervals stop being an abstraction the moment a threshold is involved.

Gifted programme entry is commonly set at 130. Intellectual disability criteria reference approximately 70, alongside adaptive functioning assessment. Mensa uses the 98th percentile.

A person whose true score sits near any of these lines will fall on either side of it depending on the day. Good clinical practice treats a threshold as a region and brings in other evidence rather than deciding on a single number. Programmes that apply a hard cut-off to a single test score are, in statistical terms, sorting partly on measurement noise.

The question to ask any test

There is a quick test for whether a site understands its own measurement: does it report an interval?

A result presented as a bare number, with no range and no discussion of error, tells you that the site either does not know about measurement error or has decided that a clean number converts better. Both are useful things to know before you decide how much weight to give the result.

Reviewed by the IQTest EN editorial team Every reasoning item on this site is solved independently by two reviewers before it is published, and every number in a reference table is traced back to a named source. Where the research is contested, we say so on the page instead of picking the tidier answer. How we write and check this.

Questions

Questions people ask

What is the standard error of measurement on an IQ test?

Roughly two to three points on a modern clinical full-scale score, producing a 95 percent confidence interval of about plus or minus five points. Index and subtest scores have wider bands.

Why did I get a different IQ score the second time?

Measurement error, practice effects and day-to-day state all contribute. Differences of a few points between sittings are expected and carry no information.

How many points of difference are meaningful?

As a working rule, more than about six points on a clinical test. Below that, the difference sits inside the measurement error.

Do clinical reports show a confidence interval?

Yes. Standard practice is to report the score alongside its confidence interval, and the interval is the part that should be quoted.

Find out where you land

Ten reasoning questions, an instant band and an explanation for every answer.