IQTest EN

Method

How an IQ score is calculated

There are four steps between the answers you give and the number on the report, and the one most people have never heard of is the one that matters most.

Updated September 2026 9 min read Reviewed by the IQTest EN editorial team

Step one: the raw score

The raw score is simply how many items you got right, sometimes with weighting for difficulty and sometimes with a bonus for speed on timed subtests.

On its own it means nothing whatsoever. Twenty-eight correct out of forty is not interpretable until you know what other people scored, which is what the next step supplies.

Step two: comparison against an age group

Your raw score is compared against the distribution of raw scores produced by the norm sample for your age band.

This is where most of the real work in test construction lives. A norm sample for a major clinical instrument runs to a couple of thousand people, stratified to match census data on age, sex, education, region and ethnicity, tested under standard conditions. Building one is expensive, which is why editions are revised roughly once a decade rather than continuously.

Age banding is why a 25-year-old and a 65-year-old with identical raw performance can receive identical IQ scores. Raw performance on speeded and fluid-reasoning tasks declines with age; the norms compensate, because the score is designed to express your standing among your peers rather than your absolute capability. This surprises people, and it is the source of a great deal of confusion about whether IQ declines with age. The longer answer.

The step that decides everything

A score is a rank against a norm group. Change the group and the same performance produces a different number. This is why a site that will not describe its norm sample has handed you a figure with no denominator.

Step three: scaled and index scores

Each subtest result is converted to a scaled score, conventionally with a mean of 10 and a standard deviation of 3. Subtests are then grouped into indices — on a modern Wechsler, typically verbal comprehension, visual–spatial, fluid reasoning, working memory and processing speed — and each index is expressed on the familiar 100/15 scale.

The full-scale IQ is a composite of the indices. It is the most reliable number on the report, because it aggregates the most items.

The profile across indices is often clinically more interesting than the total, and it is also where over-interpretation begins. Index scores are shorter and therefore noisier, with confidence intervals often eight to ten points wide. A twelve-point gap between two indices looks like a meaningful strength-and-weakness pattern and frequently is not.

Step four: the confidence interval

This is the step that online tests almost universally skip, and it is the one that makes the rest honest.

Every score is an estimate. Test theory treats an observed score as a true score plus error, and quantifies the error as the standard error of measurement. For a full-scale IQ on a modern clinical battery it runs to about two to three points, producing a 95 percent confidence interval of roughly plus or minus five.

So a properly reported clinical result does not read “IQ 118”. It reads “IQ 118, 95% CI 113–123, 88th percentile”. The interval is not a disclaimer. It is part of the result.

Why the band is wider than you think.

What happened to the mental-age formula

The original calculation, proposed by William Stern in 1912, was mental age divided by chronological age, multiplied by 100. A ten-year-old performing like a twelve-year-old scored 120.

It has two problems. Mental age plateaus in adolescence while chronological age does not, so the ratio makes every adult appear to decline steadily from around twenty. And the same ratio produces different rarities at different ages, because the spread of ability is not proportional to age.

Wechsler replaced it with the deviation IQ in 1939 and every modern test followed. The mental-age ratio survives only in popular writing, and it is usually the source of the enormous childhood IQ figures that circulate online — a five-year-old with a mental age of ten produces a ratio score of 200, which does not mean what it appears to mean.

How online tests do it

The first three steps are the same in principle. The fourth is usually omitted, and the second is where the real difference lies.

A web test has no representative norm sample. What it has, at best, is its own accumulated visitors — a group that is younger, more online, more interested in test scores and more likely to have taken several other tests than the population is. Percentiles calculated against that group are not percentiles against everyone, and a site that presents them as though they were is making a claim it cannot support.

Some sites take the more defensible route of mapping raw scores onto the theoretical normal curve rather than onto their own traffic. That is what we do, and we say so on the methodology page, because it is a real limitation rather than a detail.

Reviewed by the IQTest EN editorial team Every reasoning item on this site is solved independently by two reviewers before it is published, and every number in a reference table is traced back to a named source. Where the research is contested, we say so on the page instead of picking the tidier answer. How we write and check this.

Questions

Questions people ask

How is an IQ score calculated?

Raw answers are compared against the distribution for your age band in a norm sample, converted to scaled subtest scores, aggregated into index scores and a full-scale score on a 100/15 scale, and reported with a confidence interval.

Is IQ still mental age divided by chronological age?

No. That ratio was abandoned in the 1930s because it breaks down for adults. Every modern test uses a deviation score: your position within your own age group.

Why do older people get the same IQ with lower raw scores?

Because norms are age-banded. The score expresses your standing among people your own age, not your absolute performance, and raw performance on speeded and fluid tasks does decline with age.

What is a scaled score?

A subtest result converted to a common scale with a mean of 10 and a standard deviation of 3, so that results from subtests of different lengths and difficulties can be combined.

Should an IQ report include a confidence interval?

Yes. Standard clinical practice is to report the score with its 95 percent confidence interval, which for a full-scale score is about plus or minus five points.

Find out where you land

Ten reasoning questions, an instant band and an explanation for every answer.