Methods
How accurate is an IQ test, really?
Reliability, validity and the standard error of measurement, explained without the jargon and with the numbers that actually apply.

Accuracy is two questions, not one
When people ask whether IQ tests are accurate, they are usually running two separate questions together, and the answers are different.
Reliability asks whether the test gives you the same answer twice. If you sat the same test on two consecutive Tuesdays, how close would the results be?
Validity asks whether the thing it measures consistently is the thing it claims to measure. A bathroom scale that reads three kilograms heavy every time is perfectly reliable and completely invalid.
A test can be highly reliable and measure the wrong thing. It cannot be valid without being reliable. Good instruments need both, and the two are established by completely different kinds of evidence.
Reliability: the numbers
Major clinical instruments report internal consistency in the region of 0.95 to 0.98 for full-scale scores, and test–retest correlations around 0.90 or higher over short intervals. Those are high figures by the standards of psychological measurement.
The practical consequence is expressed as the standard error of measurement. For a full-scale IQ on a modern clinical battery, it is roughly two to three points, which gives a 95 percent confidence interval of about plus or minus five points.
So a clinical result of 118 means: the best estimate is 118, and the true value is probably somewhere between 113 and 123. Clinical reports print the interval alongside the number for exactly this reason, and the interval is the part that should be quoted.
Two scores within about six points of each other are the same score. This applies to two people being compared and to the same person on two occasions. A great deal of nonsense about IQ evaporates once this is taken seriously.
Validity: what the score predicts
IQ scores correlate with a long list of outcomes. Correlations with academic achievement typically land around 0.5, with job performance around 0.3 to 0.5 depending on job complexity and on how the analysis handles range restriction, and with income somewhat lower.
A correlation of 0.5 accounts for about a quarter of the variance. That is a substantial relationship by the standards of social science and a weak one by the standards of anyone expecting a score to determine a life. Both halves of that sentence are true and people tend to keep only the half they arrived with.
There is also a structural question. IQ tests are built from subtests that correlate with each other, and the general factor extracted from those correlations — g — is the most replicated finding in the field. Whether g is a single underlying capacity or a statistical summary of several partly-shared processes is a genuine and unresolved argument among people who agree on the data.
Where online tests sit
Online tests inherit all of the above and add four problems of their own, none of which are fixable by writing better items.
No administration control. Nobody verifies who is answering, how many attempts they have had, whether they are using a calculator, or whether they are taking the test on a phone on a bus.
Self-selected norms. Clinical tests are normed on stratified samples matched to census demographics. A web test is normed, at best, on visitors who found it and finished it — a group younger, more online and more self-conscious about test scores than the general population.
Length. Shorter tests are noisier. This is arithmetic, not a criticism: the standard error falls roughly with the square root of the number of items. A ten-item test can place you in a quarter of the distribution. Thirty items narrows that usefully. Nothing short can do what an hour of testing does.
Repeatability. Anyone can take an online test five times and keep the best score, and the format practice alone is worth several points.
None of this makes a free test useless. It makes it an indication rather than a measurement, which is a perfectly respectable thing for it to be as long as the site says so.
How to judge a test before you take it
You cannot audit an item bank from the outside, but you can read the site, and five signals tell you most of what you need.
- Does it report an interval? A single number with no error band is a claim to precision no unsupervised test can support.
- Does it describe its norm group? Silence is the loudest possible answer.
- Is the score free? A result held behind a payment is a sales funnel with a quiz attached, and the incentive runs towards flattering numbers.
- Does it guarantee accuracy? Nobody can. A guarantee is a marketing sentence, and its presence is informative.
- Does it promise Mensa qualification? No online test qualifies anyone for Mensa. A site claiming otherwise is either uninformed or dishonest.
The verdict
A well-constructed clinical IQ test, administered properly, is among the most reliable instruments in psychology and predicts a meaningful slice of academic and occupational outcomes. It is also not a verdict on a person, and the professionals who administer it are usually the first to say so.
A free online reasoning test is a reasonable indication of which broad band you fall into, a decent way to find out whether the format suits you, and worthless as evidence for anything official.
Both statements can be held at once. The trouble starts when a site borrows the credibility of the first to describe the second.
Questions
Questions people ask
What is the margin of error on an IQ test?
About plus or minus five points at 95 percent confidence on a modern clinical battery. On a short unsupervised online test it is considerably wider, which is why a band is a more honest output than a number.
Are online IQ tests accurate?
They are accurate enough to indicate a broad band and not accurate enough for any official purpose. The limits come from administration conditions and norm samples, not from the questions themselves.
Is a 5-point difference in IQ meaningful?
No. Five points sits inside the measurement error of even the best instruments. Treating it as a real difference between two people, or between two sittings by the same person, is the most common error in how these numbers are used.
Which IQ test is the most accurate?
For adults, an individually administered WAIS given by a licensed psychologist. For children, the WISC. Both take one to two hours and both are normed on stratified population samples.
Keep reading
Evidence
How to increase your IQ: what the evidence supports and what it does not
Brain-training apps, dual n-back, supplements, music, sleep. Sorting the interventions with evidence behind them from the ones with marketing behind them.
11 min read
Research
The Flynn effect: why IQ scores rose for a century, and why they may be falling
Raw scores climbed roughly three points a decade for most of the twentieth century. Nobody fully agrees why, and in several countries the climb has stopped.
9 min read
Research
Does IQ predict success? What the longitudinal studies show
It predicts a real slice of academic and job outcomes, and a much smaller slice than either the enthusiasts or the dismissers claim. Here are the actual correlations.
10 min readFind out where you land
Ten reasoning questions, an instant band and an explanation for every answer.