IQTest EN

History

A short history of IQ testing

Binet built a tool to help struggling children. Within twenty years it was being used to justify immigration quotas. How the field got from one to the other.

Published 19 September 2026 10 min read By the IQTest EN editorial team
A short history of IQ testing

1905: Binet's practical problem

Alfred Binet was asked by the French education ministry to solve something concrete. Universal schooling had been introduced, and schools needed a way to identify children who would struggle in ordinary classes so they could be given extra help.

With Théodore Simon he built a scale of thirty short tasks ordered by difficulty, from following a light with your eyes to defining abstract words. A child's performance was expressed as a mental age: the age at which the average child performed as this child did.

Binet was explicit about the limits of his own instrument. He wrote that the scale did not measure intelligence as a fixed quantity, that intelligence could be increased by education, and that the scale should never be used to label a child permanently. Every one of those warnings was ignored within a decade.

1912: the quotient

The German psychologist William Stern proposed dividing mental age by chronological age. Multiply by 100 and you have the intelligence quotient: an eight-year-old performing like a ten-year-old scores 125.

The quotient was an elegant solution for children and a broken one for adults. Mental age stops rising in adolescence while chronological age does not, so the ratio makes every adult look progressively less intelligent with each birthday. Modern tests abandoned the ratio entirely in favour of the deviation IQ — a position on the distribution of your own age group — which is what every test has used since the 1930s. The name stuck anyway.

1916: Terman and the American turn

Lewis Terman at Stanford revised and renormed Binet's scale on American children, producing the Stanford–Binet, which became the dominant instrument for decades and remains in print in its fifth edition.

Terman also inverted Binet's purpose. Where Binet built a tool to direct help towards struggling children, Terman saw a measure of innate capacity that should determine a person's place. He wrote that testing would lead to the curtailment of the reproduction of feeble-mindedness, and he was not on the fringe of the field — he was its most prominent American figure.

1917: mass testing and what was done with it

The First World War gave American psychologists the opportunity to test at scale. The Army Alpha and Beta tests were administered to around 1.75 million recruits, establishing group testing as a practical technology.

The results were then used to argue that the average mental age of American recruits was thirteen, and that recruits from southern and eastern Europe scored lower than those from northern Europe. The methodological problems were severe and were pointed out at the time: the tests were administered in chaotic conditions, the supposedly non-verbal Beta version still required test-taking conventions unfamiliar to many recruits, and the scores tracked years of residence in the United States closely enough to make the innate-capacity reading untenable.

The findings were nonetheless cited in the debate leading to the Immigration Act of 1924, which imposed national-origin quotas. This is the part of the history that gets left out of test manuals, and it is the reason the field's ethical guidelines look the way they do.

Why this history is on this page

A measurement technology with real predictive validity was used, within twenty years of its invention, to support conclusions its own inventor had explicitly warned against. That is not an argument against measurement. It is an argument for stating what a score does and does not license, every time — which is why every page on this site carries that statement.

1939: Wechsler and the modern form

David Wechsler, working at Bellevue Hospital in New York, built a test for adults that broke in several ways from the Stanford–Binet.

It used the deviation IQ rather than the mental-age ratio. It reported separate verbal and performance scores rather than a single figure, on the grounds that the profile was clinically more informative than the total. And Wechsler defined intelligence as the global capacity to act purposefully, think rationally and deal effectively with the environment — a definition that treats it as an aggregate rather than a single thing.

The structure he established is still the structure of the field. The Wechsler Adult Intelligence Scale is in its fifth edition, the children's version in its fifth, and between them they remain the reference instruments for individual assessment.

1980 onwards: a quieter, more technical field

Four developments define the modern period, and none of them made headlines.

The CHC model. Work by Cattell, Horn and Carroll produced a hierarchical model of cognitive abilities with a general factor at the top and a set of broad abilities beneath it. Most contemporary tests are built to map onto it.

Item response theory. Modern scoring models item difficulty and person ability on a common scale, which is what makes adaptive testing possible and makes it meaningful to compare people who answered different items.

Bias analysis as routine. Differential item functioning analysis is now a standard step in test construction, flagging items that behave differently across groups matched on overall ability.

The Flynn effect. Documented in the 1980s, it forced the field to confront how much of measured performance is environmental — and to take renorming intervals seriously. The longer story.

What the history is for

The useful reading of this history is not that IQ testing is discredited. It plainly is not: modern instruments are among the most reliable measurements psychology has, and they do real work in clinical and educational settings.

The useful reading is that the gap between what a measurement supports and what people want it to support is where the damage happens, and that the gap does not close on its own. Binet said so in 1905 and was right.

Reviewed by the IQTest EN editorial team Every reasoning item on this site is solved independently by two reviewers before it is published, and every number in a reference table is traced back to a named source. Where the research is contested, we say so on the page instead of picking the tidier answer. How we write and check this.

Questions

Questions people ask

Who invented the IQ test?

Alfred Binet, with Théodore Simon, published the first practical intelligence scale in 1905 to identify French schoolchildren who needed additional support.

Where does the term IQ come from?

William Stern proposed the intelligence quotient in 1912: mental age divided by chronological age, multiplied by 100. Modern tests no longer use the ratio, but the name persisted.

Why are modern IQ tests not based on mental age?

Because mental age stops rising in adolescence while chronological age does not, which makes the ratio nonsensical for adults. Tests switched to the deviation IQ, a position within your own age group.

What is the oldest IQ test still in use?

The Stanford–Binet, first published by Lewis Terman in 1916 and currently in its fifth edition.

Find out where you land

Ten reasoning questions, an instant band and an explanation for every answer.