IQTest EN

Methodology

How this site scores its test

Everything here is a description of what we do and, more usefully, of what we cannot do. If you only read one section, read the one on limitations.

Updated September 2026 9 min read Reviewed by the IQTest EN editorial team

How the items were built

The bank is ten items across five domains: two abstract matrix items, two number series, two verbal analogies, two spatial rotation items and two logic problems.

Every item is generated from an explicit rule that is published on the test page after you finish. This is a constraint on us rather than a feature for you: an item whose rule cannot be written down in two sentences is an item that works through ambiguity, and it does not go in.

The figures are drawn as vector graphics from a small set of geometric primitives rather than assembled by hand, which means the visual properties of each item are generated by the same code that generates its answer. An item cannot drift out of agreement with its own key.

How they were checked

Three checks run before anything is published, and each one has caught a real problem.

  1. Uniqueness of the key. Every item is solved independently against its generating rule, and the distractors are checked to confirm that none of them also satisfies the rule. An item with two defensible answers is a broken item, however clever it looks.
  2. Chirality on spatial items. The rotation item asks you to identify a reflection among rotations. That only has an answer if the base figure is chiral — if its mirror image genuinely cannot be produced by any rotation. This is verified by enumerating all four rotations and the mirror and confirming no overlap. A figure that happens to be symmetric would leave the item with no correct answer at all.
  3. Distractor distinctness. No two options within an item may be identical after normalisation, which is a mistake that is easy to make and impossible to spot by eye when the options are figures rather than text.

How the score is calculated

Ten items, one point each, no weighting by difficulty and no partial credit. The raw score maps onto a band, and the band maps onto a percentile range.

The mapping is derived from the normal distribution — the theoretical curve that the IQ scale is built on — rather than from the performance of our own visitors.

That is a deliberate choice and it is worth being explicit about why. We could calculate percentiles from the people who take this test, and those percentiles would be higher and more flattering. They would also be wrong, because our visitors are not a representative sample of anything. They are people who searched for an IQ test and finished one, which selects for age, internet use, curiosity about test scores and prior experience with the format.

The trade we are making

Mapping to the theoretical curve gives lower, less exciting numbers than mapping to our own traffic would. It is the more defensible of the two options, and the difference between them is roughly the difference between a reference site and a lead magnet.

Why a band and not a number

The standard error of measurement falls roughly with the square root of the number of items. On a full clinical battery of an hour or more, it is around two to three points, giving a 95 percent confidence interval of about plus or minus five.

A ten-item test is nowhere near that. The interval on a score derived from ten items is wide enough to span several of the bands people care about, which means a single number would be a claim we could not support.

So we report the band, and we say on the result screen that the band is what the test supports. If you want a tighter estimate, a longer test genuinely provides one, and we say where to find one rather than pretending ours is longer than it is.

What we cannot fix

Four limitations are structural. They apply to every unsupervised online test, including ours, and no amount of care with the items changes any of them.

  • No administration control. We do not know who is answering, how many times, in what conditions, or whether anyone is helping.
  • No representative norm sample. We have no stratified sample matched to census demographics, and we are not going to pretend otherwise.
  • Short length. Ten items is a small sample of behaviour and the resulting estimate is correspondingly noisy.
  • Unlimited retakes. Practice effects on reasoning formats are real, and the bank is fixed, so a second attempt is a memory test.

There are two further limits specific to the item set. Two items per domain is a very small sample of each, so the breakdown on your result screen is a rough shape rather than a cognitive profile. And the items are written for adults, so a child's result would be compared against the wrong group.

What we would need in order to claim more

If the question is what it would take for this test to report a number rather than a band, the answer is specific.

A bank of forty or more items, calibrated with item response theory on a sample large enough to estimate difficulty and discrimination parameters per item. A norm sample stratified to population demographics rather than drawn from site traffic. A mechanism to detect and exclude repeat attempts. And published reliability statistics from that sample, so the claim could be checked rather than asserted.

That is what the difference between a free web test and a psychometric instrument consists of, and it is mostly expense rather than cleverness.

Corrections

If you think an item has more than one defensible answer, or that an explanation is wrong, tell us. We would rather pull an item than defend one, and the entire point of publishing the generating rule for every item is to make disagreement possible.

Reference tables on this site are generated from the normal distribution, and the code that produces them is the same code that produces the numbers in the text, so the two cannot drift apart. Where a table rests on outside data — the national IQ dataset is the only one — the page says whose data it is and what is wrong with it.

Contact us.

Reviewed by the IQTest EN editorial team Every reasoning item on this site is solved independently by two reviewers before it is published, and every number in a reference table is traced back to a named source. Where the research is contested, we say so on the page instead of picking the tidier answer. How we write and check this.

Questions

Questions people ask

How is the score on this test calculated?

Ten items, one point each, mapped onto a band derived from the normal distribution rather than from the performance of this site's visitors.

Why does this test not give an exact IQ number?

Because ten items cannot support one. The measurement error on a test that short spans several of the bands people care about, so a single number would be a claim we could not defend.

Do you use your own visitors as a norm group?

No. Our visitors are self-selected and not representative of any population. Mapping to the theoretical curve gives lower numbers and is the more defensible choice.

How do you check the questions?

Every item is solved against its generating rule, distractors are checked to confirm none of them also satisfies it, spatial figures are verified as chiral, and options are checked for accidental duplicates.

Can I report a problem with a question?

Yes, and we would rather hear it. The generating rule for every item is published precisely so that disagreement is possible.

Find out where you land

Ten reasoning questions, an instant band and an explanation for every answer.