This page explains exactly how the number at the end of the test is produced. If you think a site offering IQ tests ought to be able to answer these questions, you are right — and you will notice how few do.
Where the questions come from
Every question was written for this site. None is taken from Raven’s Progressive Matrices, the WAIS, the WISC or any other published instrument: those are protected works, and reproducing them on a website is unlawful. The underlying reasoning principles — series, matrix analogies, mental rotation — are in the public domain, and that is what the bank is built on.
One practical consequence: these particular questions do not circulate elsewhere, which limits how many people arrive already knowing them.
The model
The score is not a percentage of correct answers. That would treat a very easy question and a very hard one as equivalent, which makes no sense: getting the hardest item right does not carry the same information as getting the easiest one right.
The calculation uses a three-parameter item response model. For each question, characterised by a difficulty, the model gives the probability that a person of a given ability answers correctly:
P = c + (1 − c) / (1 + exp(−a × (ability − difficulty)))
The c term is the probability of guessing right — one in four on a four-choice question. Modelling it explicitly matters: without it, someone guessing their way through the whole test would be credited with ability they never demonstrated.
Your ability is then estimated as the value that makes your particular pattern of answers most likely, and reported on the usual scale, mean 100 and standard deviation 15.
Why a range
The same model reports the precision of its own estimate. The more consistent your answers are with a single ability level, the tighter the estimate; the more scattered they are — hard items right, easy items wrong — the wider it gets.
The range shown is the 95% interval. It is rarely narrow, and that is honest: no twenty-minute test places anyone to within two points. A site that gives you a bare number is not measuring better, it is communicating less honestly.
The sub-scores
The four sub-scores are computed the same way, but on nine questions each (two on the express test). Nine questions give an indication of profile, not a measurement. A ten-point gap between two of your areas means little; a thirty-point gap is worth noting.
The limitation you should know about
The difficulty values attached to the questions started as estimates made by the item writer — not values measured on a representative sample. That is the real limitation of this test, and of every free online test, whatever they claim.
To narrow it, the site records anonymously which questions are answered correctly. No IP address, no cookie, no identifier: only the pattern of right and wrong answers. This lets the difficulty of each question be re-estimated from how people actually perform, so the score becomes better anchored over time.