BlogKnowledge

Raven's Progressive Matrices: The Non-Verbal IQ Test

Raven's Progressive Matrices: The Non-Verbal IQ Test

Raven's Progressive Matrices is the best-known non-verbal reasoning test in psychology. It presents an incomplete grid of abstract shapes and asks the test-taker to work out which piece completes the pattern — no words, no arithmetic, no general knowledge. First published in 1938, it is still used in research, occupational screening, and cross-linguistic assessment, and it is the direct ancestor of the matrix-reasoning subtests found in most modern IQ batteries. This guide explains where it came from, how its items are built, what it does and does not measure, and how its scores should be read.

1. Where Raven's Progressive Matrices came from

The test was created by the British psychologist John C. Raven, who published the first version in 1938. Raven had worked as a research assistant to Lionel Penrose on a study of the origins of intellectual disability, and he needed an instrument that could be administered quickly, in the field, to people with very different levels of literacy and schooling. Existing tests of the era leaned heavily on vocabulary, general knowledge, and verbal instructions — all of which travelled badly.

Raven's theoretical starting point came from Charles Spearman, whose two-factor theory of intelligence proposed a general factor (g) running through all cognitive tasks. Spearman had described the "eduction of relations and correlates" — the mental act of perceiving how things relate to one another and extending that relation to new cases. Raven designed the matrices as a deliberately pure exercise in exactly that operation.

He paired the matrices with a vocabulary test (the Mill Hill Vocabulary Scale), which measured what he called reproductive ability: the store of learned information a person can recall. The matrices measured eductive ability: making sense of material that has no ready-made meaning attached. Together, the pair was intended to give a two-sided picture rather than a single number — a design idea that survives today in the distinction between fluid and crystallized intelligence.

The test's longevity is unusual. Very few psychological instruments from the 1930s are still commercially published and actively normed nearly ninety years later.

2. What a matrix item actually looks like

Raven's items are copyrighted, so no genuine item is reproduced here. The general structure, however, is public knowledge and easy to describe.

A typical item shows a matrix of abstract figures — commonly a 3 × 3 grid, though the easiest sets use simpler 2 × 2 or single-strip layouts. The bottom-right cell is blank. Below the matrix sits a set of candidate pieces, usually six or eight. The task is to identify which candidate belongs in the blank cell.

Several features of the design matter:

  • No language is required. Instructions can be mimed if necessary. The items themselves contain no text.
  • No prior knowledge helps. The shapes are arbitrary — arrows, dots, hatching, geometric outlines — and carry no cultural or educational content.
  • Difficulty is "progressive." Items are ordered from very easy to very hard, so early items effectively teach the logic that later items depend on. This is where the "Progressive" in the name comes from.
  • Distractors are systematic. The wrong answers are not random. They typically represent partially-correct rule applications: the right shape with the wrong shading, the right progression applied to the wrong dimension, and so on. This is what makes guessing unreliable on the harder items.

An illustration of our own making: imagine a row where the number of dots in each cell increases by one from left to right, while the shape enclosing them rotates a quarter-turn at each step. Two independent rules operate at once, and the missing cell must satisfy both. Harder items compound several such rules.

3. The main versions of the test

Raven never published a single test. The matrices exist as a family of instruments aimed at different ability ranges, and the ability range is what determines which one is appropriate.

Version Abbrev. Items Structure Intended for
Coloured Progressive Matrices CPM 36 3 sets of 12 (A, Ab, B) Young children, older adults, people with cognitive impairment
Standard Progressive Matrices SPM 60 5 sets of 12 (A–E) The general population, ages roughly 6 to adult
Standard Progressive Matrices Plus SPM+ 60 5 sets of 12 Same range as SPM, with harder items to reduce ceiling effects
Advanced Progressive Matrices APM 12 + 36 Set I (practice), Set II (scored) Adults and adolescents in the upper ability range
Raven's Progressive Matrices, 3rd ed. Raven's 2 Variable Digital, adaptive item selection Modern clinical and occupational use

The CPM uses coloured backgrounds, which hold attention better with young children and reduce the visual load. The SPM is the classic form most people mean when they say "Raven's." The APM was developed because the SPM compresses the top of the distribution — too many able adults hit the ceiling — so its items start where the SPM's hardest items leave off.

The 2018 revision published by Pearson, commonly referred to as Raven's 2, moved the test to a digital, computer-adaptive format, in which the difficulty of the next item depends on performance so far. Adaptive delivery reaches a comparable precision in fewer items and reduces the ceiling and floor problems of the fixed-form versions.

Administration is typically 20–45 minutes. Raven originally intended the SPM to be untimed, treating it as a test of reasoning power rather than speed; timed administration became common later, and it changes what the score reflects.

4. What the test actually measures

In factor-analytic studies, Raven's matrices load very heavily on fluid intelligence (Gf) and, through it, on g. Across the psychometric literature it is routinely described as one of the strongest single-test markers of general intelligence available — which is why researchers who need a short proxy for g so often reach for it.

Cognitively, what it demands is:

  • Rule induction — inferring an abstract regularity from a small number of examples.
  • Working memory / goal management — holding several partially-solved sub-goals in mind at once while testing candidate rules.
  • Attention to relevant dimensions — deciding which of the many visual attributes (shape, number, shading, orientation, size, position) actually carries the pattern.

The classic cognitive analysis is Carpenter, Just and Shell (1990), who modelled performance on the APM in detail and concluded that the tougher items were distinguished chiefly by the number of rules to be managed simultaneously and by the demand on goal management in working memory. They catalogued a small number of recurring rule types:

Rule type What it does
Constant in a row An attribute stays the same across a row and changes between rows
Quantitative pairwise progression An attribute increases or decreases in fixed steps along a row
Figure addition or subtraction Two cells combine, or one is removed from another, to give the third
Distribution of three Three values of an attribute each appear once per row and per column
Distribution of two Two values appear and one is null in each row

Note what this list implies: the test is not measuring "visual ability" in any narrow sense. It is measuring how efficiently someone can generate, test, and hold onto abstract hypotheses under load.

Neuroimaging work has associated matrix-reasoning performance with a distributed frontal-parietal network — the pattern summarised by Jung and Haier's parieto-frontal integration theory (2007). These findings describe correlates of performance; they do not localise "intelligence" to a spot in the brain.

5. The "culture-fair" question

Raven's matrices are frequently called "culture-free" or "culture-fair." The first term is wrong and the second is, at best, an approximation.

It is true that the test removes the most obvious cultural barriers: no language, no vocabulary, no curriculum content, no culturally specific objects. That is a genuine advantage, and it is why the test spread so widely into cross-linguistic research and into settings where the examiner and the test-taker share no language.

But culture-reduced is not the same as culture-free. Performance still depends on things that are distributed unevenly:

  • Familiarity with two-dimensional abstract figures, which is largely a product of formal schooling and of exposure to printed and screen-based material.
  • Test-taking conventions — multiple-choice formats, the assumption that one answer is "correct," working alone and in silence against a clock.
  • Amount of schooling, which correlates with matrix performance independently of the content of that schooling.

The consensus position in modern psychometrics is that no test is culture-free, that Raven's substantially reduces certain sources of bias, and that this reduction should not be overstated into a claim of universal comparability across populations.

6. Raven's and the Flynn effect

The Flynn effect — the long-run rise in average raw scores on cognitive tests through the twentieth century — was documented most dramatically on Raven's-type tests. James Flynn and others found that gains on abstract matrix reasoning were substantially larger than gains on vocabulary or general-knowledge tests, in some national datasets amounting to the order of a standard deviation across a generation.

This is a genuinely awkward result for the "culture-fair" reading of the test. If the matrices measured a fixed biological capacity untouched by environment, the scores of successive generations should not have drifted upward so sharply. The leading explanations point instead to environmental change: more years of schooling, a visual environment saturated with diagrams and symbols, and — Flynn's own preferred account — a cultural shift toward habitually using abstract, hypothetical, classificatory categories rather than concrete ones.

Two practical consequences follow. First, norms expire. A raw score interpreted against 1979 norms and the same raw score interpreted against 2018 norms give different results, and the older norms flatter the test-taker. Second, rising average performance on a test is not the same thing as a population becoming more intelligent in every respect; the gains were highly uneven across cognitive domains, which is itself the most informative part of the finding.

7. How Raven's scores are reported

The raw score is simply the number of items answered correctly — out of 60 on the SPM, out of 36 on Set II of the APM. That raw number means nothing on its own. It becomes interpretable only when compared to a norm table for the appropriate age group and population, which converts it into a percentile.

Some publishers additionally convert the percentile into a standard-score scale with mean 100 and standard deviation 15, so that results can be discussed alongside Wechsler-type IQ scores. That conversion is a convenience, not an equivalence: a Raven's-derived standard score and a full-scale WAIS IQ are built from different content and are not interchangeable.

Three cautions apply to any Raven's-derived figure:

  1. It is a single-domain measure. The matrices assess fluid, non-verbal reasoning. They say nothing about verbal comprehension, acquired knowledge, or processing speed. A full battery samples several domains for exactly this reason.
  2. It carries measurement error. Like any test, an observed score is an estimate. The 95 % confidence interval around a well-normed score typically spans several points either side of the observed value.
  3. Practice changes the score, not the person. People who have worked through many matrix puzzles get better at matrix puzzles. That improvement reflects familiarity with the item format and the recurring rule types — task-specific learning — and should not be read as a change in underlying general ability.

8. Where you will meet matrices today

Even if you never sit the Raven's test itself, you have almost certainly met its design. Matrix Reasoning is a core subtest of the WAIS-IV and WISC-V. Comparable figural-matrix tasks appear in the Stanford-Binet, in the Cattell Culture Fair test, in the Naglieri Nonverbal Ability Test, in occupational aptitude batteries, and in a large share of the free tests circulating online.

Some national branches of high-IQ societies have accepted supervised scores from the Advanced Progressive Matrices as part of their admission evidence, though criteria vary by country and change over time — anyone pursuing that route should check current requirements with the organisation directly rather than relying on secondhand lists.

Online matrix tests, including the pattern-reasoning items in the Brambin profile, use the same underlying logic but are not the trademarked Raven's instrument and are not administered under standardised conditions. They are useful for curiosity and self-exploration. They are not clinical assessments and cannot substitute for one.

Frequently asked questions

Is Raven's Progressive Matrices an IQ test?

It is a test of non-verbal fluid reasoning that correlates very strongly with general intelligence, and its results are often converted onto an IQ-like scale of mean 100 and standard deviation 15. But it measures one domain rather than the several that a full battery like the WAIS covers, so a Raven's result and a full-scale IQ are not the same quantity even when they are printed on the same scale.

How long does the test take?

It depends on the version and the administration protocol. The Standard Progressive Matrices typically takes 20 to 45 minutes; a 45-minute limit is common in timed group settings. Raven originally designed it to be untimed, on the view that reasoning power rather than speed was the object of interest, and untimed administration is still used in some research contexts.

Is Raven's really culture-free?

No. It is better described as culture-reduced. Removing language and curriculum content eliminates some large sources of bias, but performance still depends on familiarity with abstract two-dimensional figures, on formal schooling, and on test-taking conventions, all of which vary between environments. Most psychometricians now regard the term culture-free as an overstatement.

Can you get better at Raven's-style puzzles by practising?

Performance on matrix puzzles does improve with practice, because the recurring rule types become familiar and the format stops being surprising. What the research supports is that this is task-specific learning: the gain shows up on the practised task and transfers poorly to unrelated cognitive tasks. It should not be interpreted as a change in general reasoning ability, and it is one reason clinicians discount scores from recently repeated tests.

Why are there so many different versions?

Because a single 60-item form cannot measure accurately across the whole ability range. The Coloured version exists so that young children and adults with cognitive impairment are not stuck at the floor of a test that is too hard; the Advanced version exists so that able adults are not stuck at the ceiling of one that is too easy. The modern adaptive edition addresses both problems at once by choosing items in response to performance.

Where can I see real Raven's items?

Not legitimately in public. The items are copyrighted material published under licence, and their security matters: an item that has circulated widely no longer measures what it was normed to measure. Practice material that imitates the format exists, but it is not the test, and its scores are not comparable to normed Raven's results.

Summary

Raven's Progressive Matrices has lasted since 1938 because it does one thing exceptionally well: it isolates abstract rule induction from language, knowledge, and curriculum. That makes it one of the purest available markers of fluid reasoning and one of the most widely used research instruments in the field.

Its limits are equally clear. It measures a single cognitive domain. It is culture-reduced rather than culture-free. Its norms drift, as the Flynn effect demonstrated more vividly on this test than on any other. And a score from it is an estimate with a confidence band, not a fixed property of a person.

Read alongside other evidence, a matrix-reasoning result is genuinely informative. Read alone, as a verdict, it is being asked to carry far more than it was ever built to carry.


Brambin offers an eight-dimension cognitive profile, including pattern-reasoning items, designed for self-exploration. It is not the Raven's Progressive Matrices test, is not affiliated with its publisher, and is not a clinical assessment. Treat any online score — ours included — as a starting point for curiosity, not a verdict.

Want to explore more?

Download Brambin for 8 types of cognitive challenges with detailed score breakdowns.

Download Brambin
Download App