BlogKnowledge

The Stanford-Binet Intelligence Scale: History and Modern Use

The Stanford-Binet Intelligence Scale: History and Modern Use

The Stanford-Binet Intelligence Scale is the oldest IQ test still in professional use. It began as a practical screening tool built in Paris in 1905, was rebuilt at Stanford University in 1916, and gave the world the number we now call an IQ. The current edition — the fifth, published in 2003 — covers ages 2 through 85 and over, and measures five cognitive factors in both a verbal and a non-verbal form. This guide traces where the test came from, how it is put together today, what its scores mean, and where its history still casts a long shadow.

1. Paris, 1905: the original Binet-Simon scale

In 1904 the French ministry of public instruction faced a practical problem. Schooling had recently been made compulsory, and teachers needed some way of identifying children who were not keeping up and would need a different kind of instruction. The ministry asked Alfred Binet, a psychologist at the Sorbonne, to find a method that did not rely on a teacher's impression.

Binet, working with the physician Théodore Simon, took an approach unusual for the era. Rather than measuring reaction times, head size, or sensory acuity — the fashionable methods of the day, inherited from Francis Galton — they assembled short everyday tasks: following a moving object, naming familiar things, repeating a sequence of digits, explaining a difference between two words, putting items in order. The tasks ran from easy to hard, and the child worked upward until they became too difficult.

The 1905 scale was revised in 1908 and again in 1911. The 1908 revision introduced the idea that made the test famous: each task was tied to the age at which a typical child could pass it. A child who succeeded on the tasks that most eight-year-olds passed was said to have a "mental level" of eight, regardless of birthday. Binet's own framing was carefully hedged — he treated the scale as a rough practical instrument for spotting children who needed extra help, and warned explicitly against reading the result as a fixed or permanent capacity.

Binet died in 1911, before he could see what happened to his scale after it crossed the Atlantic.

2. Stanford, 1916: how the test became the IQ

Henry Goddard brought the Binet-Simon scale to the United States, translating it around 1908 and distributing it widely. But the version that stuck was produced at Stanford University by Lewis Terman, who published The Measurement of Intelligence and the Stanford Revision of the Binet-Simon Scale in 1916.

Terman did three things that changed the instrument fundamentally:

  1. He expanded and rewrote the item pool and extended the scale upward into adulthood, so it was no longer only a children's screening tool.
  2. He standardized it on a large American sample, so results could be read against a normative reference rather than a clinician's judgement.
  3. He adopted the scoring formula proposed by the German psychologist William Stern in 1912 — mental age divided by chronological age, multiplied by 100 — and named it the intelligence quotient.

That third decision is why the phrase "IQ" exists at all. A ten-year-old performing like a typical twelve-year-old scored 120; one performing like a typical eight-year-old scored 80. Simple, memorable, and — as it turned out — structurally broken.

Terman also parted company with Binet on interpretation. Where Binet regarded a low score as a signal to intervene, Terman treated scores as largely inherited and stable, and campaigned for their use in sorting pupils. This view fed directly into the American testing boom of the 1920s, and Terman himself was involved with organizations of the eugenics movement of that period — a fact the field now records plainly rather than skirting. Later editions of the test were built by different people under different assumptions, and modern professional standards for cognitive assessment were written in explicit reaction to these early misuses.

3. The ratio IQ and why it had to go

The mental-age formula works reasonably well for children and collapses for adults. Mental-age scores stop climbing in the late teens while chronological age keeps going, so a perfectly typical 40-year-old would compute a steadily falling quotient for no reason other than the arithmetic. The formula also produced inconsistent spreads at different ages: 130 did not mean the same degree of rarity at age six as at age fourteen.

The fix arrived with the third edition, Form L-M, in 1960, which replaced the ratio with the deviation IQ David Wechsler had been using since 1939. Here the raw performance of each age group is converted onto a standard scale with a fixed mean and a fixed spread, so a score describes where someone sits relative to their own age peers — nothing more, nothing less.

Edition Year Key change Score scale
Binet-Simon 1905–1911 Age-graded tasks, "mental level" None (mental level)
Stanford Revision 1916 US norms, IQ formula adopted Ratio IQ (MA ÷ CA × 100)
Second revision (Forms L & M) 1937 Terman with Maud Merrill; parallel forms Ratio IQ
Third edition (Form L-M) 1960 Deviation IQ replaces ratio IQ Mean 100, SD 16
Fourth edition (SB4) 1986 Point scale, four area scores Mean 100, SD 16
Fifth edition (SB5) 2003 Five factors × two domains Mean 100, SD 15

Two details in that table matter. The 1986 fourth edition, developed by Robert Thorndike, Elizabeth Hagen and Jerome Sattler, abandoned the age-scale format for a conventional point scale organized around four areas — verbal reasoning, abstract/visual reasoning, quantitative reasoning, and short-term memory. And the 2003 fifth edition shifted the standard deviation from 16 to 15, bringing Stanford-Binet scores onto the same scale as the Wechsler tests. Old Stanford-Binet figures quoted from before 2003 sit on a slightly wider scale, which is one reason historical scores should not be compared casually with modern ones.

4. How the SB5 is built

The fifth edition, developed by Gale H. Roid and normed on a sample of 4,800 people matched to United States census data, is organized as a five-by-two grid. Five cognitive factors are each assessed twice — once verbally and once non-verbally — giving ten subtests in total.

Factor Non-verbal form Verbal form
Fluid Reasoning Object series and matrix-style items Verbal reasoning problems
Knowledge Picture-based absurdities Vocabulary
Quantitative Reasoning Numerical items presented visually Word-based quantitative problems
Visual-Spatial Processing Form board and pattern tasks Position and direction problems described in words
Working Memory Block-based memory tasks Sentence and memory-for-detail tasks

No genuine item is reproduced here — the content is copyrighted and its security is part of what keeps the norms valid. The descriptions above are of the general task type only.

The five-factor structure is not arbitrary. It maps onto the Cattell-Horn-Carroll (CHC) model that has become the common language of modern cognitive assessment, which makes SB5 results easier to read alongside other CHC-aligned batteries.

The non-verbal domain deserves a note. "Non-verbal" here means the response requires little or no speech — pointing, moving, arranging — not that the session is silent. Instructions are still given orally, though they can be gestured where necessary. This makes the non-verbal side useful with children who have limited expressive language, with people who are deaf or hard of hearing, and with test-takers assessed outside their first language, without pretending that language has been eliminated entirely.

5. Routing, basals and ceilings: how the test is actually given

The Stanford-Binet is administered one-to-one by a trained examiner, typically over 45 to 90 minutes depending on age and how many subtests are needed. Its distinguishing feature is adaptive routing, an idea it has used in some form since Binet.

Two routing subtests come first: an object-series/matrices task on the non-verbal side and a vocabulary task on the verbal side. Performance on those two determines the difficulty level at which every remaining subtest begins. A six-year-old and a talented twelve-year-old will therefore sit different sets of items.

From that entry point the examiner works with basal and ceiling rules. The basal is the level at which the test-taker passes consistently, establishing that easier items can safely be assumed correct; the ceiling is the point at which failures accumulate and testing stops. Everything below the basal is credited and everything above the ceiling is not administered.

The benefit is efficiency and comfort — nobody spends twenty minutes on items far below or far above their level. The cost is that administration and scoring take real training, which is one reason this test cannot exist in a self-administered online form.

6. How SB5 scores are reported

The SB5 produces more than one number, and the composite figure is only the top layer.

Score type Mean SD What it covers
Full Scale IQ 100 15 All ten subtests
Non-verbal IQ 100 15 The five non-verbal subtests
Verbal IQ 100 15 The five verbal subtests
Factor indices (×5) 100 15 One cognitive factor across both domains
Subtest scaled scores 10 3 A single subtest

The SB5 manual also provides an Extended IQ (EXIQ) scoring option, which stretches the reportable range well beyond the conventional 40–160 band for research and for assessments at the extremes, and change-sensitive scores, an absolute scale designed to track an individual's performance across repeated assessments over time without the distortions of age-normed comparison.

Score bands come with descriptive labels — average, high average, superior, and so on. Those labels describe where a score falls in a distribution. They are not diagnoses, and no diagnosis of any condition is ever made from a score band alone; that requires a qualified professional weighing the full assessment, developmental history, and everyday functioning together.

Three interpretive cautions apply, as they do to every cognitive test:

  1. A score is an estimate with a confidence interval. The manual reports the standard error of measurement precisely so that results can be read as a band rather than a point.
  2. The composite can hide the profile. Two people with the same Full Scale IQ may have completely different patterns across working memory, knowledge, and visual-spatial processing — and the pattern is usually the more useful information.
  3. Norms age. The Flynn effect — the long-run drift in average raw performance documented across the twentieth century — means an old edition's norms will tend to produce flattering scores. This is a routine reason clinicians work from current editions.

7. Terman's longitudinal study and what it found

In 1921 Terman began the Genetic Studies of Genius, screening California schoolchildren with the Stanford-Binet and following roughly 1,500 who scored around 135 and above for the rest of their lives. It became one of the longest-running longitudinal studies in psychology and is still being followed by later researchers.

Its most durable finding was a negative one. The stereotype of the high-scoring child as sickly, socially awkward and physically frail did not survive contact with the data; on average the group was healthy, socially competent and academically successful. That result did real work in dismantling a nineteenth-century myth.

The study's limits are equally instructive. The sample was selected through teacher nomination before testing, was not representative of the wider population, and Terman remained personally involved with participants in ways that would not pass a modern ethics review. And while the group did well professionally, it produced no Nobel laureate — whereas two physicists who later won the prize, William Shockley and Luis Alvarez, are reported to have been screened and to have fallen short of the cut-off. Whatever one makes of that anecdote, the broad pattern is well replicated: a high test score is a real statistical advantage and a poor individual prophecy.

8. Stanford-Binet or Wechsler? Where each is used

For most of the twentieth century the Stanford-Binet was the default individual IQ test. Today the Wechsler scales — WAIS-IV/V for adults, WISC-V for school-age children — are more commonly administered in general clinical practice, while the Stanford-Binet retains particular strengths.

Stanford-Binet (SB5) Wechsler (WAIS/WISC)
Age range 2 to 85+ in one instrument Separate tests by age band
Structure 5 factors × verbal / non-verbal 4–5 index scores
Item selection Adaptive routing, basal/ceiling Fixed subtest order
Full non-verbal composite Yes, a complete parallel domain Partial
Common uses Very young children, assessments at the extremes, limited-language contexts General clinical and educational assessment

The SB5's single continuous age range makes it convenient for re-assessment across a lifespan, and its low floor and high ceiling — extended further by EXIQ scoring — make it a common choice when performance is expected to fall far from the average in either direction. In practice the instrument is chosen to suit the referral question, and a good assessment draws on more than one source of evidence anyway.

Frequently asked questions

Is the Stanford-Binet still used today?

Yes. The fifth edition, published in 2003, remains in active professional use for ages 2 through 85 and over. The Wechsler scales are more common in routine clinical practice, but the Stanford-Binet is still widely chosen for very young children, for assessments where performance is expected to fall well outside the average range, and where a full non-verbal composite is needed.

What is the difference between the Stanford-Binet and an online IQ test?

Almost everything except the vocabulary used to describe the results. The Stanford-Binet is administered individually by a trained professional, uses adaptive routing with basal and ceiling rules, and is interpreted against a large standardization sample under controlled conditions. An online test is self-administered in an uncontrolled environment with no examiner, no clinical norms, and no way to verify who took it. Both may report a number near 100, but the two numbers are not the same kind of measurement.

Why did the scoring scale change from SD 16 to SD 15?

Earlier Stanford-Binet editions used a standard deviation of 16, while the Wechsler tests used 15. The fifth edition adopted 15 so that scores from the two most widely used batteries would sit on the same scale. The practical consequence is that Stanford-Binet scores quoted from before 2003 are slightly more spread out than modern ones, and figures from old and new editions should not be compared directly without conversion.

Can you prepare for a Stanford-Binet assessment?

There is no legitimate practice material, because the items are copyrighted and their security is what keeps the norms meaningful. A score obtained soon after similar testing does tend to come out higher, but that reflects familiarity with the format — a practice effect and a measurement artifact — not a change in underlying ability, which is exactly why examiners record and discount recent prior testing. The useful preparation is ordinary: adequate sleep, food, and arriving unhurried.

What does a Stanford-Binet score actually tell you?

It places performance on a set of specific reasoning, knowledge, memory and spatial tasks relative to a same-age normative sample, on the day of testing, under those conditions. That is genuinely informative, particularly the profile across the five factors. It is not a measure of a person's worth, character, creativity, judgement, or likely life outcome, and no single figure should be read as a verdict on any of those.

Who was Alfred Binet and would he recognise the modern test?

Binet was a French psychologist who built the first practical intelligence scale in 1905 to identify schoolchildren needing additional support. He would recognise the adaptive structure and the age-referenced logic, both of which survive. He would likely be less comfortable with the rest: he warned against treating the score as a fixed quantity and never used the term intelligence quotient, which was applied to his scale by others after his death.

Summary

The Stanford-Binet has lasted more than a century because it kept being rebuilt. Binet's age-graded tasks became Terman's normed American scale, the broken ratio formula gave way to the deviation IQ, the age-scale format gave way to a point scale, and the modern SB5 reorganized everything around five CHC-aligned factors measured in parallel verbal and non-verbal forms.

Its history is not tidy. The same instrument that gave psychology its first usable measure of reasoning was also central to an era of testing whose assumptions the field has since rejected. Both facts belong in any honest account of it.

What a modern Stanford-Binet result offers is a careful, professionally administered estimate of performance across five cognitive domains, reported as a range rather than a point, most useful when read as a profile. What it does not offer — and never claimed to — is a fixed, final number that explains a person.


Brambin provides an eight-dimension cognitive profile designed for entertainment and self-exploration. It is not the Stanford-Binet Intelligence Scale, has no affiliation with its publisher, and is not a clinical assessment. Any online score, ours included, is a starting point for curiosity — not a diagnosis, a placement decision, or a verdict.

Want to explore more?

Download Brambin for 8 types of cognitive challenges with detailed score breakdowns.

Download Brambin
Download App