How it works

How we score, and what we will not claim

A result is only as good as the evidence behind it. This page says where the questions come from, how each test is scored, and what we show at each level of evidence. It is kept in step with the instrument sheet we maintain for every test.

The evidence ladder

What a result page may show depends on how much evidence sits behind the test, in your language. The rule runs on the server, so no result page, share card or preview can show more than its level allows.

LevelWhat it meansReasoning Challenge showsPersonality tests show
L0 · DevelopmentThe item bank has not been calibrated on our visitors yet.Called “Reasoning Challenge Beta”. Correct answers in total and per section. No IQ, no percentile, no interval.Big Five: scale scores with a band that says where your answers sit on the scale, not a ranking. 16 Types: a preference profile, stated as self-report.
L1 · PreliminaryAt least 300 completions in the language, an item analysis with weak items removed, reliability of 0.70 or more.Still Beta. After review, a band showing where you sit among the people who took it here, with the sample described.After review, a band within the same-language sample, with the sample described.
L2 · CandidateAt least 1,000 completions in the language, a documented sample, retest reliability of 0.80 or more.Not renamed automatically. Only after reviews of validity, of whom the sample represents, and of measurement error can it be called an IQ estimate with a standard score, percentile and interval.Only after the same three reviews: percentiles, and the norm version they come from.

The numbers are internal checkpoints, not a licence. Measuring consistently and measuring intelligence are two different things, and a thousand self-selected visitors do not automatically represent anyone else.

Each test, plainly

The current definition of every test, in the words of its instrument sheet.

Big Five

  • Definition v2
  • 50 questions
  • No time limit
  • Evidence level L0
  • Items in English
Where the questions come from
The 50-item IPIP Big-Five Factor Markers (Goldberg, 1992), from the International Personality Item Pool, which places them in the public domain. Items are used verbatim, in the published order, ten per trait.
How it is scored
Positively keyed items score 1 to 5 and reverse-keyed items 6 minus the answer; each trait sums its ten items to a scale score from 10 to 50. Every question must be answered; nothing is filled in for you.
What you see today
Five scale scores, each with a band that says where your answers sit on the scale, and a short reading per trait. No percentiles. Emotional stability carries a note that it is not a mental-health measure.
Known limits
Self-report, so answering style matters. The items are English originals; a Chinese version goes through translation, back-translation and interviews before release. We have no reliability figures from our own visitors yet; published values for this scale are typically 0.77 to 0.86, on other people's samples.

16 Types

  • Definition v2
  • 60 questions
  • No time limit
  • Evidence level L0
  • Items in English
Where the questions come from
Sixty statements we wrote ourselves, fifteen per dimension, alternating between the two poles, answered on a five-point agreement scale.
How it is scored
Each answer moves its dimension toward one pole; neutral answers do not count. A dimension shows the share of possible movement toward its first pole. Above one half gives one letter, below gives the other, and exactly one half gives no letter: the result says the dimension is undecided and lists the possible types instead of forcing one. Leanings within ten points of the middle are marked as slight.
What you see today
A four-letter code with our own type name, a description, three strengths and two things to watch, four bipolar tracks with the leaning, and the types one letter away.
Known limits
A type is shorthand for self-reported preferences, not a norm score. Answers move with mood and situation, and a retake a week later can land differently. The percentages describe your answers, not a comparison with other people. No reliability data from our visitors yet.

Reasoning ChallengeBeta

  • Definition v3
  • 30 questions
  • 20 minutes, timed
  • Evidence level L0
  • Items in English
Where the questions come from
Forty items in four sections. Matrix reasoning and mental rotation are generated as figures by our own code and checked geometrically; number series and verbal analogies are written by us. Each attempt draws thirty, stratified by section and difficulty, easiest first.
How it is scored
One point per correct answer; once the time is up, skipped items count as wrong. Internally the code also computes a provisional standard score, but at level L0 the server removes it before anything leaves the database.
What you see today
Correct answers in total and per section, and a plain note about why there is no IQ, percentile or interval yet.
Known limits
The items have not been through an item analysis; difficulty labels are the author's judgement. Verbal items depend on English vocabulary and cannot simply be translated. It is timed, and device and network affect the score. The figure items are checked by code and still await an independent human solve, which will be recorded in the instrument sheet.

Rules we hold ourselves to

  1. 1None of these tests is a diagnosis, and none is a substitute for one.
  2. 2We do not use the words “Mensa”, “official” or “genius”. Our type names are our own; MBTI® is a trademark of The Myers-Briggs Company, and this site is not affiliated with it.
  3. 3A stored result is always read with the version of the test that produced it. Older results are marked as such, never re-scored by newer rules.
  4. 4Answers are used for item analysis only after they are separated from accounts. Weak items are retired by a recorded decision, never silently.
  5. 5Every change to items, scoring or norms is written into the test's instrument sheet with a version and a date.

Take a test with that in mind

Every test is free, and its result says exactly as much as the evidence allows.