We Published Reliability Numbers We Hadn't Measured. Then We Measured Them.
If you have ever scrolled to the bottom of an online test and seen a line like "reliability: α = 0.90," you probably read it as a promise about the thing you just took. It almost never is. Until September 2026 our own pages did exactly this: next to the Multiple Intelligences test we printed a reliability range of 0.70–0.85, and next to the Emotional Intelligence test we printed 0.85–0.95. Neither figure had been measured on our data; both were lifted from the published literature about the framework. Then we computed the numbers on our own completions and they came in lower — Multiple Intelligences measured 0.55–0.77 per scale, Emotional Intelligence 0.53–0.76. We changed both pages to the measured figures, which means the number you see now is worse than the one you saw before. This article is about why that happened, why it is so common, and why the lower number is the more useful one.
What a borrowed alpha is actually describing
Cronbach's alpha measures internal consistency: whether the items in a scale behave as if they are all picking up one underlying thing. If the people who agree strongly with item three also tend to agree with seven and eleven, alpha is high; if the items scatter, it is low. That is the entire claim. Alpha is not a measure of validity: it says nothing about whether the scale measures what its name says, only that its items hang together. And a very high alpha often means the items are near-duplicates rather than that the scale is good, so a 0.95 can be a warning sign as easily as a badge.
Now consider what happens when a site copies that statistic out of a paper. Alpha is not a property of a framework. It is a property of a specific set of items, worded a specific way, answered by a specific sample. The published figure was computed on someone else's items, in someone else's language, on someone else's respondents, and very often on a much longer form — alpha rises with scale length more or less mechanically, so an academic instrument will almost always beat a short web version of the same ideas. When a site prints the literature range beside its own brief quiz, the number is not slightly optimistic; it is describing a different instrument. Nobody has to be dishonest for this to happen. The paper is real, the citation is real, and the practice is so widespread it looks like the responsible thing to do. It just tells you nothing about the test in front of you.
What our own numbers look like
Measured on our own completions, Multiple Intelligences ran 0.55–0.77 across its eight scales, on a norm table built at n = 283. Emotional Intelligence runs 0.53–0.76 on its five, at n = 96. Those are honest working numbers for short, self-report web scales, and the spread matters more than the headline: some scales hold together respectably and some hold together loosely, which is exactly what a single borrowed range hides. A result from a scale at the lower end deserves reading as a rough indication rather than a precise one — a judgment you can only make if someone tells you which end your scale sits on.
There is a wrinkle in that paragraph worth pulling on, because it is the same lesson one level deeper. The Multiple Intelligences figures describe the version of the test we measured, not the version you would take today. After measuring, we added items to every scale to give each one enough questions to separate the construct it names — and the moment we did that, the measurement stopped describing the live form. Each scale is now back to a handful of responses on its current wording, which is far too few to compute anything, so the honest report is a past-tense range plus an explicit note that no current figure exists yet. Improving a test destroys its own evidence base for as long as it takes people to answer the new version. A publisher unwilling to accept that trade never improves anything, and one that quietly keeps quoting the old number is describing a test nobody can take.
There is a second limitation on both pages that matters as much as the alpha. The people we compare you against are a self-selected sample of online test-takers, not a representative population sample. Nobody drew them at random from a country; they arrived because they were curious, which is a real and specific kind of person. So a percentile here means "relative to others who chose to take this test on this site," not "relative to the general public." That sentence is less flattering than its absence, and it changes how you read every comparative number we show you. If the gap between a measurement and a descriptive label is new to you, what personality tests actually measure covers the underlying layers.
The honest cost of fixing a scale
Our Big Five test shows what it costs to do this properly. After its item rebuild, three scales have measured figures: Conscientiousness at .71, Extraversion at .80, and Emotional Reactivity — the dimension usually called Neuroticism — at .72, all on n = 289. Openness and Agreeableness show no number at all. Their items were rebuilt too and no sitting has yet answered the new form, so there is nothing to compute. Any scale below n = 25 is reported as "not computed," never as a number, because a figure from a dozen people is noise wearing a decimal point.
That is the trade you make when you fix something. The old number described the old items, so the moment the items changed it stopped being true, and the only defensible move was to delete it and wait. A blank is uncomfortable to publish — it looks like an oversight, and it is the one thing a marketing instinct will push you to fill. But a stale figure on rewritten items is worse than a gap, because it says something confident and false. If you take Openness or Agreeableness today you are helping build the number that will eventually sit there; until then, read those two scales as provisional.
Where alpha isn't even the right question
Some of our tests cannot have an internal-consistency figure at all, and the difference between unmeasured and undefined is worth being precise about. On forced-choice forms your scores are shares of one fixed total: every point you give to one option is a point you did not give to another. Learning Styles works this way, and so do 16 Personalities, DISC, Love and Affection Styles, and Conflict Style. Because the scores are forced to trade off against each other, alpha means nothing on that form — not that we have not got round to it, but that the figure would be an artifact of the format rather than a fact about the scale. Anyone printing a clean alpha beside a forced-choice profile is printing something uninterpretable. These forms still tell you useful things about relative preference; they just cannot be audited with this tool.
What we still haven't done
Being specific about what is missing is part of the same job. No across-sitting test–retest study has ever been run on this platform, for any test, so we cannot tell you how stable your scores would be if you came back in six weeks. Not a low figure — none at all, because the data does not exist. We have run no convergent-validity comparison against an external published instrument, so we cannot claim our scale for a construct tracks the established one. No test here has been calibrated with item response theory or Rasch modelling, and no differential-item-functioning analysis has been run, so we have not formally checked whether any item behaves differently for different groups. None of those gaps is closed by a good alpha. Internal consistency is about the cheapest thing you can measure; it is the floor, not the ceiling.
None of this makes the tests useless, and none of it makes them clinical. They are instruments for thinking about yourself — noticing patterns, comparing frameworks, finding language for something you half-knew. They are not diagnostic or medical tools and nothing here is health advice. Their value survives an alpha of 0.55 perfectly well as long as you know it is 0.55. It does not survive being told 0.90 and finding out later.
What changed for you
Concretely: two pages now show lower numbers than they did, both state that the comparison sample is self-selected, three Big Five scales show measured figures, two show nothing, and the forced-choice tests say internal consistency is undefined for their format instead of skipping the topic. The use of all this is calibration. When a scale sits at the low end, hold the result loosely and check it against something else. When it sits at .80, lean on it a little harder. When the number is absent, treat the result as directional. It is also the cheapest way to judge any test site you come across, paid or free — see free vs paid personality tests — because a platform willing to print a 0.55 measured something, and one printing a borrowed 0.90 did not.
The tests are free, so if you want to see what this looks like in practice, take one and read the reliability line at the bottom of the page along with your result. The full picture — every scale, every sample size, every analysis we have not run — is laid out on our methodology page.