Skip to main content

What a New Test Cannot Know About Itself Yet

6 min readMy Path Research

Somewhere on this platform there is a page that tells you thirteen people have completed the test you are looking at. The natural reaction to that number is suspicion — thirteen sounds like a prototype, or a hobby project, or something that was put online before it was finished. But the number is not describing the test. It is describing the audience. Those are two different facts, and most of the confusion people have about test quality comes from collapsing them into one. A test can be assembled with real care and still have almost nothing measured about it, because measurement requires respondents and respondents arrive on their own schedule. So before you decide what a small number means, it is worth separating the two questions that number sits between.

Build quality and evidence are two different axes

The first question is how well the instrument is built: whether the items are clearly written, whether each scale has enough of them, whether ties in the scoring are handled honestly instead of broken by alphabetical accident, whether the result text says what the score can and cannot support. The second question is how much is known about it: whether enough people have taken it that its internal consistency can be computed, whether its score distribution has been observed, whether anything at all has been measured about how it behaves in the wild. Our own internal scoring keeps these on separate axes and refuses to average them. Across the self-assessment shelf, build quality averages 8.2 out of 10; the evidence about that same shelf averages 4.7. A single combined number in the middle of those two would have been worse than useless — it would have hidden exactly the distinction that tells you what to do next.

Because the two axes call for completely different responses. A low build score is a to-do list: the items are vague, a scale is too short, the result copy overreaches, and all of that can be fixed by someone sitting down and fixing it. A low evidence score usually just means the instrument is young. Treating the second as a defect is how a publisher ends up rewriting a perfectly sound test that simply nobody has taken yet — throwing away good work in response to a number that was never about the work. Build quality can be repaired in a week with no new respondents at all. Evidence takes months and needs people. No amount of code produces a reliability coefficient.

What a small count actually tells you

Here is the shelf, measured, by lifetime completions: Career interests 8,919; Big Five 298; Multiple Intelligences 289; Learning Styles 158; Emotional Intelligence 98; 16 Personalities 87; Love & Affection Styles 101; DISC 56; Enneagram 32; Social Skills 19; Attachment Style 18; Grit 15; Character Strengths 13; Conflict Style 12. Six of the fifteen tests sit below 35 lifetime completions. At that sample size, any judgement about the items or the scoring is a judgement about design, not about measured behaviour — and no reliability coefficient is computable at any quality of code. This is why a scale with fewer than 25 responses is always reported on our pages as "not computed" rather than as a number. A coefficient calculated on twelve people is not a weak estimate of the truth; it is noise wearing the costume of a statistic, and printing it would be the dishonest option dressed as the rigorous one.

Only four tests carry enough traffic for their reliability figures to be more than indicative: Career interests at 8,919, the Big Five at 295, Multiple Intelligences at 287, and Learning Styles at 158. That is the honest ceiling of what is currently known here, and it is a short list.

Improving a test destroys its own evidence

Now the part that is genuinely uncomfortable, and the reason the two-axis idea matters more than it first appears. Three of those four well-measured tests have had their item banks revised since the figures were gathered. Which means the trustworthy numbers describe an earlier version of the test — not the one you would take today. A test that was measured and then improved has no current measurement. The evidence does not transfer across a rewrite, because the thing being measured is no longer the same thing.

The Big Five shows the cost concretely. After its item rebuild, Conscientiousness measures .71, Extraversion .80 and Emotional Reactivity .72 on n = 289 — while Openness and Agreeableness are not computed at all, because their items were rebuilt and nobody has answered the new form yet. The weak number went away and no number replaced it. That is what improvement looks like from the inside: you fix the scale that was underperforming, and in doing so you delete the only figure you had for it, and then you wait. A publisher unwilling to accept that trade never improves anything. They keep the old coefficient because it looks reassuring on the page, and the items stay as they were, and the test is permanently as good as it was the day someone first measured it.

The quadrant nobody expects

The most striking case on our shelf runs the other way entirely. Our Attachment Style test has the best tie-handling on the platform — the part of the scoring most tests quietly get wrong — and it has been taken 18 times. There is nothing to fix there. The build is the strongest thing about it. What it lacks is visitors, which makes it a promotion problem wearing the mask of a quality problem, and the two look identical if you only ever look at one number. The same is true of Grit, Character Strengths, Conflict Style and Social Skills: carefully built, barely taken. If you are the nineteenth person to complete one of those, your result is computed by exactly the same scoring engine as the 8,919th career-interests result. What differs is not the care taken with your answers. It is how much we can tell you about how the instrument behaves across people.

What is missing, stated plainly

Some gaps are not about sample size at all and will not close with traffic. No across-sitting test-retest study has ever been run here, on any test — that data does not exist, so we do not report stability over time. For the forced-choice forms, where your scores are shares of one total rather than independent ratings, internal consistency is undefined rather than merely unmeasured; the arithmetic that produces the coefficient does not apply to that shape of score, and reporting one would be a category error. Two tests now have norm tables built on our own completions, Multiple Intelligences at n = 283 and Emotional Intelligence at n = 96, and both of those samples are self-selected online test-takers rather than a representative population, which is worth holding in mind when a percentile tells you where you sit. None of these instruments are clinical or diagnostic; they are tools for thinking about yourself, and the numbers attached to them describe how much is known, not how much you should trust your own reading of the result.

So when a page says thirteen people have taken something, read it as a fact about the instrument's age and reach, and read the build quality separately. If you want the ranking of what is worth taking regardless, what personality tests actually measure is the better starting point, and the full scorecard — per test, both axes, with the not-computed cells left visibly empty — lives at our methodology page.

We use essential cookies to keep you signed in and optional analytics cookies to improve the product. See our Privacy Policy.