Skip to main content

How to Measure a Forced-Choice Test's Consistency Without a Second Sitting

8 min readMy Path Research

Every honest test platform eventually has to answer a question about itself: if you gave someone the same form twice, would it say the same thing? The standard tools are Cronbach's alpha, which asks whether a scale's items hang together within one sitting, and test–retest correlation, which asks whether the same person gets the same result weeks later. Neither is available for a forced-choice form on a platform that has never brought anyone back for a second sitting — and we haven't, for any test here. Closing that gap meant deriving a measurement rather than looking one up, and walking into a trap that produced a confidently wrong answer first. The trap is the more useful half.

Why alpha has nothing to say about a forced-choice form

A forced-choice form — the technical word is ipsative — asks you to pick between two options. DISC does this, and so do Conflict Style, Love & Affection Styles and 16 Personalities. You are not rating how much a statement sounds like you on a scale of one to five; you are choosing, and every choice in favour of one type is simultaneously a choice against the other. The consequence is structural: the scores are shares of one fixed total, so if your dominance share rises, another share has to fall.

That is exactly what alpha cannot survive. Alpha is a ratio built from the variances of individual items and the variance of their sum, and on an ipsative form that sum is the same number for everyone by construction, so the denominator does not vary. Alpha for a forced-choice score is therefore undefined: not unmeasured, not awaiting more data, not something we could report once enough people have taken the test. It has no value at all, and printing a Likert-style alpha beside a forced-choice score would be a category error dressed up as rigour. The usual fallback is test–retest, which needs the one thing we do not have.

The form already asks you twice

Here is the observation that opened the door. A forced-choice form is built from pairs of types, and a form with few types has to reuse those pairings. Look at how few are in play here: DISC asks 28 items over only 6 distinct type-pairs, and 16 Personalities asks 60 items over only 4. That is not an accident — four types make only six pairs — and it means you are asked about the same opposition repeatedly, in different words.

Every repeated pairing is a chance for you to agree or disagree with yourself. If a form asks you seven times whether you lean assertive or steady, your answers either point mostly the same way or they scatter — measurable from one sitting, with no new items. The check was already inside the form, unread. So we computed it, per respondent and per type-pair — what share of their choices went the same way as their own majority on that pair — then aggregated across pairs and everyone who finished.

The trap: the majority of a coin flip is never below half

The first version of this analysis was wrong, and wrong in the way that feels right. We computed the agreement figure, got something in the low seventies, compared it to the fifty percent you would expect from random answering, and concluded the forms showed solid consistency. That was an artefact of the statistic, not a finding about the tests.

Think about what the number is. A respondent's agreement on a pair is the share of their choices matching their own majority on that pair — not a fixed target or a pre-registered direction, but a majority computed after the fact from the very answers being scored. And the majority side of a set of coin flips is never the minority side. Flip a coin seven times and the more common outcome appears at least four times out of seven, always, by definition. Someone clicking at random therefore does not score fifty percent here; they score the expected majority share for that many flips, which sits above half and climbs as groups get smaller. On the group sizes DISC uses, random answering scores 68.7 percent. So DISC's 73.5 percent agreement is not beating chance by more than twenty points. It is beating it by 4.8.

This is the heart of the method, and the part anyone reproducing it is most likely to get wrong. Any statistic defined against a quantity estimated from the same data carries a built-in floor, and comparing it to the baseline you would use for a different statistic flatters your instrument. The fix is not to drop the measure but to work the floor out exactly and report the distance above it.

What the corrected null actually is

The corrected baseline is computed, not simulated, because the problem has a closed form. For one type-pair answered n times under random choice, the count on either side is binomial and the majority count is the larger of that count and its complement: the distribution of max(k, n−k), exactly writable for any n. A respondent's agreement is their total majority count over their total answers, so the null for the whole form is the convolution of those per-pair distributions across every pair it uses, at that form's own group sizes — no seed and no Monte Carlo error, just an exact reference distribution to test the observed figure against.

With that in place you can finally see how much of the agreement is signal. Each form below reports its measured agreement, its chance floor, the lift between them, and how much of the remaining distance to 100 percent it covers:

  • Conflict Style, n = 12: 85.8% against a 75.0% floor, a lift of +10.8, covering 43% of the available headroom
  • 16 Personalities, n = 87: 70.0% against 60.5%, a lift of +9.5, covering 24% of the headroom
  • DISC, n = 54: 73.5% against 68.7%, a lift of +4.8, covering 15% of the headroom
  • Love & Affection Styles, n = 100: 78.7% against 75.0%, a lift of +3.7, covering 15% of the headroom

All four clear p < 0.001 against the exact null, so the signal is real: respondents do answer the same opposition the same way more often than chance forces them to. Headroom is the more honest summary of how strongly, and it is where this turns from self-congratulation into a to-do list — 16 Personalities covering 24 percent of its headroom is the same fact as its shortage of decided axes, restated in another currency. If you have ever finished a type test sitting almost exactly on a boundary, this is the measurement underneath that experience, which why your results change describes from the inside.

It describes the form, not your sitting

The most important thing about this figure is where it may be applied: to the instrument, never to one person's sitting. The temptation runs the other way, because a per-respondent score looks exactly like an attention check — surely someone near the floor was clicking through carelessly? 40 of DISC's 54 sittings cannot individually be told apart from chance. With six pairs and a handful of answers each, one sitting cannot separate a genuinely mixed profile from random clicking, and no amount of clean arithmetic will change that.

There is a deeper reason not to try: inconsistency here is often the correct answer. Someone genuinely torn between two options should split those choices. A person who leans assertive in a crisis and steady in a planning meeting is reporting something true that the form is too coarse to hold. Reading a low individual score as a carelessness flag would punish exactly the people the instrument is least equipped to describe. These are self-understanding instruments, not clinical or diagnostic tools, and that cuts both ways: it limits what a result can claim about you, and what a low number may accuse you of. What personality tests actually measure has the broader version of that boundary.

What this is and isn't

None of this replaces test–retest. A second sitting weeks later measures something this cannot: whether the result is stable across time, mood and circumstance, which is what most people mean by reliable. Repeated-pair agreement is narrower — whether you answer the same opposition the same way within one sitting. It is, precisely, the most such a form can honestly say about itself without a retest, in a slot that otherwise holds nothing. It is not an IRT calibration, a differential-item-functioning analysis, or a convergent-validity study against an outside instrument; none of those have been run here.

The two forms covering only 15 percent of their headroom can earn more of it, and the mechanism is not mysterious: sharper items on the pairs that scatter most, and more distinct pairings so each carries less of the load. How each form here is scored, and what is and isn't known about it, is written up in the methodology notes.