Skip to main content

We Keep a Record of Who Read Every Test Question

7 min readMy Path Research

Pick any free personality test online and try to find out who checked its questions. Not who wrote the landing page, not which academic framework the test borrows its name from — who read item fourteen, in the language you are reading it in, and confirmed that it asks what the scale it sits on claims to measure. You will almost never find an answer, and usually not because someone is hiding it. Nobody did the reading, and no record exists either way, so there is nothing to disclose. We keep that record instead, and this article is about what it contains — including the part that makes it less impressive than the headline number sounds, which is exactly why it is worth publishing.

What is actually in the table

There is a table in our database whose only job is to remember that a test question was read and checked. It currently holds 3,744 recorded reads, covering 624 items across 6 languages in 14 test banks. Every row points at one item in one language, and records the date and the identity of whoever read it. None of those rows are stale — every one still matches the exact wording of the item it approved, which is the subject of a later section. That is the whole artefact: not a badge, not an endorsement, just a durable note saying this sentence, in this language, was looked at by someone asked to check whether it belongs on the scale it sits on.

Who did the reading, stated plainly

The reader was an AI model, working at the instruction of the platform owner. Every recorded row stores that fact in a field naming the reviewer, so the provenance travels with the record rather than sitting in a footnote. This is not a native-speaker sign-off. It is not external expert adjudication. It is not review by a human subject-matter expert. If you hoped the number 3,744 meant a panel of psychologists had signed off on our item pool, it does not, and no phrasing on our side should be allowed to leave that impression. What it means is narrower: 3,744 times, a capable reader was pointed at one question in one language and asked whether the wording does the job, and the answer was written down where it can be checked and contradicted later.

Why labelling it beats blurring it

The temptation, once you build a review register, is to describe it in language that invites the most flattering interpretation — "reviewed," "validated," "quality-checked," all perfectly true and all quietly inviting you to picture a committee. That trade is a bad one, and not only on ethical grounds. A review record that hides what performed the review is worth less than no record at all, because it converts a piece of genuine evidence into a marketing claim you cannot audit. If the reviewer field says what it says, you can weigh it yourself: you can decide that an AI read of an item's content coverage is meaningful for some purposes and insufficient for others, and you would be right on both counts. If it says only "reviewed," you have learned nothing except that we wanted you reassured. Strip the provenance out and the number becomes decoration.

The hash is the clever part

Here is the mechanism that makes the register more than a spreadsheet of good intentions. Each row is hashed against the item text that was read. If anyone later edits that question's wording — a rephrased clause, a softened adverb, a single comma — the stored hash no longer matches the live text, the row stops counting, and the item is reported as needing review again. The approval cannot silently outlive the sentence it approved. This is deliberate, because a review marker that cannot go stale is a lie waiting to happen: it will eventually sit beside text that nobody in its history ever saw, still radiating confidence. Ours cannot do that. When we say zero rows are stale, we are not saying nobody has edited an item since; we are saying the register has already dropped every row whose text moved underneath it.

It costs us something, which is the point

The uncomfortable consequence follows directly. A copy edit stales its own review row. If we improve a question's wording tomorrow — make it clearer, remove an ambiguity a reader complained about — our measured review coverage goes down until that item is read again. Improving the instrument lowers the number we would most like to report. That is intended, and we would rather absorb the cost than own a metric that only ever climbs. A quality number that cannot fall is not a measurement; it is a scoreboard. Our coverage figure is a snapshot of a moving item pool, and a dip in it after a content pass is evidence the machinery is connected to something real.

What reading items actually catches

It is fair to ask whether reading questions one by one finds anything a spell-checker would miss. It does, and the examples are not subtle. On our Social Skills bank, the French translation was asking a materially different question from the English on 35 of 36 items — not mistranslated words, but items that had drifted into probing something adjacent, which quietly makes the French and English versions of one scale non-comparable. On our Character Strengths bank, 288 items were written in the second person where the instrument requires the first — a grammatical detail that changes the respondent's task: rating whether a description fits you is not the same as endorsing a statement about yourself. Neither problem shows up in a scoring test, a unit test or a user complaint. They show up when somebody reads the items.

Five banks where the work is done and the note is not

Because the register counts only filed rows, our own scorecard is currently unflattering in a slightly absurd way. Five test banks have their narrative review work finished and deployed, and still score below the top mark purely because the note recording the review has not been filed:

  • Grit, DISC, Social Skills, Conflict Style, and Love & Affection Styles.

The work exists; the artefact does not. We could close that gap by writing prose claiming the reviews happened, which is what the register exists to prevent. So the number stays where it is until the rows are written, and our published coverage understates what has been done — the opposite of the usual failure mode, and a more comfortable one to live with.

What this covers, and what it does not

There is a reason item review is the strand we could build first. Content validity — whether a question actually covers what it claims to cover — is the one part of test quality that does not need respondents; it can be established before a single person takes the test, by reading. Nearly everything else needs people: reliability estimates need answer data, norms need a reference population, retest stability needs the same people twice. The rest is genuinely absent, and worth saying in one sentence: no external expert review, no native-speaker sign-off, no differential-item-functioning analysis, and no test–retest study has been carried out on any of our test banks. For the general picture of which strands matter for which purpose, what personality tests actually measure is the wider frame, and free accurate personality tests online covers what "accurate" can and cannot mean for something you take for free in ten minutes. Our tests, Grit included, are instruments for self-understanding rather than clinical, diagnostic or medical tools, and the register does not move that boundary.

What the register gives you is narrow and checkable: for any question, we can tell you whether it was read, in which language, when, and by whom — and the record will tell you itself when it has gone out of date. The full account of how our tests are built, scored, and checked lives at methodology.

We use essential cookies to keep you signed in and optional analytics cookies to improve the product. See our Privacy Policy.