Practice The recall test Four grades
The Companion Notebook emblem The Companion Notebook Essays on the characters who travel with you

Written out in full so that anyone can repeat it, disagree with it, or tell us where it is broken.

Practice

The recall test, step by step

Everything this journal claims rests on one procedure, and the procedure is simple enough to be suspicious of. Here it is in full, including the four places where we think it is weak.

Nothing here requires special equipment. If you want to run this on your own year-old games, you need a pack of index cards, a pen, and the discipline not to look anything up.

Stage one: play and take ordinary notes

Each of us plays the game normally, at whatever pace suits, on whatever difficulty we would have chosen anyway. We take the sort of notes anyone might take — observations, complaints, screenshots — and we do not try to note companions specifically, because knowing what the test will be about would contaminate it.

The notes are then sealed. In practice that means a single document per game, dated, closed, and not reopened for a year. We do not reread them, quote from them, or discuss them.

Stage two: wait twelve months

No contact with the game. No replaying, no watching video of it, no reading about it. If a game comes up in conversation with someone outside the journal we note that it happened, and if the conversation went into any detail we disqualify that game from the volume. This has happened twice.

Stage three: the cards, from memory only

At the twelve-month mark, each of us sits down with blank index cards and writes, for each game, one card per companion we can recall. On each card, in this order:

Nothing may be looked up, including the game's own store page. If a name arrives ten minutes later it goes on the card with a mark showing it was late, and late names are graded as partial rather than full.

Stage four: open the notes and compare

Only now do the sealed notes come out. The comparison produces three things: a recall grade per companion, a list of everything in the notes that did not survive, and the discrepancies in any counted figures. That last category is where the more interesting findings have come from — the twelfth pause in essay one exists only because the count was wrong by one and the missing instance turned out to be the only scripted one.

The four grades

GradeRequiresVolume four count
FullName, behaviour and one specific scene, unprompted and immediate14
PartialBehaviour or scene retained; name absent, wrong or late7
Role onlyFunction retained and nothing else: the medic, the driver6
NoneNo trace, including after a written prompt is offered4

Stage five: the essay, and the veto

An essay is drafted by whichever of us has the most to say, and the other two check it against their own cards. Any claim that only one set of cards supports is either removed or explicitly attributed to one reader. Any companion where our three sets of cards flatly contradict each other is dropped, and one of the three unwritten games in volume four was dropped for exactly that reason.

Where this method is weak

Four problems, in descending order of how much they worry us.

The sample is tiny. Nine games and thirty-one companions per volume, three readers. Any pattern we report should be read as a hypothesis that survived one small test, not as a finding. When we write that eleven of fourteen full recalls had player-directed habits, that is eleven cases.

We know what we are testing. After four volumes we cannot un-know that this journal is about companions, and it is likely that we now attend to them more closely in play than an ordinary player would. This inflates recall across the board and we have no clean way to correct for it.

Selection is not blind. We choose the nine games, and we choose games with companions in them. Volume five deliberately selects for inward habits to test one clause of the argument, which is good practice for that clause and bad practice for everything else.

Counting is unreliable. A figure recalled from memory a year later, like eleven pauses, feels precise and is not. We report the discrepancy with the notes every time for that reason, and readers should treat every count in this journal as approximate even where it happens to be nearly right.

What we would change if we had more people

With ten readers rather than three, the obvious improvement is a control group: half the readers told the test is about companions, half told it is about level design, cards written blind. That would separate the attention effect from the memory effect and would probably shrink several of the numbers in volume four.

We do not have ten readers, and we would rather run a small honest test with its weaknesses printed than a large one we cannot actually staff. If you run something similar and get different results, we would consider that more useful than agreement.