← The Signal
Qualitative Research

Your Codebook Is an Instrument Too

A survey scale gets tested for reliability. The codebook behind every qualitative finding usually does not. Codebook reliability, intercoder agreement, and the audit trail that keeps qualitative coding consistency defensible.

On a Tuesday afternoon in week three, two coders open the same six interviews and compare what they marked. Passage 41 in interview 14 is "gatekeeping" to one of them and "institutional delay" to the other. Neither reading is wrong. The codebook defines gatekeeping as a person or policy that blocks access, and a three month wait for an evaluation is arguably both. They talk it through for twenty minutes, settle on a rule, and keep going. The rule never gets written down.

By week nine, gatekeeping has quietly absorbed most of what institutional delay was for. Nobody decided that. It happened one twenty minute conversation at a time.

A survey scale would never be treated this way. Alpha, item-total correlations, a line in the methods section about how the instrument performed. The codebook behind a qualitative study does the same job as that scale, turning what people said into something you can make a claim about, and it usually gets no test at all.

The Instrument Nobody Tests

A codebook is a measurement instrument. Each code is a decision rule: this passage counts, that one does not, and here is the boundary between them. When the rule is applied the same way across nineteen interviews and two coders, the findings mean what the chapter says they mean. When the rule moves, the findings still look tidy, and they are quietly describing something that changed underneath them.

Drift shows up in a few recognizable ways. A code widens until it stops excluding anything, which is the gatekeeping problem above. Two codes with adjacent definitions start trading passages depending on who is coding that day. A code invented at interview 12 never gets applied to the eleven that came before it, so its count reflects when you thought of it rather than how often the thing occurred.

Solo coding is not exempt. You in week one and you in week nine are effectively two coders with a poorly documented handoff.

What Agreement Numbers Can and Cannot Tell You

Intercoder agreement is the standard check, and it is worth running. Percent agreement reads easily and inflates easily, since two coders who both apply a code to almost nothing will agree almost all the time. Cohen's kappa corrects for chance agreement and gives you a number that survives a reviewer's question.

Here is the honest limit. High agreement means two people are applying a definition the same way. It does not mean the definition is any good. Two coders can be consistently vague together, and kappa will look excellent while they are. Agreement is necessary and it is not sufficient.

There is also a tradition, common in constructivist and interpretive work, that treats intercoder agreement as the wrong test entirely, because the analysis is understood as interpretive rather than replicable. That position is defensible. It also raises the burden on everything else: the definitions, the memos, and the trail that shows a reader how you got from a transcript to a claim.

Coding on the Mac

Mac software has an old habit worth keeping: put the work in front of the person and stay out of the way. Open the project, open the transcript, code. Nothing to configure first, no browser tab that logs you out at hour two of a coding session.

Apple Silicon settled the practical side. Transcription, search across nineteen transcripts, and agreement calculations all run on the machine in front of you, at speed, with the audio and the consent forms staying exactly where the IRB protocol says they live. Nothing about checking a codebook requires a server.

What Quala Keeps

Quala is our qualitative analysis app for the Mac, and it treats the codebook as a document with a history rather than a list of labels. Each code carries its definition, what it includes, what it excludes, and the example passages that anchor it. Edit a definition in week nine and the earlier version stays in the record with the date it changed.

Every application of a code keeps its own trail: which passage, which coder, when, and the memo written at the moment of the decision. Run a second coder over the same transcripts and agreement comes back per code, not as one summary number. A study with an overall kappa of .78 and one code sitting at .41 has a specific problem in a specific place, and the per-code view says which place.

ReliCheck Intelligence can suggest codes drawn from the codebook you wrote. It never applies one. That line matters more here than almost anywhere else, because an auto-applied code is a finding nobody made. Quala does not think for you. It removes the translation between your question and your evidence, and that is all it removes.

When a committee member asks how coding stayed consistent across nineteen interviews, the answer comes from the record instead of from memory. Definitions as written, when they changed, why they changed, agreement per code, and the memo from that Tuesday afternoon when two people decided what gatekeeping would mean.

The codebook still has to be yours. Testing it is what makes it an instrument instead of a habit.

Quala is available for the Mac at relicheck.com and on the Mac App Store, with a free 30-day trial.