Qualitative Coding and Analysis
Qualitative work is dismissed when it looks like somebody read some interviews and reported what struck them. The dismissal is usually unfair and almost always unanswerable, because the author cannot show how they got from thirty transcripts to six themes. There is no equivalent of a regression table to point at, and in its absence a reader who is inclined to doubt the findings has nothing to check and no reason to stop doubting.
The remedy is not to imitate quantitative work by counting things that should not be counted. It is to make each step visible: how the codes were arrived at, how consistently they were applied and by whom, how themes were built from codes rather than asserted over them, which cases contradicted each theme and what that did to the claim, and why these quotations rather than others. Every one of those is answerable, and answering them converts an impression into evidence.
The failure has a predictable cost. In a thesis it is the examiner's question about how the themes were derived, asked in the first ten minutes, which sets the tone of everything after it. In a journal it is a reviewer who writes that the analysis is descriptive and the themes are not grounded, which is a rejection dressed as a comment. And in mixed methods work it is the quiet demotion of the qualitative strand to illustration, where it stops carrying any inferential weight and becomes decoration between the tables.
When to use this, and when not to
Use it for any analysis of textual or observational material that will produce a claim: semi-structured interviews, focus groups, open-ended survey responses, policy documents, meeting minutes, media coverage, field notes, or free-text fields in an administrative system. Use it whether the study is wholly qualitative or the qualitative strand of a mixed design.
Use it when the analysis is already written and has been challenged. The method runs retrospectively: the codebook can be reconstructed and applied properly, agreement can be assessed on a sample, and negative cases can be searched for. Sometimes the reconstruction confirms the original themes, which is a strong result. Sometimes it does not.
Do not use it to design the data collection. The interview guide, the sampling strategy and the recruitment route belong to research-design for the logic and research-ethics-and-data-protection for the consent, anonymisation and access. Do not use it to write closed survey items, which is survey-and-instrument-design, although this skill is what analyses the open-ended items that instrument collects.
Do not use it for a review of published studies, which is systematic-review-protocol, even though the extraction and synthesis there share some of the same discipline. Do not use it to interpret coefficients or to write the results section of a quantitative paper, which is results-writing. Do not use it as the ethics and anonymisation authority: what can be quoted, what has to be redacted and what consent permits are decided by research-ethics-and-data-protection, and this skill takes those constraints as given.
What you need before starting
The research question, in the form the qualitative material could answer. Qualitative data answers questions about meaning, process, mechanism and variation, and does not answer questions about magnitude or population prevalence. Missing: write the question first, and where it is really a prevalence question, say so, because no amount of coding will convert twenty-six interviews into a population estimate.
The full corpus, in a stable form, with a fixed identifier per source. Missing: analyse what exists but freeze the set before coding starts, because material arriving mid-analysis changes what the earlier coding would have been and there is no honest way to reconcile that afterwards except recoding.
Transcripts checked against the recordings, or a stated transcription standard. Transcription conventions determine what is analysable: whether hesitation, overlap, pause length and non-verbal sounds are preserved. Missing: state the standard used, check a sample of transcripts against audio, and note the checking rate.
The approach and the level of interpretation, decided rather than drifted into. Missing: decide now, using the rules in step 1, because a corpus half coded inductively and half deductively produces a codebook nobody can apply.
A second coder, or an explicit decision to proceed without one. Missing: proceed alone, disclose it, have a second reader audit a sample of the coding, and write a reflexive account of what you brought to the reading. Neither substitutes for double coding; both are better than silence.
Whatever tool will hold the coding. Dedicated qualitative analysis software helps considerably with retrieval and with the audit trail, but nothing in this method requires it. Missing: use a spreadsheet with one row per coded extract holding source, location, code, and the extract itself, plus a document per code holding its definition. That is slower for retrieval and identical in rigour.
The anonymisation rules and what consent permits. Missing: assume the strictest reading, code with pseudonyms from the start, and settle the question before any quotation is drafted, since quotations are where identification usually leaks.
Whether prevalence claims will be made. This decides whether double coding is required, and it has to be decided before coding rather than after. Missing: assume yes, because a paper almost always ends up wanting to say that most participants described something.
The method
Decide the approach and state it. Inductive coding, where codes are built from the material, suits an area without an established account and produces categories nobody imposed. Deductive coding, where codes come from a theory or an existing framework, suits testing whether an established account holds in this setting. Hybrid, a deductive frame with room for emergent codes, is the most common in practice and is entirely legitimate when stated. The rule for choosing: if you can name the framework you would code against and would be willing to report that it failed, code deductively; if naming it would mean inventing it, code inductively. Then state the level of interpretation separately, because these are two different decisions: are you reporting what participants said, or interpreting what it means in terms they would not use? Both are valid. Conflating them is what makes a section feel slippery, since the reader cannot tell whether a claim is a summary or an argument.
Read the whole corpus once before coding any of it. Take notes, do not code. The purpose is to know the shape of the material so that the codebook built on the first five transcripts is not built on the five least typical. Where the corpus is too large to read entirely, read a purposive spread across whatever dimensions the sampling varied on.
Build the codebook on a subset, and write the exclusion rule for every code. Three to five transcripts, or a comparable slice of documents. For each code: a short name, a one-sentence definition, an inclusion rule saying what belongs, an exclusion rule saying what looks like it belongs and does not, and an anchor example drawn from the data. The exclusion rule is what makes a codebook usable by a second person, and writing it forces a boundary to be decided rather than felt. The judgement call is granularity, and the rule is that a code should be applicable to material from more than one source and should distinguish something that matters to the question; a code applied once is a note, and a code applied to everything is a topic heading.
Keep the structure flat enough to use. Sixty codes across three levels usually means coding has become description, and it produces a scheme where two coders cannot agree because they are choosing between siblings that overlap. The working limit for a two-person team is roughly fifteen to thirty-five codes in at most two levels. Where more seem necessary, the usual cause is that a dimension has been coded as many codes rather than as one code with attributes.
Pilot the codebook with the second coder before coding the corpus. Both code the same transcript independently, then compare, extract by extract. Disagreement at this point is information about the codebook, not about the coders. Revise the definitions, particularly the exclusion rules, and repeat on a second transcript. The rule: do not begin full coding until a pilot round produces agreement you would be willing to report.
Code the whole corpus, in a deliberately mixed order. Coding only the parts that look interesting guarantees finding what you expected. Mixing the order, rather than working through by date or by recruitment sequence, prevents the drift where later transcripts are coded to a subtly different standard than earlier ones. Expect the codebook to change: version it, date each version, record what changed and why, and when a code changes materially, recode the earlier material rather than leaving two standards in the same dataset.
Write analytic memos while coding, not afterwards. A short note whenever something is noticed: a possible relationship between codes, a case that does not fit, a moment where the code was applied reluctantly. These memos are where themes actually come from, and they cannot be reconstructed later because the noticing happens once. They are also the most convincing part of an audit trail.
Measure agreement where prevalence claims will be made, with a statistic that accounts for chance. Two people code the same material independently, and disagreement is reported with a chance-corrected coefficient rather than as raw percentage agreement, which overstates consistency because it counts agreement that chance alone would produce. Which coefficient:
Coefficient Use it when Notes Cohen's kappa Exactly two coders, nominal codes, both coding every unit in the overlap The default for interview and document coding with a two-person team. Weighted kappa where the codes are ordered Fleiss's kappa Three or more coders, nominal codes, every unit coded by the same number of coders Generalises Cohen's kappa; do not use it for two coders Krippendorff's alpha Any number of coders, missing data, unequal overlap, or ordinal, interval or ratio codes The most general and the safest default when the coding design is untidy. Reported by content analysis journals as standard Percentage agreement Always, alongside one of the above, never alone Descriptively useful and easy to read, but it counts chance agreement and is uninterpretable by itself Gwet's AC1 Prevalence is very skewed and kappa collapses despite high agreement The kappa paradox: 95 percent agreement with kappa near zero on a rare code. Report it with the raw prevalence so the reader sees why The thresholds this file applies, which are Krippendorff's and are the ones most widely used across coefficients: 0.80 and above supports firm claims; 0.667 to 0.80 supports tentative claims and must be described as tentative; below 0.667 is not reportable and means the codebook is repaired and the material recoded, not that the number is presented with an apology. Report the statistic per code as well as overall, because a respectable overall figure routinely hides one code on which the coders disagree constantly. A per-code figure below 0.667 sitting under an acceptable overall figure is reported with the reason, not dropped. Resolve disagreements by discussion and record what the discussion changed. The rule: recurring disagreement on one code means the definition is wrong, not that a coder is careless, so revise the definition and recode that code across everything.
Where the analysis makes no prevalence claim and the tradition is explicitly interpretive, agreement statistics may be the wrong instrument entirely; the substitutes are named in the adaptation section below, and the choice is stated rather than left silent.
Build themes from codes, then test them. A theme is not a code with a longer name. It is a pattern that holds across the material and does work in the argument, and it should be expressible as a claim that could be wrong. Build it by grouping codes, then go back to the raw extracts under those codes and check the grouping holds, because groupings made from code names rather than from data are how themes drift away from the material. Then attack it: search deliberately for extracts that contradict it. A theme that has never been challenged is an impression with a name.
Report negative cases and what they do to the claim. They bound it, refine it, or break it, and all three are results. The rule: every theme carries at least one sentence about what did not fit, and where nothing did not fit, say that explicitly rather than leaving it ambiguous, because a reader assumes silence means nobody looked.
Match prevalence language to what the data supports. "Most participants" is a claim about counts and needs the counts, reported as a number out of the sample rather than as a percentage of twenty-six people. "Several participants described" is honest where a count would imply a precision the sampling cannot support. Never report percentages from a purposive sample as though they estimate anything beyond the sample. Where a count is reported, be clear whether it counts participants or extracts, since one voluble participant can produce eleven extracts.
Select quotations by a stated rule. Typical rules: the clearest statement of a common position; the case that shows the boundary of a theme; a negative case; the extract that shows a process rather than a conclusion. Say which rule applies to which quotation, or at least state the rules in the methods. Selecting for eloquence is what produces a quotation set drawn from four articulate participants out of thirty, and readers notice the identifiers repeating.
Present quotations honestly. Pseudonymous identifiers, with the characteristics that matter for interpretation attached. Enough surrounding context that the meaning is not created by the cropping. Every edit marked, including removed hesitation, and never edited for fluency in a way that changes register, because smoothing a participant's speech is a small act of misrepresentation that accumulates across a paper. Check the distribution of your quotations across participants before submitting, and if fewer than half your sources are quoted anywhere, ask why.
Assess saturation only if you claim it, and show the evidence. If the claim is made, say how it was judged: at what point new material stopped producing new codes, and how much material was collected after that point to confirm it. A saturation claim with no evidence is a convention rather than a finding, and reviewers increasingly say so.
Assemble the audit trail as you go. Codebook versions with dates, coding decisions, the agreement statistics, the memos, and the notes on how themes were formed. The test is that any sentence in the finished paper can be traced back to coded extracts and from there to the transcripts. Where extracts can be shared without identifying anyone, prepare them for sharing; where they cannot, say why and preserve the trail internally.
Reflexivity, written as a claim rather than a confession
A reflexivity statement that lists the researcher's demographic characteristics and stops there does nothing. A useful one names the specific way the researcher's position plausibly shaped this analysis and says what was done about it.
Three questions produce it. What did participants likely assume about you, and which topics would that make harder to raise? Which finding were you hoping for before you started? Where in the coding did you feel resistance, and what does the memo from that moment say?
Then the mitigation: who read the coding independently, which alternative reading was tested and rejected, and on what evidence. The statement belongs in the methods, at about a paragraph, and it should contain at least one thing that is uncomfortable to write.
Worked example
Situation. A researcher was studying why a new central procurement rule, introduced across a national health system to standardise the purchase of medical consumables, was being applied in some regional units and quietly bypassed in others. Thirty-one semi-structured interviews had been conducted with procurement officers, clinical leads and finance managers across nine units, averaging fifty-two minutes. A colleague was available to double code for about six days in total. The paper was aimed at a public administration journal whose reviewers would expect a stated analytic approach.
Task. An account of the bypassing that explained variation across units, defensible against a reviewer who suspects the themes were chosen to fit an argument, within ten weeks.
Action. The approach was declared hybrid: a deductive frame drawn from an established account of policy implementation, plus room for emergent codes, and the level of interpretation stated as interpretive rather than descriptive, since the object was to explain behaviour that participants did not describe as bypassing.
The wrong turn came early and cost about two weeks. The first codebook was fully deductive, built from the implementation framework, with fourteen codes across four framework categories. It was applied to six transcripts and it worked, in the sense that every extract found a code. The problem surfaced in the memos: three separate notes recorded discomfort about extracts that were being coded as resource constraint but which were really about something else, namely that officers were protecting long-standing relationships with local suppliers who delivered at short notice during shortages. The framework had no place for that, so it was being absorbed into the nearest available category, and it was the most interesting thing in the material.
The codebook was rebuilt as hybrid, with the framework categories retained and four emergent codes added, of which the supplier relationship code became central. Everything already coded was recoded, which was the two-week cost. The lesson recorded in the audit trail was that a purely deductive scheme cannot signal its own inadequacy, since it always finds a home for every extract, and only the memos revealed the problem.
The final codebook held 24 codes in two levels. Pilot double coding on one transcript gave a Krippendorff's alpha of 0.58, chosen over Cohen's kappa because the overlap was uneven across transcripts. That was below the 0.667 floor and therefore not reportable at all, tentatively or otherwise. Disagreement was concentrated almost entirely on two codes whose boundary was unclear: informal workaround and local discretion. The two were collapsed into one code with an attribute distinguishing whether the action was sanctioned by a unit manager, which was the distinction that actually mattered. A second pilot on a different transcript gave 0.81, above the 0.80 line, and the full corpus was then double coded on a third of transcripts, stratified across units, with an overall alpha of 0.79 and a per-code table reported in the appendix. At 0.79 the paper described its prevalence statements as tentative, which is what the 0.667 to 0.80 band requires. The lowest per-code figure was 0.64 on a code about perceived clinical risk, below the floor, which was reported rather than hidden with a sentence about why that code is harder to apply and a note that no prevalence claim rests on it alone.
Three themes were built from the codes. The central one held that bypassing was concentrated where a unit had an established local supplier who had performed during a past shortage, and that the rule was experienced as removing an insurance policy rather than as an administrative burden. Negative case search found two units that fitted the supplier condition and complied anyway. Both had had a recent audit finding, which bounded the theme rather than breaking it, and that boundary became a substantive part of the paper.
Prevalence was reported as counts: the supplier-relationship account appeared in interviews with 19 of 31 participants across 7 of 9 units. Quotations were selected by three declared rules, and a check of the distribution showed 23 of the 31 participants quoted at least once, with no participant quoted more than three times.
Result. The paper reported three themes, one bounding condition, and two negative cases with an explanation. The agreement table and the codebook went into the appendix, and the reviewers, one of whom was explicitly sceptical of interview-based implementation studies, did not raise a methods objection. The revision they did ask for was about the framework's fit, which was the right argument to be having.
Total analysis time was about nine weeks. The recoding cost two of them, and the memo discipline that made the recoding necessary is what saved the paper.
A second scenario, where it goes differently
The same researcher later had to analyse 4,180 open-ended responses to a single survey question asking staff what one thing would most improve procurement in their unit. Most responses were under twenty words.
Almost every parameter changed. There is no interview context to interpret against, so the level of interpretation moved to descriptive: the analysis reports what was said, and where meaning is ambiguous, the ambiguity is recorded rather than resolved. The unit of coding became the whole response rather than a passage, with a rule permitting up to two codes per response and a record of how often that happened.
The codebook was built on a random sample of 300 responses rather than on the first 300, because early responders differ. Coding was done by two people on a random 15 percent overlap rather than on a third, since the volume made a third impractical and short responses make agreement easier to achieve. Agreement was 0.86 by Krippendorff's alpha, with 91 percent raw agreement reported beside it, which reflects the simplicity of the material rather than superior work, and the paper said so.
Prevalence claims are legitimate here in a way they are not with interviews, because the responses come from a defined sample with a known response rate, so code frequencies were reported as counts and shares with the response rate stated alongside. No saturation claim was made or needed. Quotations were selected to illustrate each code's range rather than to show boundaries, and because responses are short, more of them were shown, twelve per major code in an appendix table.
What did not change: the codebook still had exclusion rules, the whole corpus was still coded, disagreement still drove definition changes rather than coder correction, and the audit trail still ran from each reported figure to the coded responses.
Output
The codebook, versioned and dated:
| Code | Definition | Include when | Exclude when | Anchor example (source, location) | Version added | Notes |
The coding record, one row per coded extract:
| Extract ID | Source ID | Location in source | Code | Coder | Date | Extract text |
The agreement report:
| Code | Extracts double coded | Coefficient and value | Percentage agreement | Main source of disagreement | Resolution |
with the overall figure, which coefficient was used and why that one given the number of coders and the code type, the threshold being applied and its source, the proportion of the corpus double coded, and how disagreements were resolved. Any per-code figure below the threshold is listed with an explanation rather than omitted.
The theme table:
| Theme | Codes it draws on | Claim, in one sentence | Sources contributing | Participants (n of N) | Negative cases and what they do | Quotations used |
The quotation ledger, which is an internal document and the thing that prevents cherry-picking:
| Quotation ID | Participant pseudonym | Characteristics shown | Theme | Selection rule applied | Edits made | Cleared for publication |
The methods paragraph, which should contain, in order: the approach and level of interpretation, the corpus and how it was prepared, how the codebook was built and on what subset, the number of coders and the proportion double coded, the agreement statistic with its value, how themes were built, how negative cases were sought, the quotation selection rule, and where the audit trail is held.
The audit trail, as a file list: codebook versions, coding record, memos with dates, agreement calculations, theme development notes, and the mapping from paper claims to extract IDs.
Failure modes
Themes announced rather than built. Recognise it when the themes could have been written before the interviews, and when no code list connects them to the data. Fix by coding properly and rebuilding, which usually changes at least one theme.
Coding only the interesting transcripts. Recognise it by an uneven number of extracts per source, concentrated in the sources the researcher remembers. Fix by coding everything, in mixed order.
Percentage agreement reported as reliability. Recognise it by a figure above ninety percent with no coefficient named. Fix by computing kappa or alpha, which will be lower and honest, and report the percentage beside it rather than instead of it.
A high overall agreement figure hiding one broken code. Recognise it by the absence of a per-code table. Fix by reporting per-code figures and revising the definitions of the worst.
The codebook that grew to sixty codes. Recognise it when codes cannot be told apart in a sentence. Fix by collapsing overlapping codes and converting dimensions into attributes.
Themes that are topics. Recognise it when a theme is a noun phrase with no claim in it, such as "communication". Fix by rewriting each theme as a sentence that could be false.
Negative cases not sought. Recognise it by their absence, which is near-universal in weak qualitative sections. Fix by searching for them explicitly and reporting what was found, including nothing.
One articulate participant carrying an argument. Recognise it by counting quotations per pseudonym. Fix by finding corroborating extracts or by weakening the claim.
Quotations edited for fluency. Recognise it when the transcript and the quoted text differ in register. Fix by restoring the original and marking every removal.
Counts presented as percentages of a small purposive sample. Recognise it when a paper says 42 percent of a sample of twenty-six. Fix by reporting counts.
Saturation asserted. Recognise it when the word appears with no evidence behind it. Fix by showing when new codes stopped appearing, or by dropping the claim, which costs nothing.
Coding drift over a long project. Recognise it by comparing early and late coding of the same code. Fix by recoding an early sample against the current codebook and reporting the check.
Edge cases
A single coder and no possibility of a second. Disclose it in the methods, have a colleague independently code two transcripts as an audit and report the comparison qualitatively, write a substantive reflexivity statement, and avoid prevalence claims that rest on consistent application. This is a common and acceptable position; concealing it is not.
Material in a language the analyst does not read fluently. Code in the original language wherever possible, since coding a translation codes the translator's choices. Where translation is unavoidable, have a bilingual colleague check the coding of a sample against the original, and quote in both languages.
Documents rather than speech. Genre and authorship matter: a policy document is a negotiated artefact, not a person's account. Code with the document's purpose and audience recorded as attributes, and never treat a document's silence as evidence of an absence.
Very short responses, such as open survey fields. Code the whole response, cap codes per response, build the codebook on a random rather than a first sample, and move the level of interpretation towards descriptive.
Focus groups. The unit is contested: an extract is produced in interaction and may not represent an individual view. Code interaction as well as content where the question warrants it, and never count participants in a focus group as though they were independent respondents.
A corpus too large to code by hand. Sample purposively for depth and code that sample fully, or use computational assistance for retrieval and structure with a human coding a validated sample. Where any automated classification is used, report the tool, its settings, and the human-validated error rate on a random sample of its output; treat it as a search aid, not a coder.
The material contains disclosure of harm or wrongdoing. Stop and follow the study's disclosure procedure before continuing the analysis. That procedure belongs to research-ethics-and-data-protection and should have been written before fielding.
Anonymisation makes a quotation useless. Where the detail that makes an extract meaningful is also what identifies the speaker, do not publish it. Paraphrase without quotation marks, state that it is a paraphrase, or describe the pattern and cite the extract ID internally.
Reanalysing somebody else's coded data. Treat their codebook as data about their analysis rather than as a given. Recode a sample independently before deciding whether to adopt it, and report that check.
Quality bar
- The approach, the level of interpretation and the unit of coding are stated in the methods.
- The codebook carries a definition, an inclusion rule, an exclusion rule and an anchor example for every code, and is versioned with dates.
- The whole corpus was coded, in mixed order, not a selection.
- Agreement is reported with a chance-corrected statistic, overall and per code, or single coding is disclosed and compensated.
- Every theme is stated as a claim that could be false, and every theme reports what did not fit.
- Quotations are selected by a stated rule, edits are marked, and the spread of quotations across participants has been checked.
- Prevalence language matches the design: counts for counts, and no percentages from small purposive samples.
- An audit trail runs from any claim in the paper to a coded extract to a source, and the file list exists.
Adapting this to your context
This file leans post-positivist: codebooks, double coding, agreement coefficients, prevalence language. That fits framework analysis, content analysis and the qualitative strand of mixed methods. It does not fit every tradition, and what does not fit should be swapped, not forced.
- Agreement statistics. They assume the codebook should be applied consistently by different people. Reflexive thematic analysis, constructivist grounded theory, narrative and discourse analysis reject that on principle, and reporting kappa there signals you have misread the tradition. Substitute a reflexivity statement, an audit trail, member checking or peer debriefing, and say which and why.
- The software. None is named here. NVivo, MAXQDA and ATLAS.ti all handle codebooks, extracts, memos and inter-rater statistics; Dedoose suits a distributed team. A spreadsheet works below roughly thirty transcripts if the extract register keeps its IDs.
- Two coders and a third of the corpus. A social science default. Health services research often double codes everything; large open-text corpora use 10 to 15 percent. State the proportion and why.
- Stopping. The file assumes a fixed corpus. If you are still collecting, use information power or a saturation rule fixed before fieldwork, not "until nothing new appeared".
- What not to change. Code the whole corpus, not the interesting parts, and keep an audit trail that runs from any sentence in the paper back to a coded extract.
Related skills
research-design sets the sampling and the interview logic that produce this material, and research-ethics-and-data-protection governs consent, anonymisation and what may be quoted. survey-and-instrument-design writes the closed items whose open-ended companions this skill codes. systematic-review-protocol applies a related discipline to published studies rather than to primary material. data-profiling-and-cleaning handles the structured half of a mixed methods dataset, and descriptive-statistics-tables describes the sample this analysis draws on. results-writing turns the theme table into the paper's results section, and discussion-and-conclusion sets how far the claims may be pushed. peer-review-simulator is the useful last step, since qualitative sections attract a predictable set of objections that this method is built to answer.