HSCI 841 – Lesson 8

Content Analysis

Qualitative Research Methods & Analysis in Public Health

Learning objectives for this lesson:

  • Trace content analysis from Lasswell's WWII propaganda studies through Berelson's 1952 codification to Krippendorff's contemporary synthesis
  • Distinguish manifest from latent content and explain why most contemporary content analysis blends both
  • Treat codes as variables, the operational move that turns content analysis into a hybrid quantitative/qualitative method
  • Sample text systematically: choose sampling units, recording units, and context units that match the research question
  • Develop a content-analytic coding scheme that is exhaustive, defensibly exclusive (or explicitly multi-coded), and reliably applied
  • Compute Krippendorff's alpha as the standard reliability statistic and interpret its acceptable thresholds
  • Test hypotheses on content-analytic data using chi-squared, distributional comparisons, and trend analysis
  • Apply dictionary-based and computational content analysis to the loneliness corpus, from a Taguette export to a tested subgroup comparison

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Bernard, H. R., Wutich, A., & Ryan, G. W. (2017). Analyzing Qualitative Data: Systematic Approaches (2nd ed.). SAGE. This lesson covers Chapter 11 (pp. 243–268).

Section 1 of 5

What Content Analysis Is: History, Manifest/Latent Content, and Codes as Variables

⏱ Estimated reading time: 30 minutes
Section 1 of 5

What Content Analysis Is

History, manifest vs. latent content, and codes as variables.

Origins

Lasswell and the propaganda studies

Harold Lasswell's wartime work at the Library of Congress operationalised content analysis: counting symbols, tracing their distribution, and drawing inferences about communicative intent.

His framing question still organises the field:

Who says what, to whom, in which channel, with what effect?

Harold Lasswell, political scientist and communication theorist.
Harold Lasswell (1902–1978). Public domain, via Wikimedia Commons.
Codification to synthesis

Berelson (1952) to Krippendorff (2018)

Berelson, 1952

First textbook definition: objective, systematic, quantitative description of manifest content. Deliberately restricted to the literal surface.

Krippendorff, 1980–2018

Expanded to latent meaning; formalised reliability requirements; developed the alpha statistic. Now the canonical reference.

Bernard, Wutich, and Ryan (2017) position content analysis as part of a continuum of text-analytic methods, which is the framing this course uses.

The key distinction

Manifest vs. latent content

Manifest Literal surface of the text Word counts, explicit references Higher inherent reliability "mentions pet" Latent Underlying meaning Requires inference from context Reliability must be demonstrated "grief framed as moral debt" ↔

Contemporary content analysis blends both. The analyst declares which is which and demonstrates reliability for latent codes.

The defining move

Codes as variables

Thematic analysis

A code is a label attached to a passage. Its job is to organise interpretation for one study.

Content analysis

A code is a variable that takes a value (0/1, or a count) for every recording unit. It is a column in a data matrix.

Once codes are variables, frequency tables, cross-tabulations, and chi-squared tests follow naturally. Both the qualitative and quantitative halves are load-bearing.

Carry forward

When content analysis is the right choice

  • Your question asks how often, or whether subgroups differ in what they say.
  • The codes-as-variables move is what separates content analysis from thematic work.
  • Both qualitative and quantitative halves are necessary; neither is decoration.

A later section covers the design decisions that make a content analysis defensible.

Introduction and Overview

Content analysis is the oldest systematic method for analysing text in the social sciences, and it is also the bridge between the qualitative and quantitative traditions of this course. Where the thematic analysis you learned earlier in the course stops once you have a defensible set of themes, content analysis keeps going: it counts the themes, distributes the counts across subgroups, and tests whether the differences are larger than chance. The method is qualitative in its first move (identifying what the relevant categories are) and quantitative in its second (counting their occurrence). It is the technique you reach for when your research question asks both what kinds of loneliness do people describe and how often, and does that distribution differ by age, gender, or caregiver status.

This section traces how content analysis became the workhorse it is today. We start with Harold Lasswell's WWII analysis of Axis propaganda, watch Bernard Berelson codify the method in 1952, follow Klaus Krippendorff's modernisation into the canonical 2018 monograph, and end with the contemporary synthesis offered by Bernard, Wutich, and Ryan (2017, Ch. 11). Along the way we will pull apart the distinction between manifest and latent content, the line that separates "what the text literally says" from "what it means", and we will explain the single operational move that turns content analysis into the hybrid method it is: treating codes as variables. By the end of the section you will be able to say what content analysis is and why it is methodologically different from the thematic work you have already done.

Learning Objectives for this section

  • Locate content analysis historically: from Lasswell (1927, 1942) and Berelson (1952) to Krippendorff (2018) and the contemporary synthesis.
  • Distinguish manifest from latent content and recognise that most contemporary content analysis blends both.
  • Explain the operational move that makes content analysis a hybrid method: codes as variables.
  • Identify the kinds of public-health research questions for which content analysis is the right tool.

1.1 A Short History: From Lasswell to Krippendorff

Harold Lasswell’s wartime propaganda analysis at the Library of Congress operationalized content analysis: who says what to whom in which channel with what effect. The first attempt to quantify symbolic content systematically. Many of today’s content-analysis categories descend directly from his coding sheets.

Bernard Berelson’s textbook defined content analysis as 'the objective, systematic, and quantitative description of the manifest content of communication.' This narrow definition dominated for two decades and is still the working definition in much applied communications research.

Klaus Krippendorff’s textbook broadened the field to include latent meaning and explicitly engaged the reliability and validity questions that remain central. Krippendorff’s alpha statistic, developed across multiple editions, is now the standard inter-coder reliability measure for content analysis.

Content analysis today spans manual coding, computational text analysis, supervised machine learning, and now LLM-assisted approaches. All trace back to the Lasswell-Berelson-Krippendorff lineage. The methods have changed; the underlying logic, turning text into countable, comparable data, has not.

The modern history of content analysis begins with Harold D. Lasswell, the political scientist who in 1927 published Propaganda Technique in the World War, the first systematic effort to study communication content as a window onto political intent. Lasswell's question was not "what does this propaganda mean to its readers?" but "what categories of appeal does it use, and in what proportion?" The methodological commitment was to systematic counting of explicit features, mentions of leaders, mentions of enemies, deployment of symbols, rather than impressionistic close reading. The pay-off was that two analysts working independently could produce comparable numbers.

During the Second World War, Lasswell led the Experimental Division for the Study of War-Time Communications at the U.S. Library of Congress. The Division analysed German, Italian, and Japanese propaganda in volume that no individual reader could have made sense of impressionistically. By systematic coding of explicit features, the frequency with which Axis broadcasts mentioned specific Allied generals, the proportion of broadcast time devoted to economic versus military themes, the rise and fall of named enemies week by week, the Division produced intelligence assessments that helped guide both counter-propaganda efforts and broader policy decisions. The work was content analysis as inference about a producer: from the text, back to the propagandist's strategic state of mind. The method got its first major public-policy validation in this period (Lasswell, Lerner, & Pool, 1952).

The post-war codification belongs to Bernard Berelson. His Content Analysis in Communication Research (1952) gave the field its first textbook and its most-cited definition: content analysis is "a research technique for the objective, systematic, and quantitative description of the manifest content of communication" (Berelson, 1952, p. 18). Each of the four adjectives in that definition matters. Objective meant that the analysis was repeatable by other analysts. Systematic meant that the same procedure was applied across the entire sample. Quantitative meant counting, not impression. Manifest meant the literal surface of the text, what was said, rather than what the analyst imagined the author meant. Berelson's stance was, in effect, the high-modernist version of content analysis: positivist, quantitative, and resolutely uninterested in latent meaning.

The next major recodification came from Klaus Krippendorff, whose Content Analysis: An Introduction to Its Methodology (1980, 2004, 2013, and now in its 4th edition, 2018) brought the method into the contemporary social-science mainstream. Krippendorff's most important moves were two. First, he redefined content analysis as "a research technique for making replicable and valid inferences from texts (or other meaningful matter) to the contexts of their use" (Krippendorff, 2018, p. 24), broadening Berelson's "objective description" to "inference about context," which made room for latent content. Second, he gave the field its statistical centre of gravity by developing Krippendorff's alpha, the reliability statistic that has since become the methodological standard for content analysis (we will use it in a later section). Krippendorff's text remains the field's most-cited methodological reference and the standard against which contemporary content-analytic studies are judged.

The synthesis you are reading in this course comes from Bernard, Wutich, and Ryan (2017), who present content analysis as part of a continuum of text-analytic methods rather than as a standalone enterprise. Their stance is that the strict separation Berelson drew between qualitative and quantitative analysis was always more rhetorical than real, and that contemporary content analysis is best understood as a method that blends the two: qualitative judgement determines the coding scheme; quantitative procedures determine how the codes are distributed and whether the differences are reliable. This is the position that Hsieh and Shannon (2005) articulated influentially in the health-research literature when they distinguished conventional, directed, and summative content analysis, a typology that has since structured most qualitative-content-analytic work in nursing, health services research, and public health.

TraditionApproximate datesDefining commitmentWhat you would publish
Lasswell propaganda analysis 1927–1948 Systematic counting of explicit features to infer producer intent Frequency tables of symbols, names, and themes, usually classified documents
Berelson classical content analysis 1952–1970s Objective, systematic, quantitative description of manifest content Tables of code frequencies, with explicit coding rules and percent-agreement reliability
Krippendorff contemporary content analysis 1980–present Replicable and valid inference from text to context; explicit room for latent content; rigorous reliability statistics Code distributions, inferential tests, Krippendorff's alpha as the reliability statistic, explicit sampling logic
Qualitative content analysis (Hsieh & Shannon, 2005) 2000s–present Conventional / directed / summative variants; categories may emerge from data, be imposed from theory, or be derived from word counts Coded extracts plus a frequency table; mixed-method publications in health journals
Bernard/Wutich/Ryan synthesis (this course) 2017 Content analysis as a hybrid method that explicitly turns codes into variables and analyses them with descriptive and inferential statistics A coded corpus, a frequency-by-subgroup table, a defensible inferential test, and an interpretive narrative

Why the history matters for a methods section

A methods section that reports a content analysis positions the work in this lineage. Most contemporary health-research applications cite Krippendorff (2018) for the reliability machinery, Hsieh & Shannon (2005) for the typology, and Bernard, Wutich, and Ryan (2017) or Schreier (2012) for the practical workflow. Knowing where each one sits in the genealogy lets you make defensible methodological choices, for example, whether to treat your codes as mutually exclusive (Berelson) or to allow multi-coding (Krippendorff/Schreier).

1.2 Manifest Versus Latent Content

Key insight - Manifest vs latent is a continuum, not a binary

The classical distinction holds that manifest content is what is literally on the page ('the word vaccination appears') and latent content is what is implied or interpreted ('this passage frames vaccination as a moral obligation'). In practice, every coding decision involves some interpretation, even counting the word 'vaccine' requires deciding whether 'vaccinated' counts. The more latent the code, the more important reliability evidence and explicit decision rules become. A defensible content analysis names where its codes sit on the manifest-to-latent continuum.

A frequently asked methodological question about content analysis is: are you coding what the text literally says, or what it means? Bernard, Wutich, and Ryan's answer, and the answer of the modern field, is that the choice is yours, but it must be explicit.

Manifest content is the literal, surface-level content of the text. If your code is "mentions of pets," a manifest coder counts every appearance of the words pet, dog, cat, Rufus, Marie's cat, parrot, and so on. The advantage is that two manifest coders working from a clear word-list will produce nearly identical counts; the disadvantage is that manifest content misses everything the participant is doing with the text other than literal naming.

Latent content is the underlying meaning that requires interpretation to surface. If your code is "loneliness framed as the cost of having loved," a latent coder reads a passage and decides whether the participant is articulating the idea, regardless of whether the specific words "cost" or "love" appear. Linda's account of Bill's empty chair (P05) is a latent expression of "loneliness-as-residue-of-marriage", she does not use that phrase, but the meaning is unmistakable to a competent reader. Latent coding captures more of the meaning but is harder to do reliably.

Berelson's classical definition (1952) restricted content analysis to manifest content. Krippendorff and the subsequent generation rejected the restriction, on the grounds that inference about context, the contemporary purpose of content analysis, almost always requires reading for latent meaning. The modern compromise is that content analysis routinely codes both, but the analyst must declare which is which and must demonstrate that latent codes can be applied reliably.

FeatureManifest codingLatent coding
Unit examined Words, phrases, named entities Passages, propositions, implicit meanings
Decision procedure Match against a dictionary or word-list Read for meaning; apply rule from codebook
Reliability difficulty Lower, identical word lists give identical counts Higher, requires training and an explicit codebook
Interpretive depth Shallow but defensible Deeper but contestable
Typical reliability target (Krippendorff's alpha) α ≥ 0.80 readily achievable α 0.67–0.80 acceptable for tentative inference; α ≥ 0.80 for definitive claims

A worked example from the loneliness corpus

Consider the question: how often do participants describe loneliness in terms that involve a piece of household furniture?

Manifest version: Count every occurrence of the words chair, couch, sofa, bed, table across all 20 transcripts. Linda mentions "chair" 8 times in P05; Helen mentions "chair" 3 times in P11. Manifest count: straightforward, replicable, and largely uninformative on its own.

Latent version: Code any passage in which a piece of furniture stands in for an absent person or a lost role. Linda (P05) describes Bill's chair as the empty space he used to fill, that is a clear latent instance. Helen (P11) describes her armchair as the place where she does not have anyone to read with, also latent, but with a different propositional structure (her loss is structural, not bereavement-specific). Linda's mention of the kitchen table where she ate dinner with Bill, latent, same theme. The latent code furniture-as-trace-of-absent-other may apply to 9 transcripts (vs. 11 transcripts for the manifest word-list), but the latent code is doing meaningful analytic work while the manifest word-list is not.

In an analysis of the loneliness interviews, you will frequently want both: the manifest word-list as a fast first pass, and the latent code as the substantive analytic move.

1.3 Codes as Variables: The Operational Move

Here is the single move that makes content analysis what it is. In thematic analysis (an earlier module), a code is a label you attach to a passage; the code's job is to organise interpretation. In content analysis, the same code becomes a variable: a column in a data frame, with a value for every unit in the corpus. Once codes are variables, the entire apparatus of descriptive and inferential statistics is available.

The reframing is subtle but transformative. Consider the code loneliness-as-residue-of-marriage. In thematic analysis, the code is a finding: "this is one of the kinds of loneliness participants describe." In content analysis, the code becomes a variable that takes the value 1 for every transcript in which the theme appears and 0 elsewhere. Now you can ask: in what proportion of transcripts does the theme appear? Does that proportion differ between widowed and non-widowed participants? Is the difference larger than chance? Does the proportion grow over the course of the interview, that is, do participants need to talk for a while before they articulate the theme? Each of these is a statistical question that the codes-as-variables move makes available.

Bernard, Wutich, and Ryan are explicit that this move is what places content analysis in the hybrid position it occupies. It is not "qualitative analysis with numbers attached" (which is what Berelson thought it was), nor "a quantitative method applied to text" (which is what nineteenth-century philology was). It is a method in which qualitative judgement is required to define the variables, and quantitative procedures are used to analyse them. Both halves are necessary; neither is optional.

MoveThematic analysis (an earlier module)Content analysis (this module)
What a code is A label organising interpretive material A variable with a value for every unit
What you report The code, with illustrative quotes The code's frequency, distribution, and inferential tests
What you defend That the code names something real in the data That the code is reliably applied AND names something real in the data
How you defend it By showing illustrative passages By showing inter-coder agreement (Krippendorff's alpha) AND illustrative passages
What "more data" buys you Greater confidence that you have saturated the themes Statistical power to detect distributional differences

1.4 When Content Analysis Is the Right Tool

Not every qualitative research question calls for content analysis. The method's distinctive payoff is in questions that involve distributional comparison: how often something appears, whether subgroups differ in their use of a code, whether something changes over time. Where the question is genuinely interpretive, what does this single passage mean, or how does this participant make sense of their experience, thematic analysis (an earlier module), schema analysis (a later module), or narrative analysis (also a later module) will give you more analytic traction than content analysis will.

Concretely, you should reach for content analysis when one or more of the following are true:

  • You have a comparison question. Caregivers versus non-caregivers; younger versus older; immigrant versus native-born; pre-pandemic versus post-pandemic. Content analysis is built for these.
  • Your corpus is large enough that close reading alone would be impractical. Twenty transcripts is on the lower end, you can certainly do content analysis on a corpus this size, and this lesson does, but for corpora of 50, 200, or 5,000 documents, content analysis is one of the few options that scales.
  • You need replicable counts that other researchers can audit. Where policymakers or journal reviewers will demand "how many" rather than "how rich," content analysis is the methodological answer.
  • Your research question is longitudinal. Trends, shifts, before/after comparisons, content analysis is the standard tool for them in the qualitative literature.
  • You are writing for a quantitative audience. Public-health journals routinely ask for the kind of distributional evidence that content analysis produces. Where thematic analysis would be rejected as too impressionistic, content analysis can be persuasive.

Reflection

Pick one candidate theme from a codebook for the loneliness dataset, such as the codebook developed in the themes and codebooks lesson. Now answer two questions about it: (1) Is the theme manifest, latent, or both? (2) If you turned it into a variable, what comparison or distribution would you most want to examine? Be specific, name the subgroups, the cases, or the dimension of variation you would test.

Model answerA strong response is specific to one named theme and gives a clear, testable comparison. Example: "My code loneliness-as-residue-of-marriage is predominantly latent, participants do not use the phrase, but the propositional content is clear in passages where they describe an absent partner. If I turned the code into a variable, I would expect the count to be much higher for the 6 widowed participants in the corpus than for the 4 single-and-never-partnered participants. The interesting comparison is less the obvious widow/non-widow contrast than whether the never-partnered participants articulate any functionally equivalent latent code (perhaps loneliness-as-residue-of-roads-not-taken) that fills the same analytic slot." A weak answer names a theme but cannot operationalise the variable or specify a comparison, that is the sign that the code has not yet been promoted to variable status.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which historical figure is most associated with the WWII propaganda-analysis programme that established content analysis as a serious social-science method?

Lasswell led the U.S. Library of Congress Experimental Division for the Study of War-Time Communications and earlier authored Propaganda Technique in the World War (1927). Berelson codified the field in 1952; Krippendorff modernised it from 1980; Pennebaker developed LIWC for dictionary-based content analysis.

Question 2: The operational move that makes content analysis a hybrid quantitative/qualitative method is best described as:

Codes-as-variables is the hinge: qualitative judgement defines the codes, but once defined, each code becomes a variable that can be analysed with descriptive and inferential statistics.

Question 3: Which of the following best characterises the relationship between manifest and latent content in contemporary content analysis?

Berelson restricted content analysis to manifest content; Krippendorff and the subsequent field opened the door to latent. The modern compromise is to use both, but to be explicit about which is which and to demonstrate reliability for latent codes.
Section 2 of 5

Designing a Content Analysis: Sampling Text, Recording Units, and Reliability

⏱ Estimated reading time: 30 minutes
Section 2 of 5

Designing a Content Analysis

Sampling text, recording units, and reliability.

Three unit types

Sampling, recording, and context units

Sampling unit

The chunk drawn from the population. In the loneliness corpus: each interview transcript.

Recording unit

Where the variable takes its value. May be the whole transcript, a paragraph, a sentence, or a speaker turn.

Context unit

The surrounding text a coder may consult. Wider context supports latent coding; narrower context increases reliability.

The coding scheme

Three requirements

Exhaustive

Every recording unit can be assigned at least one code, or an explicit "none applies" category.

Defensibly exclusive

Code boundaries are clear enough to guide consistent decisions. Multi-coding is permitted, but declared.

Operationally defined

Each code has a description, an inclusion rule, and an exclusion rule. Another coder can apply it without asking you.

Reliability standard

Krippendorff's alpha

Krippendorff's alpha
\[ \color{#0B7B6B}{\alpha} = 1 - \frac{\color{#C2410C}{D_o}}{\color{#6D28D9}{D_e}} \]
α reliability coefficient Do observed disagreement De disagreement expected by chance

\(D_o\) = observed disagreement; \(D_e\) = expected disagreement by chance.

α ≥ 0.80

Definitive claims

0.67–0.80

Tentative inference; flag in write-up

< 0.67

Revise codebook before proceeding

End-to-end

The workflow in order

  1. Formulate a distributional / comparative research question
  2. Define the text population and sample from it
  3. Declare sampling, recording, and context units in writing
  4. Develop the codebook (a priori + emergent)
  5. Train coders on a practice subset
  6. Compute reliability on 10–20% of corpus
  7. Revise codebook if α < 0.67
  8. Apply final scheme to full corpus
  9. Analyse frequency matrix; run inferential tests
  10. Report transparently (corpus, units, codebook, reliability)
Carry forward

What to take into the next section

  • Unit choices shape both what the analysis shows and how reliable it will be.
  • Krippendorff's alpha: 0.80 definitive, 0.67 tentative, below 0.67 revise.
  • A low alpha is a diagnosis of a codebook problem, not a verdict on the data.

A later section covers what to do with the coded matrix once reliability is acceptable.

Introduction and Overview

An earlier section framed content analysis methodologically. This section is about the design decisions that determine whether your content analysis will be defensible. There are three of them, in order: what counts as the text to be analysed (sampling), what counts as the unit being coded (recording units and context units), and how the codes are applied with sufficient consistency that the resulting counts mean something (reliability). Get any one of these wrong and the analysis is suspect; get all three right and you have a piece of content-analytic work that a sceptical quantitative reviewer will accept.

Bernard, Wutich, and Ryan are unusually explicit on these matters because content analysis, more than any other qualitative method, can be done badly in ways that produce numbers that look authoritative but are not. The history of the field is littered with frequency tables built on inconsistent coding of poorly specified units sampled from non-representative corpora. The remedy is procedural: you commit, in writing, to a sampling rule, a unit-definition rule, a coding scheme, and a reliability target, before you begin coding. You then execute the procedure transparently, report what happened, and let the reader audit it.

Learning Objectives for this section

  • Distinguish sampling units, recording units, and context units, and choose each defensibly for your research question.
  • Build a content-analytic coding scheme that is exhaustive, defensibly exclusive (or explicitly multi-coded), and clearly defined.
  • Distinguish a priori (deductive) from emergent (inductive) coding schemes, and recognise the typical hybrid in practice.
  • Compute Krippendorff's alpha and interpret it against the field's reliability thresholds.

2.1 Sampling Text: The Three Kinds of Units

Krippendorff's most enduring methodological contribution after the alpha statistic is the distinction between three levels of unit in any content-analytic design:

  • Sampling units, the chunks of text that are drawn from the population. In a corpus of 20 loneliness transcripts, each transcript is a sampling unit. In a corpus of newspaper coverage of the overdose crisis, each article is a sampling unit. In a Twitter dataset, each tweet is a sampling unit.
  • Recording units, the unit you actually code. The recording unit is where the variable takes its value. Recording units may be the same as sampling units (you code each transcript as a whole) or smaller (you code each sentence, each paragraph, each turn-at-talk, or each named theme-bearing passage).
  • Context units, the surrounding text the coder is permitted to consult when deciding how to code a recording unit. If your recording unit is a sentence and your context unit is the paragraph, the coder reads the sentence in the context of the paragraph but assigns the code to the sentence alone.

The trio matters because the choice of each one shapes both what the analysis can show and how reliable it will be. A study that uses the whole transcript as the recording unit produces a frequency-per-transcript table; a study that uses the sentence as the recording unit produces a frequency-per-sentence table. These two analyses can give different answers to the same surface question, and the difference is methodological, not substantive.

Recording unit choiceWhat it lets you measureWhat it makes harderTypical use in loneliness corpus
The whole transcript Presence/absence of each code per participant Intensity, frequency within a transcript, longitudinal pattern within an interview "Of the 20 participants, how many invoke the loneliness-as-residue-of-marriage code at any point?"
The paragraph or speaker turn Within-transcript distribution; co-occurrence of codes Word-level features; reliability is harder than transcript-level "In how many turns of Linda's interview does she invoke the absent-Bill theme, and where in the interview do those turns cluster?"
The sentence Fine-grained distribution; sequence; intensity by participant Higher coding burden; lower reliability for latent codes "What proportion of Linda's sentences contain spatial metaphors for absence?"
The word or phrase (dictionary-based) Massively scalable counts Latent meaning entirely; ambiguity resolution; context "How often does each transcript contain the word 'chair', and does the count correlate with widowhood status?"

For the loneliness dataset, recommended: transcript-level recording units with paragraph-level context

In a 20-transcript corpus, the most defensible choice is to code each transcript as either containing or not containing each code (a binary recording-unit-per-transcript design), with each paragraph as the context unit you read to make the decision. This gives you a clean codes × participants matrix that supports chi-squared comparisons across subgroups. The trade-off is that you cannot measure within-transcript intensity, but for a corpus this size, between-transcript distributional analysis is the analytically productive level.

2.2 Developing a Coding Scheme

Identify sampling unitsv

What counts as a case? Documents? Paragraphs? Speaker turns? Articles? The sampling unit is what your eventual frequencies will be ratios of. Decide it before coding.

Identify recording unitsv

What counts as a single coded instance? A word? A sentence? A paragraph? A clause? The recording unit determines how granular the codes are and how much data each code generates.

Develop the coding schemev

Three sources: (1) prior literature/theory (deductive codes), (2) pilot reading of data (inductive codes), (3) iterative refinement from the first 10-20% of the corpus. The final scheme almost always combines deductive and inductive codes.

Train, pilot, refinev

Two or more coders apply the scheme to a pilot set (typically 10-15% of the corpus). Compare disagreements line-by-line. Refine code definitions, add inclusion/exclusion rules, add exemplars. Re-pilot until reliability is acceptable.

Apply to the full corpusv

Full-corpus coding by trained coders. Spot-check reliability on a held-out sample (typically 10%). Re-train if reliability drifts.

Aggregate and analyzev

Frequency tables, cross-tabulations, chi-squared tests, dictionary expansions, time trends. The countable output is the comparative advantage of content analysis over thematic analysis.

The coding scheme, the codebook, in the language of an earlier module, is the spine of any content analysis. In an earlier module you built one for thematic analysis; here you adapt it for the variable-treatment of content analysis. Three properties matter more in content analysis than they did in thematic analysis.

Exhaustive coverage. Your coding scheme should cover the content you care about. For content analysis specifically, exhaustive coverage means that every recording unit can be assigned at least one code (or the explicit "not-coded" residual category). The classical Berelson position is that codes should be mutually exclusive, each unit gets exactly one code, on the grounds that overlapping codes make the resulting frequencies hard to interpret. Krippendorff and the contemporary field reject the mutual-exclusivity requirement: a passage can simultaneously instantiate "loneliness-as-existential-fact" and "loneliness-coped-with-by-pet-companionship" and there is no good reason to force the coder to choose. The modern compromise is multi-coded passages are permitted if the codebook says so and the multi-coding is consistent.

Clear operational definitions. Each code in your scheme needs a definition that another coder could apply without consulting you. The definition has three parts: a brief substantive description, an inclusion rule (what counts), and an exclusion rule (what doesn't count). Inclusion and exclusion rules are where reliability is won or lost. "Mentions of loneliness" is a definition without inclusion/exclusion rules. "Any statement in which the participant attributes loneliness to a specific event or relationship (inclusion); excluding generic statements about loneliness in the population or society at large (exclusion)" is a definition that can be applied reliably.

A priori versus emergent. Content-analytic codebooks come from one of two directions, or (usually) both. A priori codes come from the literature, your conceptual framework, or your interview guide; you decide before you read the data that you will code for, say, "stigma," "social-comparison," and "coping-by-substance-use" because the literature on loneliness says these are the right categories. Emergent codes come from the data, you read the corpus, notice what is there, and let the codebook grow to fit. Hsieh and Shannon (2005) call the first directed content analysis and the second conventional content analysis; their summative third type starts from word counts and works upward. Elo and Kyngäs (2008) and Mayring (2000) offer complementary process descriptions widely cited in nursing and European health-research traditions.

ApproachCodebook sourceAdvantageRisk
Conventional (inductive / emergent) Codes emerge from close reading of the corpus Sensitive to participants' own categories; surfaces unexpected themes May reproduce the analyst's preconceptions; can be unsystematic
Directed (deductive / a priori) Codes drawn from theory, prior literature, or instrument Replicable; comparable across studies; testable against theory May miss what the data are actually doing if the codebook is wrong
Summative Codes start as word counts, then expand to latent meaning Scales; obvious replicability; computational tractability Manifest-only by default; risks counting words without coding meaning
Hybrid (most contemporary work) A priori scaffold plus emergent expansion Best of both; transparent about origin of each code Requires explicit documentation of which codes came from where

Recommended scaffold for a content analysis of the loneliness dataset

For a first content analysis of these data, the recommended workflow is hybrid: take 5–7 codes from an existing Taguette codebook, declare them as an a priori scaffold, then apply them systematically across 8–12 transcripts. As you code, allow up to two emergent codes if you find a category that cannot be accommodated by the original 5–7. Document each emergent code with the same care as a priori ones, brief description, inclusion rule, exclusion rule.

2.3 Inter-coder Reliability: Stricter Than Thematic Analysis

Simple agreementClick to explore
Cohen’s kappaClick to explore
Krippendorff’s alphaClick to explore
Beyond the numberClick to explore

Reliability is the test of whether two coders, working independently with the same codebook, would assign the same codes to the same passages. In thematic analysis (an earlier module), reliability matters but it is one consideration among several; in content analysis, reliability is load-bearing, the frequencies you report are only as good as the consistency with which the codes were applied. If two coders applying the same codebook produce different counts, the counts are noise.

The field's standard reliability statistic is Krippendorff's alpha (Krippendorff, 2004, 2018; Hayes & Krippendorff, 2007). Alpha has three properties that recommend it over the older Cohen's kappa: it accommodates any number of coders (where Cohen's kappa handles two); it accommodates any level of measurement (nominal, ordinal, interval, ratio); and it handles missing data gracefully. The arithmetic compares the observed disagreement among coders to the disagreement expected by chance, and produces a coefficient that runs from 1 (perfect agreement) through 0 (chance) to negative values (worse than chance). Put simply, alpha asks how much better than random guessing the coders did: when they never disagree the observed disagreement is zero and alpha is 1, and when they agree only as often as chance would predict the two disagreements are equal and alpha is 0.

Krippendorff's alphaInterpretationTypical recommendation
α ≥ 0.80 Strong agreement Acceptable for definitive content-analytic claims (Krippendorff, 2018)
0.67 ≤ α < 0.80 Acceptable for tentative inference Report; flag as tentative; discuss areas of disagreement
α < 0.67 Inadequate Revise the codebook and re-train; do not publish counts from this code

The thresholds are convention, not law, and they vary slightly between sources. The 0.80/0.67 dichotomy comes from Krippendorff's own writing and is the most commonly cited in the field. Some health-research applications adopt a stricter 0.70 floor and a 0.80 publication target. Whatever you adopt, declare it in your methods section before you compute it.

Reliability is computed on a subset of the corpus, typically 10–20%, that two coders code independently. The remaining 80–90% is then coded by one coder alone, with the assurance that the reliability of the system is documented. The reliability subset is usually drawn randomly from the corpus to ensure that the reliability estimate is generalisable.

What to do when reliability is low

A low alpha is not the end of the analysis; it is a diagnosis. The cause is almost always one of: (a) a code definition that is too vague (revise inclusion/exclusion rules); (b) a code definition that covers too much (split into sub-codes); (c) insufficient coder training (run additional training sessions on a separate set of practice passages); or (d) the code is genuinely contested in the corpus (which is itself a finding, report it). The workflow is iterative: train, code a subset, compute alpha, revise the codebook, re-train, re-code, re-compute. Two cycles is typical; three is not unusual.

2.4 The Workflow End-to-End

ACTIVITY Try it - Pilot a coding scheme

Take 5-10 short documents (e.g., recent news headlines about a public health topic). Develop a small coding scheme:

  1. Define 3-5 codes that capture features you care about (e.g., 'mentions vaccination', 'cites a Canadian source', 'uses risk framing').
  2. Write a one-sentence definition + a one-sentence inclusion rule for each code.
  3. Code the documents yourself once. Set the codes aside.
  4. The next day, re-code the same documents without looking at your original codes. Compute your own intra-coder reliability.

If you can’t agree with yourself, two independent coders will agree even less. Most content-analysis projects need 2-3 iterations before reliability is acceptable.

A complete content-analytic study, in Bernard, Wutich, and Ryan's framing, follows these steps in order:

  1. Formulate the research question, explicitly distributional or comparative.
  2. Define the population of texts, what corpus is the analysis drawn from?
  3. Sample the texts, if the population is small enough, the entire population may be sampled; otherwise, define a sampling frame and a sampling procedure.
  4. Decide on sampling, recording, and context units, declare them in writing.
  5. Develop the codebook, a priori, emergent, or hybrid; with operational definitions for every code.
  6. Train coders, on a practice subset that does not enter the final analysis.
  7. Compute reliability on a 10–20% subset, Krippendorff's alpha; revise codebook if needed.
  8. Apply the final codebook across the corpus, one coder for the remaining 80–90% is typical for a small-corpus study.
  9. Analyse the resulting frequency table, descriptive statistics, then inferential tests (a later section).
  10. Report transparently, corpus, sampling, units, codebook (in an appendix), reliability, analysis, interpretation.
The content-analysis workflow A comparative question leads to sampling text, defining units, and building a codebook. Coders are trained and inter-coder reliability is computed; if Krippendorff's alpha is below threshold the codebook is revised and reliability rechecked. Once reliability is acceptable the codebook is applied to the full corpus, the frequency table is analysed, and results are reported. Comparativequestion Sample thetext Define units Build codebook Train coders & compute Krippendorff's α Apply to fullcorpus Analyse table Report α ok α low:revise
The workflow is sequential, with one feedback loop: a low reliability coefficient sends the analyst back to revise the codebook before the scheme is applied to the full corpus.

Reflection

Suppose you are designing a content analysis of the loneliness dataset. Declare the design decisions you would make. Which 5–7 codes would you use? What is your recording unit (whole transcript / paragraph / sentence)? What is your context unit? What reliability target would you adopt, and how would you handle codes that do not meet it?

Model answerA strong answer names the codes, the unit choices, and the reliability rule. Example: "I would use the eleven codes in the codebook supplied with the dataset: loneliness-as-embodied-ache, loneliness-amid-company, loneliness-of-not-being-known, loneliness-as-shrinking-world, loneliness-as-sole-responsibility, loneliness-hidden-by-shame, loneliness-as-cost-of-love, loneliness-as-existential-fact, loneliness-and-technology, loneliness-coped-by-pet, and loneliness-as-untranslatable-word. My recording unit will be the whole transcript: each transcript is coded as containing or not containing each code, producing a 20 × 11 binary matrix. My context unit is the paragraph, if a passage is ambiguous, the coder reads the surrounding paragraph before deciding. I will adopt Krippendorff's α ≥ 0.70 as my acceptability threshold and α ≥ 0.80 as my publication target. Codes that fall between 0.67 and 0.80 will be reported but flagged as tentative; codes below 0.67 will be revised and re-coded or dropped." A weak answer either does not commit to specific codes or does not name a reliability rule.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In Krippendorff's terminology, what is a recording unit?

The recording unit is where the code is applied and the variable takes its value, commonly the whole transcript, the paragraph, or the sentence. Sampling units are drawn from the population; context units are the surrounding text consulted while coding.

Question 2: Which reliability statistic is the contemporary standard for content analysis, accommodating any number of coders and any level of measurement?

Krippendorff's alpha is the field's standard because it accommodates any number of coders, any level of measurement, and missing data. Cohen's kappa (Cohen, 1960; thresholds: Landis & Koch, 1977) is restricted to two coders and nominal data; percent agreement does not correct for chance.

Question 3: What is the typical interpretation of a Krippendorff's alpha of 0.74 on a content-analytic code?

Krippendorff's convention places 0.67 ≤ α < 0.80 in the "tentative" range, acceptable for reporting but with the code's reliability flagged. Below 0.67 the code should be revised and re-coded; at or above 0.80 the code supports definitive claims.
Section 3 of 5

Analyzing Coded Text: Hypothesis Testing, Comparisons, and Dictionaries

⏱ Estimated reading time: 30 minutes
Section 3 of 5

Analysing Coded Text

Frequency tables, cross-tabulations, chi-squared tests, and dictionary-based approaches.

The starting point

Frequency tables

Code Transcripts (n=20) % Technology as double-edged1470% Shame prevents disclosure1365% Role loss as trigger1050% Physical proximity not equal to connection945% Loneliness as residue of marriage840% Competence-based identity735% Anticipatory grief630%
The comparison move

Cross-tabulation by subgroup

Code Caregiver (n=10) Non-caregiver (n=10) Technology as double-edged7 (70%)7 (70%) Shame prevents disclosure9 (90%)4 (40%) Role loss as trigger6 (60%)4 (40%) Anticipatory grief5 (50%)1 (10%) Competence-based identity6 (60%)1 (10%) Physical proximity not equal to connection3 (30%)6 (60%)
Inferential tests

Chi-squared and Fisher's exact

Chi-squared test of independence
\[ \color{#0B7B6B}{\chi^2} = \sum \frac{(\color{#C2410C}{O} - \color{#6D28D9}{E})^2}{\color{#6D28D9}{E}} \]
χ² chi-squared statistic O observed cell count E expected count under independence

Chi-squared

Use when all expected cell counts are at least 5. Returns χ², degrees of freedom, and p-value.

Fisher's exact

Use when any expected cell count is below 5, which is common in 20-transcript corpora. More conservative; preferred for small samples.

Automated approaches

Dictionary-based content analysis

A pre-specified word list is applied computationally. Outputs are frequency counts per category, treated as variables like any other code.

LIWC (Linguistic Inquiry and Word Count, Pennebaker et al.) classifies words into 80+ psychological and linguistic categories.

Common sentiment dictionaries: Bing Liu's Bing lexicon, AFINN, and the NRC Emotion Lexicon.

Best practice

Dictionaries for manifest, replicable counts. Human coding for latent meaning. Both together are stronger than either alone.

Carry forward

What to take into the next section

  • Frequency tables are the starting point; cross-tabulation by subgroup is where findings emerge.
  • Fisher's exact is usually the right test in 20-transcript corpora with small expected cells.
  • Dictionaries and human coding are complementary, not competing approaches.

A later section walks through the complete workflow: Taguette export to coded matrix to published result.

Introduction and Overview

Once your coding is done and your reliability is acceptable, you have a coded matrix: rows are recording units (transcripts, paragraphs, or sentences), columns are codes, and the cell values are either binary indicators (1 if the code applied, 0 otherwise) or counts (the number of times the code applied to that unit). This matrix is the input to the inferential half of content analysis. This section walks through what you can do with it: descriptive frequency tables, cross-tabulation by subgroup, chi-squared tests for distributional differences, and trend analyses where the corpus is longitudinal. It then introduces dictionary-based content analysis, LIWC and sentiment dictionaries, and previews the computational content analysis that a later module will develop in depth.

The intellectual point of this section is that content analysis is a hypothesis-testing method when it wants to be. Whatever you tested in your earlier regression work, you can test on content-analytic codes: differences in proportions, associations, trends, interactions. The variables are codes rather than survey items, but the inferential machinery is the same. This is what Bernard, Wutich, and Ryan mean when they say content analysis is the bridge between the qualitative and quantitative traditions.

Learning Objectives for this section

  • Compute and present a content-analytic frequency table at three levels: overall, by subgroup, and over time.
  • Apply chi-squared tests of independence (and Fisher's exact for small cells) to codes × subgroup matrices.
  • Interpret distributional differences in content-analytic data substantively as well as statistically.
  • Recognise when dictionary-based content analysis (LIWC, sentiment dictionaries) is appropriate.
  • Locate computational content analysis (a later module) as the scalable extension of the dictionary approach.

3.1 Frequency Tables: The Starting Point

Every content-analytic study starts with descriptive frequencies. For the loneliness corpus, a typical first-pass frequency table looks like the one below: each of the eleven codes in the codebook supplied with the dataset, the count of transcripts in which the code appears (out of 20), and the percentage. These are the numbers the supplied export produces; a different codebook would produce different ones.

CodeTranscripts (n / 20)Percentage
loneliness-as-embodied-ache1890%
loneliness-amid-company1575%
loneliness-and-technology1575%
loneliness-of-not-being-known1365%
loneliness-hidden-by-shame1260%
loneliness-as-existential-fact1155%
loneliness-as-cost-of-love1155%
loneliness-as-shrinking-world1050%
loneliness-as-sole-responsibility420%
loneliness-coped-by-pet315%
loneliness-as-untranslatable-word315%

Already this table is doing analytic work. The most prevalent code, loneliness-as-embodied-ache, appears in 18 of 20 transcripts, suggesting that loneliness is described in bodily terms (ache, weight, exhaustion, sleeplessness) by almost everyone, regardless of who the participant is. The least prevalent, loneliness-as-untranslatable-word, appears in 3 transcripts, and you can already predict who those three participants are: the ones who came to Canada from elsewhere (Aarav, Amira, Chen). The frequency table tells you both what is shared across the corpus and what is concentrated.

3.2 Cross-Tabulation by Subgroup: The Comparison Move

Content analysis's real value emerges when you cross-tabulate the codes by a participant-level variable: age band, gender, caregiving status, immigration status, life-stage. Here is the same code distribution disaggregated by caregiver status in our worked corpus (5 caregivers, 15 non-caregivers):

CodeCaregivers (n=5)Non-caregivers (n=15)Total (n=20)
loneliness-as-embodied-ache5 (100%)13 (87%)18 (90%)
loneliness-amid-company5 (100%)10 (67%)15 (75%)
loneliness-and-technology4 (80%)11 (73%)15 (75%)
loneliness-of-not-being-known2 (40%)11 (73%)13 (65%)
loneliness-hidden-by-shame3 (60%)9 (60%)12 (60%)
loneliness-as-existential-fact1 (20%)10 (67%)11 (55%)
loneliness-as-cost-of-love3 (60%)8 (53%)11 (55%)
loneliness-as-shrinking-world2 (40%)8 (53%)10 (50%)
loneliness-as-sole-responsibility4 (80%)0 (0%)4 (20%)
loneliness-coped-by-pet0 (0%)3 (20%)3 (15%)
loneliness-as-untranslatable-word0 (0%)3 (20%)3 (15%)

Now the analytic action begins. The loneliness-as-sole-responsibility code appears in 4 of the 5 caregivers and in none of the 15 non-caregivers, a categorical difference that is the kind of finding content analysis is built to produce. The loneliness-as-existential-fact code shows the reverse pattern: 10 of 15 non-caregivers (the older, more settled participants like Helen and Frank) against 1 of 5 caregivers, whose loneliness is tied to a situation they expect to end. The loneliness-and-technology code is close to invariant by subgroup (80% against 73%), supporting the earlier hypothesis that this is a near-universal feature of contemporary loneliness regardless of life-stage. Note also the code that does not move: loneliness-hidden-by-shame sits at 60% in both groups, so the intuition that caregivers are the ones who hide their loneliness is not supported here. These patterns are descriptive; the next move is to ask whether they are larger than chance.

3.3 Chi-Squared Tests on Codes × Subgroup

The classical inferential test for a 2 × 2 (or larger) contingency table of categorical data is the chi-squared test of independence. Applied to a content-analytic frequency table, the chi-squared test asks: is the distribution of this code statistically different between the subgroups? The mechanics are familiar from earlier courses; here we apply them to qualitative codes treated as variables.

Take the loneliness-as-sole-responsibility code from the table above: 4 of 5 caregivers vs. 0 of 15 non-caregivers. The chi-squared test on this 2 × 2 table gives χ² = 10.42 with 1 degree of freedom, p = 0.001 (Fisher's exact p = 0.001; with cells this small Fisher is the number to report). The difference is statistically significant at conventional thresholds. Practically, and this is the part the qualitative analyst contributes, the analysis suggests that caregiver participants in this corpus describe a loneliness of being the only adult responsible for a dependent person, in a way that non-caregivers never do. The substantive interpretation might be that caregiving concentrates decision-making in one person, and that the absence of anyone to share the load is itself the lonely part. The chi-squared test is the warrant for the claim; the qualitative reading is the claim.

When to use Fisher's exact instead

Chi-squared tests assume that expected cell counts are at least 5 (some sources say 1). In small qualitative corpora, especially those with rare codes, this assumption frequently fails. The standard remedy is Fisher's exact test, which gives an exact p-value regardless of cell size. Statistical software that offers the chi-squared test generally offers Fisher’s exact test alongside it; for a 20-transcript corpus with codes appearing in 4–14 transcripts, Fisher's exact is almost always the more defensible choice. The shame example above shows why this matters: its expected counts are only 3.5 in two of the four cells, so although the uncorrected chi-squared gives p = 0.019, Fisher's exact returns p of about 0.06. Under the test you should trust here, that caregiver difference is better reported as suggestive than as conclusively significant. Report both if you wish, but lead with Fisher's.

3.4 Trend Analysis: Codes Over Time

Content analysis is at its strongest in longitudinal applications. Where your corpus has a time dimension, coverage of a topic across years of newspaper articles, posts on a forum across the pandemic, or transcripts collected at multiple time points, the codes can be plotted over time and tested for trend. The classical example is Pool's (1952) analysis of editorial coverage of the Soviet Union across decades; the contemporary example is the proliferation of computational content analyses of social-media discourse before, during, and after specific events (elections, crises, vaccine roll-outs).

The loneliness dataset is cross-sectional, so trend analysis is not the dominant move, but you may have a quasi-temporal variable. For example, you could ask whether participants who were interviewed earlier in the corpus (the first 10 transcripts) differ systematically from those interviewed later (the last 10) in any code, as a way of probing whether your codebook was over-fitted to early transcripts. Or you could examine whether participants who lived through the pandemic as adults (versus those who were teenagers at the time) describe loneliness differently. These are pseudo-trend analyses on a cross-sectional corpus, but the inferential logic is the same.

3.5 Dictionary-Based Content Analysis

So far we have assumed that humans are applying the codes. Dictionary-based content analysis is the variant in which a pre-specified word list (a "dictionary") is applied to the corpus by a computer, and the resulting counts are treated as content-analytic variables. The approach is best understood as the manifest-content extreme of content analysis: it counts what is explicitly there, fast and at scale, at the cost of any latent-meaning sensitivity.

The best-known dictionary in the social sciences is the Linguistic Inquiry and Word Count (LIWC) dictionary, developed by James Pennebaker and colleagues from the early 1990s (Pennebaker, Boyd, Jordan, & Blackburn, 2015; Boyd, Ashokkumar, Seraj, & Pennebaker, 2022). LIWC is a curated set of about 90 categories, positive emotion, negative emotion, anxiety, sadness, body, health, family, social processes, cognitive processes, and so on, each defined by a list of words and word stems. Run LIWC over a transcript and it returns the percentage of words in each category. A transcript that scores high on "sadness" contains a high proportion of words like sad, lonely, grief, cry, miss, lost; one that scores high on "social processes" contains a high proportion of pronouns referring to other people.

LIWC has been used extensively in health-related text analysis: predicting depression from writing samples (Pennebaker, Mehl, & Niederhoffer, 2003), characterising trauma narratives, distinguishing the writing of patients on different medications, predicting suicide risk from social-media posts. It is the methodological ancestor of contemporary sentiment-analysis systems and remains in active use; LIWC-22 is the current version.

Sentiment dictionaries are the lighter-weight cousin of LIWC. The best-known are Bing Liu's Bing lexicon (about 6,800 words classified as positive or negative), the AFINN lexicon (a scored version, with words assigned valence from -5 to +5), and the NRC Word-Emotion Association Lexicon (which classifies words into Plutchik's eight basic emotions plus positive/negative). All three are freely available word lists that text-analysis software can apply to a transcript corpus automatically.

Dictionary-based vs. human-coded content analysis

The choice is not exclusive. Most rigorous contemporary content-analytic studies use both: dictionary-based methods for the manifest, scale-sensitive, replicable counts, and human-coded latent analysis for the meaning-sensitive, ambiguity-tolerant, theory-engaged interpretation. The two halves answer different questions and are reported alongside each other. The danger is treating dictionary output as the whole analysis, LIWC says a transcript scores 4.2% on "sadness," but it cannot tell you that the sadness is bereavement-specific, that it co-occurs with relief, or that the participant is describing the sadness ironically. Only human coding can.

3.6 Computational Content Analysis: A Preview of a later module

Computational content analysis is the scalable extension of the dictionary approach, plus the addition of unsupervised methods that discover categories rather than counting pre-specified ones. The three families you will meet later in the course are: topic models (Latent Dirichlet Allocation and its successors, which discover thematic structure in large corpora without a codebook); word-embedding methods (which represent words as vectors in a high-dimensional space, enabling semantic similarity and analogy operations); and large-language-model-based coding (GPT-4-class systems applied as zero-shot or few-shot coders).

The methodological status of computational content analysis is still being worked out. The advantages are obvious: speed, scale, replicability of the computational pipeline. The disadvantages are real: opacity (LDA and LLMs cannot show their work the way a human coder's audit trail can), brittleness (a topic-model solution is sensitive to hyperparameter choices in ways the analyst rarely audits), and the recurring failure to validate computational outputs against human-coded gold standards. Bernard, Wutich, and Ryan's stance, and the stance of this course, is that computational content analysis is a useful complement to human content analysis, not a replacement for it. A later module will give you the operational machinery; this module is establishing the framework that machinery extends.

VariantWhat it countsScalabilityReliabilityLatent-content sensitivity
Human-coded latent content analysis (this module) Codebook-defined themes ~50–200 documents per project Krippendorff's alpha computed and reported High
Human-coded manifest content analysis Word or phrase occurrences ~500–5,000 documents per project Near-perfect Low
Dictionary-based (LIWC, sentiment) Pre-specified word lists Millions of documents Perfect within the dictionary Very low
Topic models (LDA; a later module) Unsupervised thematic structure Millions of documents Depends on hyperparameter validation Moderate but opaque
LLM-based coding (a later module) Anything specifiable in a prompt Limited by API cost, not by labour Variable; requires gold-standard validation High but unverifiable

Reflection

Consider the comparison in section 3.3: loneliness-as-sole-responsibility appeared in 4 of 5 caregivers vs. 0 of 15 non-caregivers, χ² p = 0.001. Imagine a reviewer who is sceptical of qualitative work pushes back: "you have 20 transcripts, can you really make a chi-squared claim?" Write your response. What are the legitimate cautions, and what are the legitimate defences?

Model answerA strong response acknowledges the legitimate cautions and then articulates the defensible position. Cautions: 20 transcripts is small for chi-squared; Fisher's exact would be the more defensible test given likely small expected cell counts; and the sample is purposive, not probabilistic, so the inferential claim cannot generalise to the British Columbian population, it generalises only to the comparison within this corpus. Defences: the test is doing exactly the work it should be doing, it answers the question "is the difference between caregivers and non-caregivers in this corpus larger than what we would expect by chance alone, given the corpus size?" That is a legitimate question even when the sample is small and non-probabilistic. The chi-squared test is not the warrant for a population-level prevalence claim (an earlier course territory); it is the warrant for a within-corpus distributional claim that triangulates with the qualitative reading. A defensible methods section will say: "Within this purposive corpus of 20 transcripts, the difference in code prevalence between caregivers and non-caregivers exceeded what would be expected by chance (chi-squared p = 0.019; the more conservative Fisher's exact test gives p of about 0.06, just short of the conventional 0.05 cutoff). Given the corpus size the finding is best reported as suggestive and as warranting further investigation in a larger study, not as a claim that generalises."

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A content-analytic study finds that a code appears in 9 of 10 caregivers and 4 of 10 non-caregivers. Which test is the most defensible inferential choice given the small expected cell counts?

Fisher's exact test gives an exact p-value regardless of cell size and is the standard recommendation when chi-squared's expected-count assumption may fail, common in small qualitative corpora.

Question 2: Which best characterises the relationship between LIWC and human-coded latent content analysis?

LIWC operates at the manifest, word-list level. It is fast and scalable but cannot detect ambiguity, irony, or latent meaning. Rigorous contemporary studies use it alongside, not instead of, human-coded latent analysis.

Question 3: Why is a chi-squared (or Fisher's exact) test on a content-analytic codes × subgroup matrix a legitimate inferential move, even when the sample is small and purposive?

The chi-squared / Fisher's test answers a within-corpus question: is the distributional difference larger than would arise by chance, given the corpus size? It does not warrant population-level prevalence claims, but it is a legitimate triangulation of the qualitative reading.
Section 4 of 5

Content-Analyzing the Loneliness Dataset: The Analysis Workflow

⏱ Estimated reading time: 45 minutes
Section 4 of 5

The Analysis Workflow

From Taguette export to coded matrix, frequency tables, visualisation, and inferential tests.

Step 1

Load the Taguette export

Input

A CSV from Taguette: one row per highlighted passage, columns for document filename, code (tag), and passage text.

After the join

A coding table with participant-level variables (age, gender, caregiver, immigrant, life_stage) attached to each passage row. Ready for reshaping.

Step 2

Reshape into codes × cases matrix

Long format (Taguette) doc | tag | content P01 | shame | ... P01 | tech | ... P02 | shame | ... P03 | role_loss | ... (one row per passage) → reshape Wide format (analysis) id | shame | tech | role_loss | ... P01 | 1 | 1 | 0 | ... P02 | 1 | 0 | 1 | ... P03 | 0 | 1 | 1 | ... (one row per participant)
Steps 3 & 4

Frequency tables and visualisation

Subgroup summary

Group by subgroup variable; count how many participants in each group have each code; compute percentages. One table per comparison.

Bar chart

Horizontal bars; codes sorted longest-to-shortest; two bars per code (e.g., caregiver vs. non-caregiver). Codes where bars diverge most are the primary findings.

Steps 5 & 6

Inferential tests and keyness analysis

Step 5: Contingency tests

One test per code that showed a meaningful descriptive difference. Report odds ratio alongside p-value. Fisher's exact for small expected cells.

Step 6: Keyness

Which words appear disproportionately in one subcorpus? Log-likelihood ratio; complement to code-level chi-squared.

Carry forward

What a content analysis reports

Four artefacts

Taguette export; codebook with operational definitions; reliability report (Krippendorff's alpha); frequency table, bar chart, and at least one inferential test.

Write first

Before running any test, declare in writing: which subgroup contrast, which codes, what you expect. That declaration anchors the methods section.

A later section is the final assessment, where you consolidate and check your understanding.

Introduction and Overview

Earlier sections framed content analysis methodologically. This section turns operational. You will see, end to end, how to take a Taguette codebook, apply it across the 20 loneliness transcripts, transform the resulting export into a codes × cases matrix, compute frequency tables and visualisations, run chi-squared and Fisher's exact tests on codes × subgroup, and conduct a keyness analysis that surfaces which words differ most between two subcorpora.

The workflow that follows starts from an exported Taguette CSV with one row per (passage, code) pair, columns for the document filename, the tag (code), and the passage content. A first coding pass usually covers 3–5 transcripts; a content analysis extends the codebook to 8–12 transcripts or more. The volume jump is intentional: you need enough cases for the frequency tables and chi-squared tests to be analytically informative, and 8–12 transcripts is the floor for a defensible content analysis of this dataset.

Learning Objectives for this section

  • Load a Taguette export and reshape it into a codes × cases matrix.
  • Compute frequency tables overall and by subgroup.
  • Visualise code frequencies by subgroup with a bar chart.
  • Conduct chi-squared and Fisher's exact tests on codes × subgroup.
  • Use keyness analysis to compare the vocabulary of two subcorpora.
  • Report a content analysis with its coded data, codebook, reliability statistics, and analysis output.

4.1 Step 1: Load the Taguette Export and Add Participant Variables

Your Taguette export is a CSV with one row per highlighted passage. Each row contains the document filename (which encodes the participant ID), the tag (your code), and the passage content. The first step is to open it as a table (a spreadsheet works well) and add the participant-level variables (age, gender, caregiver status, etc.) needed for subgroup comparison.

The course data bundle, HSCI_841_loneliness_data.zip, contains a worked Taguette export (taguette_export_week8.csv), the codebook with its 11 codes (codebook_week8.csv), and the participant table (participant_metadata.csv). The participant identifier is the part of each document filename before the pseudonym (for example, P01 in P01_Maya.txt), and it is the key that links each coded passage to its row in the participant table. After the join, each row is one highlighted passage, with columns for the code (tag), the passage text (content), and the participant-level variables (age, gender, life stage, caregiver status, immigration status) needed for subgroup analysis. Before going further, confirm that every participant in the export has been matched to a row in the participant table.

4.2 Step 2: Reshape Into a Codes × Cases Matrix

The codes-as-variables move is now a literal data-shape operation: convert the long-format export into a wide-format matrix where rows are participants and columns are codes. The cell values are typically binary (1 if the code applied to that participant at all, 0 otherwise) but can be counts (the number of times the code applied within the transcript).

A spreadsheet pivot table with participants as rows and codes as columns produces the count version directly; replacing every non-zero count with 1 gives the binary version. The participant attributes (age, gender, life stage, and so on) are then added as further columns. The result has one row per participant, 20 rows for the full corpus or 8–12 for a coded subset, and it is the analytic-ready frame for everything that follows.

4.3 Step 3: Frequency Tables

The overall and by-subgroup frequency tables in Sections 3.1 and 3.2 come from simple counts on the binary matrix. The overall frequency of a code is the number of participants with a 1 in that code’s column; the by-subgroup frequency is the same count taken separately within each subgroup and divided by the size of that subgroup.

Worked example: caregiver status in the loneliness corpus

Split by caregiver status (5 caregivers, 15 non-caregivers), the worked corpus in Section 3.2 shows three patterns. Loneliness-as-sole-responsibility appears in 4 of 5 caregivers (80%) and in none of the 15 non-caregivers. Loneliness-as-existential-fact shows the reverse, appearing in 10 of 15 non-caregivers (67%) and 1 of 5 caregivers (20%). Loneliness-hidden-by-shame sits at 60% in both groups (3 of 5 and 9 of 15), and loneliness-and-technology is close to invariant (80% and 73%). The same counts can be repeated for any other participant variable, such as age band or immigration status. A coded subset of 8–12 transcripts will give different numbers depending on which transcripts it includes.

4.4 Step 4: Visualise the Subgroup Comparison

A bar chart of code frequencies by subgroup is the standard published display for a content-analytic study, and any spreadsheet or statistics program can draw one.

The usual layout is a horizontal bar chart with the codes on the vertical axis, sorted so that the longest bars sit at the top, the percentage of each subgroup with the code present on the horizontal axis, and two bars per code, one for caregivers and one for non-caregivers. The chart shows percentages because the subgroups differ in size (5 and 15). In the worked corpus the bars pull apart most sharply for loneliness-as-sole-responsibility (80% against 0%) and loneliness-as-existential-fact (20% against 67%), and they sit level for loneliness-hidden-by-shame (60% in both groups). Codes for which the two bars differ markedly are the ones worth subsequent chi-squared testing.

4.5 Step 5: Chi-Squared and Fisher's Exact Tests

The inferential half of the analysis. For each code that showed a meaningful descriptive difference between subgroups in step 4, run the appropriate contingency-table test.

Worked example: loneliness-as-sole-responsibility by caregiver status

The code that separates the two subgroups most sharply in the frequency table gives the following 2 × 2 table.

Code presentCode absentTotal
Caregivers415
Non-caregivers01515

The chi-squared test with Yates’ continuity correction gives χ² = 10.42 with 1 degree of freedom, p = 0.001, and Fisher’s exact test also gives p = 0.001. Because no non-caregiver carries the code, one cell is zero and the sample odds ratio cannot be computed as a finite number; in that situation the two proportions (80% against 0%) are reported alongside Fisher’s exact p-value. Repeating the test for every code produces a results table, sorted by p-value, for the methods and findings sections.

What to report: lead with the codes that show the largest distributional differences (lowest p-values); report the odds ratio alongside the p-value, because p-values from small samples are misleading on their own; and flag any codes for which expected cell counts fall below 5 by noting that Fisher’s exact test was used.

4.6 Step 6: Keyness Analysis

Keyness is the manifest-content cousin of the chi-squared analysis above. It asks: which words appear disproportionately in one subcorpus compared to another? The classical implementation is the log-likelihood ratio test on word frequencies between two subcorpora, and corpus-linguistics software computes it for every word in the vocabulary.

For the loneliness dataset, the natural keyness contrast is one of the subgroup contrasts you have already been working with: caregiver vs. non-caregiver, immigrant vs. non-immigrant, or under-40 vs. over-65. The output is a ranked list of words, each with a chi-squared (or log-likelihood) statistic and a sign indicating which subcorpus the word over-occurs in.

Worked example: caregiver and non-caregiver transcripts

The 20 transcripts are grouped into a caregiver subcorpus (5 transcripts) and a non-caregiver subcorpus (15 transcripts). The text is lower-cased, punctuation, numbers, and common stopwords are removed, and the frequency of each remaining word in the caregiver subcorpus is compared with its frequency in the non-caregiver subcorpus using the log-likelihood ratio, with the caregiver subcorpus as the target. Words such as kids, mom, dad, husband, hospital, appointment, exhausted can be expected to load on the caregiver side, and words such as chair, walker, eyes, fading, alone, quiet on the non-caregiver (older, more isolated) side. The keyness analysis is manifest content analysis at scale; combined with your latent codebook results, the two halves triangulate to support stronger claims than either alone.

4.7 What a Content Analysis Reports

A content analysis of a coded subset ends in a frequency table and at least one defensible quantitative comparison. A complete report of it includes four artefacts: the coded transcripts (Taguette export), the frequency table (as a spreadsheet or CSV file), the chi-squared or Fisher's exact comparison with its odds ratio, and an interpretive memo of about 500 words that places the quantitative findings back into a qualitative reading.

Reflection

Before you run any part of the analysis, declare in writing what comparison you intend to test. Which subgroup contrast (caregiver/non-caregiver, immigrant/native-born, older/younger, women/men)? Which code do you predict will differ most? And what is the substantive reasoning behind your prediction?

Model answerA strong response makes the prediction before running the analysis, the pre-registration logic of an earlier course carrying into this course. Example: "I will test the contrast between participants who self-identify as caregivers (n=5) and those who do not (n=15). I predict that the code loneliness-hidden-by-shame will appear more frequently in caregiver transcripts, because the caregiving identity carries a moral demand of unselfish endurance that conflicts with admitting loneliness. I also predict that loneliness-as-existential-fact, Helen's code, will appear more frequently in non-caregivers, because the non-caregivers in this corpus include the oldest and most settled participants, for whom loneliness has become a lasting condition rather than a situation. The technology code I expect to be invariant by caregiver status." In the worked corpus the first prediction fails, the shame code sits at 60% in both groups, while the second holds, 67% of non-caregivers against 20% of caregivers, and the code that does separate the groups, loneliness-as-sole-responsibility, was not predicted at all. Writing the prediction down first protects you from p-hacking and gives you a way to discuss findings that contradicted your expectation in the memo, which is often the most analytically valuable section.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A complete report of a content analysis of the loneliness transcripts includes:

A content analysis reports the coded transcripts, the frequency table, the chi-squared or Fisher's exact comparison, and an interpretive memo that places the numbers back into a qualitative reading. A grounded-theory analysis or a theme list without counts answers a different kind of question.
Section 5 of 5

Final Assessment

⏱ Estimated time: 30 minutes

Bringing It All Together

This lesson has done two related things. First, it has located content analysis historically, from Lasswell's WWII propaganda studies through Berelson's 1952 codification to Krippendorff's contemporary synthesis, and clarified what makes it methodologically distinctive. Content analysis is the qualitative-quantitative bridge: it begins with the same kind of qualitative judgement that drives thematic analysis (an earlier module) but turns the resulting codes into variables and applies the inferential machinery of earlier courses. The single operational move, treating codes as variables, is what lets a 20-transcript qualitative corpus answer questions that thematic analysis alone cannot: how often, how distributed, larger than chance.

Second, the lesson has given you the operational machinery to do content analysis on qualitative data such as the loneliness transcripts. A codebook is the input; the Taguette export is the bridge; the six-step workflow turns long-format coding data into wide-format codes × cases matrices, computes descriptive and inferential statistics, and produces publication-quality visualisations. The worked analysis, a coded subset, a frequency table, a chi-squared or Fisher's exact comparison, and an interpretive memo, is the course's first piece of hybrid quantitative-qualitative analysis.

What you take away from this lesson sets up the lessons that follow. Schema and narrative analysis moves in a different direction: not counting codes across many transcripts but tracing the cognitive structures inside individual transcripts. Discourse analysis deepens the attention to language and power. Analytic induction, QCA, and decision models return to systematic comparison with a different toolkit. Computational text and LLM analysis extends today's dictionary-based and keyness moves to the scale of millions of documents.

Key Takeaways from this lesson

  • Content analysis has a 90-year history: Lasswell (1927, 1942) established it as a method for propaganda analysis; Berelson (1952) codified it as the systematic, quantitative description of manifest content; Krippendorff (1980–2018) modernised it to accommodate latent meaning, inference, and rigorous reliability statistics.
  • Manifest content is what the text literally says; latent content is what it means. Most contemporary content analysis blends both, with the analyst declaring which is which and demonstrating reliability for the latent codes.
  • The codes-as-variables move makes content analysis a hybrid method: qualitative judgement defines the codes; quantitative procedures analyse them. Both halves are necessary; neither is optional.
  • Three unit-types structure a content-analytic design: sampling units (what is drawn from the population), recording units (where the variable takes its value), and context units (the surrounding text consulted during coding).
  • Reliability is load-bearing: Krippendorff's alpha is the field standard; α ≥ 0.80 supports definitive claims, 0.67–0.80 is acceptable for tentative inference, and below 0.67 the codebook must be revised.
  • Inferential tests on codes × subgroup matrices, chi-squared and Fisher's exact, are legitimate when the question is whether a within-corpus distributional difference is larger than chance, not when it is generalising to a population.
  • Dictionary-based content analysis (LIWC, sentiment dictionaries) and computational content analysis (a later module) are scalable extensions of the same logic, valuable as complements to human coding rather than as replacements.

Core Concepts Reviewed

History and core concepts: The history of content analysis from Lasswell through Berelson to Krippendorff and the contemporary synthesis; the manifest/latent distinction and why contemporary content analysis blends both; the codes-as-variables operational move; when content analysis is the right tool.

Design and reliability: Sampling, recording, and context units; exhaustive, mutually exclusive, and multi-coded schemes; a priori vs. emergent codes; Krippendorff's alpha as the reliability standard with its 0.80 / 0.67 thresholds; the end-to-end workflow.

Analysis and inference: Frequency tables overall and by subgroup; chi-squared and Fisher's exact tests on codes × subgroup contingency tables; trend analysis; dictionary-based content analysis (LIWC, sentiment); preview of computational content analysis.

The analysis workflow: loading Taguette exports and joining participant variables; reshaping into codes × cases matrices; descriptive frequency tables overall and by subgroup; bar-chart visualisation; chi-squared / Fisher's exact tests; keyness analysis of two subcorpora; what a content analysis reports.

The final reflection below asks you to step out of method-mode and articulate where content analysis fits in your own methodological identity. There is no single right answer; the goal is to leave the lesson with a defensible stance on when this hybrid method belongs in your toolkit.

Reflection

Content analysis sits on the boundary between qualitative and quantitative research. Some methodological writers (especially in nursing and health-services research) treat it as a qualitative method that uses some counting; others (especially in communications and computational social science) treat it as a quantitative method applied to text. Where do you locate it, and where will you locate yourself, methodologically, after this module?

Model answerA strong response avoids the false choice. Content analysis is neither a qualitative method with numbers attached nor a quantitative method applied to text; it is genuinely hybrid, and the operational move that makes it hybrid, treating codes as variables, is what gives the method its analytic power. Bernard, Wutich, and Ryan's stance, which this course adopts, is that the boundary between qualitative and quantitative was always more rhetorical than real, and that content analysis is the methodological proof. For your own location: the productive answer is that you are becoming a methodological omnivore, a researcher who can do credible qualitative work and read quantitative work, both in the same person, and who can deploy a method like content analysis without feeling that you have betrayed either side. The other defensible answer is to commit primarily to the qualitative tradition while keeping content analysis in your toolkit for questions where distributional comparison is genuinely warranted. Either stance is defensible; neither requires you to disavow the other half of your training. What is not defensible is treating content analysis as something to be done apologetically, either as a watering down of qualitative work or as a half-hearted attempt at quantification.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment for this lesson: Content Analysis (15 Questions)

Question 1: Which figure is most associated with the development and naming of Krippendorff's alpha, the field-standard reliability statistic for content analysis?

Krippendorff developed the alpha statistic and authored the field's most-cited methodological text (1980, 2004, 2013, 2018). Lasswell led the WWII propaganda-analysis programme; Berelson codified content analysis in 1952; Pennebaker developed LIWC.

Question 2: The single operational move that makes content analysis a hybrid quantitative/qualitative method is:

Codes-as-variables is the hinge: qualitative judgement defines the codes; once defined, each code becomes a variable that can be analysed with the descriptive and inferential machinery of statistics.

Question 3: Berelson's 1952 definition restricted content analysis to:

Berelson restricted the method to manifest content. Krippendorff and the subsequent field reopened the door to latent meaning, with the condition that latent codes be applied reliably.

Question 4: In Krippendorff's terminology, the context unit is:

Context units are the surrounding text consulted while coding; recording units are where the variable takes its value; sampling units are what is drawn from the population.

Question 5: A code shows Krippendorff's α = 0.74 in a reliability subset. What is the appropriate response?

Krippendorff's convention places 0.67 ≤ α < 0.80 in the "tentative" range, report and flag, optionally revise. Below 0.67 the code is inadequate; at or above 0.80 it supports definitive claims.

Question 6: A code appears in 9 of 10 caregivers and 4 of 10 non-caregivers. Which test is the most defensible inferential choice?

Fisher's exact gives an exact p-value regardless of cell size and is the recommended alternative to chi-squared when expected cell counts may fall below the standard threshold.

Question 7: Which approach to building a content-analytic codebook starts with codes drawn from prior theory or instrument and applies them deductively to the data?

Hsieh and Shannon (2005) distinguish conventional (inductive/emergent), directed (deductive/a priori), and summative (word-count-derived) content analysis. Directed is the deductive variant.

Question 8: The Linguistic Inquiry and Word Count (LIWC) dictionary is best understood as an example of:

LIWC counts pre-specified word lists across a corpus. It is a manifest-content method that scales to millions of documents but is blind to latent meaning, ambiguity, and irony, valuable as a complement to human coding, not a replacement.

Question 9: The Bernard/Wutich/Ryan stance on the qualitative/quantitative boundary in content analysis is best summarised as:

Bernard, Wutich, and Ryan's stance is that the strict qualitative/quantitative separation Berelson drew was always more rhetorical than real, and that contemporary content analysis is best understood as a method that explicitly blends both traditions.

Question 10: A research question is "do younger participants and older participants describe loneliness differently?" Why is content analysis the right tool for this question?

Content analysis's distinctive payoff is in distributional and comparative questions: how often, in which subgroup, larger than chance. For "do X and Y describe Z differently?", content analysis is the natural tool.

Question 11: Which of the following is NOT one of the three unit-types Krippendorff distinguishes?

Krippendorff's three units are sampling, recording, and context. There is no separate "inferential unit"; inference operates over the recording-unit-level codes.

Question 12: Which of the following BEST describes the relationship between content analysis and the methods you will study later in the course (computational text and LLM analysis)?

Computational content analysis (topic models, embeddings, LLM-based coders) is the scalable extension of human content analysis. Bernard, Wutich, and Ryan's stance, and the course's, is that it complements rather than replaces human coding, and that any computational output must be validated against human-coded gold standards.
✦ Complete the final reflection above before submitting

Lesson complete

You have successfully completed this lesson: Content Analysis.

You can now locate content analysis in its disciplinary history, distinguish manifest from latent content, treat codes as variables, sample text systematically, develop a reliable coding scheme, compute Krippendorff's alpha, and test distributional hypotheses on codes × subgroup matrices. Together these skills make content analysis the first hybrid qualitative-quantitative method in the course.

Next up, a later lesson: Schema and Narrative Analysis, which moves from counting codes across many transcripts to tracing cognitive structures inside individual transcripts.

Continue to Lesson 9 →
Reference

Glossary: Key Terms, People & Methodological Stances

📚 Reference page: available throughout the lesson

This glossary collects the key concepts, people, and methodological stances introduced in this lesson. Use it as a reference while you work through the material, or as a review before the final assessment. Type in the search box to filter entries.

Core Concepts
Content Analysis A research technique for making replicable and valid inferences from texts (or other meaningful matter) to the contexts of their use (Krippendorff, 2018). In this course, content analysis is treated as the hybrid method that bridges qualitative and quantitative analysis by treating codes as variables.
Manifest Content The literal, surface-level content of a text, the words and phrases that are explicitly present. Manifest coding is fast, reliable, and limited in interpretive depth.
Latent Content The underlying meaning of a text that requires interpretation to surface. Latent coding is interpretively richer than manifest but harder to apply reliably; modern content analysis explicitly accommodates both.
Codes as Variables The operational move that defines content analysis. A code, once defined qualitatively, becomes a column in a data frame with a value for every recording unit. This unlocks descriptive and inferential statistical analysis of qualitative coding.
Sampling Unit The chunk of text drawn from the population, each interview transcript, each newspaper article, each tweet. Choice of sampling unit follows from the research question and the population definition.
Recording Unit The unit that is actually coded, where the variable takes its value. May be the whole transcript, the paragraph, the sentence, or the word/phrase. The choice shapes both what the analysis can measure and how reliable it will be.
Context Unit The surrounding text the coder is permitted to consult when deciding how to code a recording unit. A larger context unit increases interpretive richness; a smaller one increases reliability.
Coding Scheme (Codebook) The list of codes with operational definitions, inclusion rules, and exclusion rules that govern application. Must be exhaustive over the content of interest; codes may be mutually exclusive (classical Berelson) or allow multi-coding (modern compromise).
A Priori (Deductive) Codes Codes derived from theory, prior literature, or instrument before data analysis begins. Hsieh and Shannon (2005) call this directed content analysis.
Emergent (Inductive) Codes Codes that arise from close reading of the corpus. Hsieh and Shannon call this conventional content analysis. Most contemporary studies use a hybrid: a priori scaffold plus emergent expansion.
Dictionary-Based Content Analysis Manifest content analysis at scale, in which a pre-specified word list is applied by a computer to a corpus. LIWC and sentiment dictionaries are the best-known examples. Fast and replicable; blind to latent meaning, ambiguity, and irony.
LIWC (Linguistic Inquiry and Word Count) A dictionary system developed by James Pennebaker and colleagues (current version LIWC-22) that classifies words into approximately 90 categories (positive emotion, negative emotion, sadness, body, family, social processes, etc.). Widely used in psychology, health communication, and computational social science.
Keyness A manifest-content analysis that identifies words appearing disproportionately in one subcorpus compared to another, with a log-likelihood or chi-squared statistic for each word.
Reliability & Inference
Krippendorff's Alpha The field-standard reliability statistic for content analysis (Krippendorff, 2004, 2018). Accommodates any number of coders, any level of measurement, and missing data. Acceptable thresholds: α ≥ 0.80 for definitive claims; 0.67 ≤ α < 0.80 for tentative inference; below 0.67 the codebook should be revised.
Cohen's Kappa An older reliability statistic (Cohen, 1960; benchmarks from Landis & Koch, 1977) restricted to two coders and nominal data. Still reported in some literatures but largely superseded by Krippendorff's alpha for content-analytic work.
Chi-Squared Test of Independence The classical inferential test for whether the distribution of a categorical variable differs between groups. In content analysis, applied to codes × subgroup contingency tables. Assumes expected cell counts ≥ 5; use Fisher's exact instead when this fails.
Fisher's Exact Test An exact inferential test for contingency tables that gives accurate p-values regardless of cell size. The recommended replacement for chi-squared when expected cell counts may fall below the standard threshold, common in small qualitative corpora.
Methodological Variants
Conventional Content Analysis Hsieh and Shannon's (2005) name for inductive content analysis: codes emerge from close reading. Sensitive to participants' own categories but at risk of analyst preconceptions.
Directed Content Analysis Hsieh and Shannon's name for deductive content analysis: codes drawn from theory, prior literature, or instrument. Replicable and comparable but at risk of missing what the data are actually doing.
Summative Content Analysis Hsieh and Shannon's third type: starts with word counts and expands to latent interpretation. Bridges dictionary-based and human-coded approaches.
Qualitative Content Analysis An umbrella term used in nursing and health-services research for content-analytic work that emphasises interpretive depth over inferential machinery. Often refers to Hsieh & Shannon's (2005) typology or Schreier's (2012) framework; see also Elo & Kyngäs (2008) and Mayring (2000).
Computational Content Analysis The scalable extension of content analysis using topic models (LDA), word embeddings, or large-language-model-based coding. Developed later in the course. Valuable as a complement to human coding; requires validation against gold-standard human codes.
Key People
Harold D. Lasswell (1902–1978) Political scientist whose Propaganda Technique in the World War (1927) and WWII-era work at the U.S. Library of Congress Experimental Division established content analysis as a systematic social-science method. Famous formulation: "Who says what to whom in what channel with what effect."
Bernard Berelson (1912–1979) Author of Content Analysis in Communication Research (1952), the field's first textbook. His four-adjective definition, objective, systematic, quantitative, manifest, structured the method for two decades. Restricted content analysis to manifest content; this restriction was later relaxed by Krippendorff and the contemporary field.
Klaus Krippendorff (1932–2022) Communications scholar whose Content Analysis: An Introduction to Its Methodology (1980, 2004, 2013, 2018) is the field's most-cited methodological text. Broadened the method to accommodate latent content and developed Krippendorff's alpha, the field-standard reliability statistic.
James W. Pennebaker (b. 1950) Social psychologist who developed LIWC (Linguistic Inquiry and Word Count) and pioneered dictionary-based analysis of personal-narrative text. His work linking writing style to psychological state has been influential in health communication and digital-mental-health research.
Hsiu-Fang Hsieh & Sarah E. Shannon (2005) Authors of the highly cited Three Approaches to Qualitative Content Analysis (2005) in Qualitative Health Research. Their typology, conventional, directed, summative, has structured most qualitative content-analytic work in nursing and health-services research for the past two decades.
H. Russell Bernard, Amber Wutich, Gery W. Ryan Authors of Analyzing Qualitative Data: Systematic Approaches (2nd ed., 2017). Their Chapter 11 (pp. 243–268) is the source for this lesson's content. They present content analysis as the hybrid method bridging the qualitative and quantitative traditions.
No matching entries. Try a different search term.