# Lesson 11: First Steps in Analysis

*Companion-podcast transcript, Sarah and Kiffer*

---

**Sarah:** Welcome back to Office Hours. I'm Sarah.

**Kiffer:** And I'm Kiffer. This week we are on Lesson eleven, First Steps in Analysis, which is the lesson where the data come back and we have to do something with them.

**Sarah:** The lesson has two parts, one with numbers and one with words. Why put them together in a single week?

**Kiffer:** Because the first steps are surprisingly similar. With survey data, you prepare and document a file and then describe it. With interview data, you prepare transcripts, write down the rules for coding, and then do a first pass. In both cases the early work decides whether anyone can trust what comes later.

**Sarah:** We are still following the Cedar Valley Social Connection Study, which, for anyone new, is fictional.

**Kiffer:** That's right. It is an invented study of loneliness among adults aged sixty-five and older in a mostly rural health region in British Columbia. By now the team has one thousand six hundred completed surveys, twenty-four interviews and four focus groups.

**Sarah:** Let's start with Section one, on data entry and data dictionaries. What is the first thing the team should do when the survey file arrives?

**Kiffer:** Make a copy of the raw export, mark it read-only, and never edit it. Every change after that, whether it fixes a typing error or creates a new variable, happens in a script that reads the raw file and writes a separate analysis file.

**Sarah:** What goes wrong if you just fix things in the spreadsheet?

**Kiffer:** Nobody can tell afterwards what was changed or why, and a mistake cannot be undone. With a script, anyone can rerun every step and get the same analysis file, and if a cleaning decision turns out to be wrong, you change one line and run it again.

**Sarah:** The section also talks about the shape of a dataset.

**Kiffer:** Almost all statistical software expects a rectangular layout. Each row is one unit of observation, usually one participant, identified by a study number in place of a name. Each column is one variable, and each cell holds one value. Karl Broman and Kara Woo turned this into practical rules: one header row, consistent codes, dates written year first, and no colour or merged cells used to carry information, because statistical software reads only the values in the cells.

**Sarah:** Some of the Cedar Valley surveys were done on paper.

**Kiffer:** Five hundred and sixty of them, typed by hand into the same REDCap form. The standard protection for hand-typed forms is double data entry. Two people type the same forms independently, a program compares the files cell by cell, and every disagreement is checked against the paper. Cedar Valley used the lighter check from Lesson eight, re-entering a random tenth of the forms.

**Sarah:** And the online surveys?

**Kiffer:** The team used REDCap, which stands for Research Electronic Data Capture. It applies validation rules as people answer, so an age of six in a study of older adults is refused on the spot. A rule only catches the errors someone anticipated, though, so the export still needs checking.

**Sarah:** Then the data dictionary. How is that different from a codebook?

**Kiffer:** In quantitative work the two words often mean the same thing. A data dictionary describes every variable: its name, a label giving the full question, the type of data, the permitted values and what each code means, the missing codes, and where the variable comes from. We call it a data dictionary on purpose, because in Section three the word codebook means something else.

**Sarah:** Give me an example of an entry.

**Kiffer:** Take the first loneliness item. Its label is the question, how often do you feel that you lack companionship. It is ordinal, coded one for hardly ever, two for some of the time and three for often. The missing codes are minus eighty-eight for prefer not to answer and minus ninety-nine for not answered. Its source is the three-item loneliness scale from the University of California, Los Angeles, published by Hughes and colleagues in two thousand and four.

**Sarah:** The lesson makes a point about derived variables.

**Kiffer:** A derived variable is created from other variables. The loneliness total is the sum of the three items, so it runs from three to nine, and the Cedar Valley protocol counts a score of six or higher as lonely. Neither the total nor the lonely variable appears on the survey, so the dictionary states the rule that creates them, including what happens when an item is missing.

**Sarah:** Which brings us to missing codes, and there is a nice example in the reading.

**Kiffer:** Four participants are seventy-two, sixty-eight, eighty-one and seventy-five, so their mean age is seventy-four. Add a fifth person whose age was recorded with the code minus ninety-nine. If the software treats that code as a number, the mean of the five values is thirty-nine point four years.

**Sarah:** Which looks like a perfectly normal adult age.

**Kiffer:** That is the danger. An obviously impossible average would be caught at once, but thirty-nine point four could pass into a report. So the cleaning script converts every missing code to the value the software recognizes as missing. Many programs display that value as N A, short for not available. You keep the distinct reasons in the raw file and report the number missing in your tables.

**Sarah:** Then comes the part some students will be nervous about, which is working with a large real dataset for the first time.

**Kiffer:** The lesson takes it slowly. Before any numbers appear, you look at the shape of the file, how many rows and columns it has, and what each column is supposed to hold.

**Sarah:** What should a student do in their first five minutes?

**Kiffer:** Keep one folder for the project, with the raw file, the data dictionary and a written log of every change. Then check that the file you received matches what the dictionary says it should contain. I want students to keep that record from the start, because the record is what lets anyone else trust the numbers.

**Sarah:** The first worked example is simple arithmetic.

**Kiffer:** It is. You take the Cedar Valley counts, one thousand six hundred surveys and three hundred and ninety-two people classed as lonely, divide one by the other, and round the result to twenty-four point five percent. The habit to build is writing down the numerator and the denominator every time, so that a reader can check the percentage.

**Sarah:** And then the lesson turns to real data.

**Kiffer:** Because Cedar Valley is invented, we practise on the Canadian Social Connection Survey, a national online survey of social connection, loneliness and health among people living in Canada. A de-identified public version is freely available online. We keep the twenty twenty-one wave, which has four thousand and forty-five respondents and three thousand two hundred and forty-seven columns.

**Sarah:** Nobody memorizes three thousand columns.

**Kiffer:** No. The names have prefixes for each part of the survey, and the file includes a labels table that works as a data dictionary, so students look up the wording of a question by its variable name. One lookup, for household size, returns nothing for twenty twenty-one. I left that in on purpose, because it shows what an undocumented variable looks like. The right response is to find out how the value was produced before using it.

**Sarah:** And the survey has its own way of recording missing answers.

**Kiffer:** It has two. When someone saw a question and left it blank, the survey records a category called presented but no response. When there is no answer at all, for example because the person stopped earlier, the file records the value as not available. On the item about feeling left out, twenty-four people saw the question and skipped it, and four hundred and twenty-one have no recorded answer.

**Sarah:** Let's move to Section two, which ends with a Table one. What is a Table one?

**Kiffer:** It is the descriptive table that comes first in most health research articles. It tells readers who was studied, usually overall and separately for the groups being compared. Readers need that before they can judge whether a finding applies to anyone else.

**Sarah:** How does a student decide what to put in each row?

**Kiffer:** The level of measurement decides the summary. Categorical variables, such as gender or self-rated health, get counts and percentages. Numeric variables need a centre and a spread. For roughly symmetric ones like age, you use the mean and the standard deviation. For skewed ones like emergency department visits, you use the median and the interquartile range.

**Sarah:** Explain why skew matters.

**Kiffer:** The reading uses nine invented participants with zero, zero, zero, zero, one, one, two, three and eleven visits. The total is eighteen, so the mean is two. The median, the middle value, is one. One person with eleven visits pulls the mean upward, and the median gives a better picture of the typical participant.

**Sarah:** And the standard deviation, in plain words?

**Kiffer:** It tells you how far values typically sit from the mean. If everyone is close to the average, it is small, and if values are spread out, it is large. The interquartile range does a similar job for the median, holding the middle half of the data.

**Sarah:** In the worked examples, you keep people aged sixty-five and older. Why?

**Kiffer:** So the practice data resemble the Cedar Valley population. There are four hundred and eighty-six of them. Two hundred and twenty-four scored six or higher on the loneliness scale, two hundred and nine scored lower, and fifty-three had no score because at least one item was missing.

**Sarah:** So what percentage were lonely?

**Kiffer:** It depends on the denominator, and that is the point. A frequency table usually leaves out missing values, so it divides by four hundred and thirty-three and gives fifty-one point seven percent. That is a valid percentage. Keep all four hundred and eighty-six in the denominator and you get forty-six point one percent, which understates the proportion among those who answered. Your table has to say which one you used.

**Sarah:** Fifty-one point seven percent is about double the Cedar Valley figure of twenty-four point five. Should students conclude that half of older Canadians are lonely?

**Kiffer:** No, and this is one of the most useful moments in the lesson. The Canadian Social Connection Survey recruited people online and did not draw them at random from a sampling frame, so its percentage describes these respondents. It does not estimate how common loneliness is among older adults in Canada, for the reasons Lesson seven explained. And the Cedar Valley figure is invented.

**Sarah:** Then the mean of the loneliness score, which gives a surprising result the first time.

**Kiffer:** The missing values have to be set aside first, because an average cannot include unknown numbers. Among the people who answered, the mean is about five point six five and the median is six. The same idea applies to the category for people who left a question blank. For relationship status, ten older respondents did that, and if you leave the category in, it appears as a group in every table. So those values are changed to not available and the empty category is dropped.

**Sarah:** Then cross-tabulations, which seem to confuse people.

**Kiffer:** Usually about which percentage to use. A cross-tabulation counts people in every combination of two categories, here gender in the rows and loneliness group in the columns. Among respondents classed as lonely, seventy point one percent were women. That is a column percentage. Among women, fifty-three point six percent were classed as lonely. That is a row percentage.

**Sarah:** Those sound like the same fact.

**Kiffer:** They answer different questions. The column percentage tells you what the lonely group looks like, which is what a Table one needs. The row percentage tells you how common loneliness is within a gender group. Students should always ask which total a percentage is divided by.

**Sarah:** And there is a warning about small cells.

**Kiffer:** There are eight non-binary respondents in this age group, and only one was not classed as lonely. A single person moves that row's percentage by twelve and a half points, and a cell of one can make someone identifiable in a small community. Many data custodians set a minimum cell size for published tables, so combining or suppressing small cells are the usual solutions.

**Sarah:** Walk us through building the Table one itself.

**Kiffer:** There are five decisions. Define the analytic sample, here the four hundred and thirty-three respondents with a loneliness score, with the fifty-three exclusions reported in a footnote. Choose the columns, one overall and one for each loneliness group. Choose the rows, ideally guided by your causal diagram. Choose the summary for each row. And report the missing data. The STROBE statement, which stands for Strengthening the Reporting of Observational Studies in Epidemiology, asks for exactly that.

**Sarah:** What does the finished table show?

**Kiffer:** Mean age is seventy-one point three years in both groups. Respondents classed as lonely were more often single and not dating, forty-six point three percent compared with thirty-two point two. They were also more often in poor physical health by their own rating, ten point three percent compared with one point one.

**Sarah:** So loneliness makes people's health worse.

**Kiffer:** The table cannot tell us that. It describes differences in this sample, and it contains no statistical tests. It cannot show whether loneliness affects health, whether poor health leads to loneliness, or whether something else drives both. Inferential comparisons come in Health Sciences three forty-one and four ten, and causal reasoning belongs to Health Sciences three forty-one. And every number in the table was taken straight from the analysis output, a habit I want students to keep.

**Sarah:** Let's turn to Part two. Section three is about building a first codebook and coding a transcript. What is a code, in the simplest terms?

**Kiffer:** A code is a short label you attach to a passage of text to record what it is about, or what it means for your question. The passage is called a coded segment, and one segment can carry several codes. Once everything is coded, you can pull together every passage about transportation from all twenty-four interviews and read them side by side.

**Sarah:** And memos go along with coding.

**Kiffer:** A memo is a short dated note recording an idea, a question or a connection you notice. Analytic memos continue the reflexive memos from Lesson ten, and they often contain the first drafts of later themes.

**Sarah:** The lesson says it stops at first-cycle coding.

**Kiffer:** Johnny Saldaña distinguishes first-cycle coding, the initial labelling of segments, from second-cycle coding, where codes are grouped into categories and themes. A theme is a pattern of meaning that runs across many participants. Building themes is what Health Sciences eight forty-one teaches. Here we do the first cycle well.

**Sarah:** Where do the codes come from?

**Kiffer:** From two directions. Deductive codes are written before coding, from the research question, the interview guide, the causal web or a theory. Robert Weiss, for example, distinguished emotional loneliness from social loneliness, and a team could write a code for each. Inductive codes are created during coding, when a passage says something important that no code captures. An in vivo code is an inductive code that uses the participant's own words as its label.

**Sarah:** Tell us about the starter codebook.

**Kiffer:** A codebook here is the list of codes with the rules for applying them. Kathleen MacQueen and colleagues recommended that each entry give a name, a short definition, a full definition, when to use the code, when not to use it, and an example. The starter codebook holds the deductive codes and is treated as version one point zero.

**Sarah:** What codes did the Cedar Valley team start with?

**Kiffer:** Seven: moving, family, community, transport, health, technology and loneliness. Each has rules. Family is not used for friends or neighbours, who go under community. Moving is not used when a change in someone's social life is linked to something else, such as the death of a spouse.

**Sarah:** You keep coming back to the rule for when not to use a code.

**Kiffer:** Because beginners leave it out, and it is what keeps similar codes apart. Without it, two coders will put the same passage under different codes.

**Sarah:** Then the lesson codes an excerpt from an interview with a participant called Ruth. Who is she?

**Kiffer:** She is invented, and the excerpt was written for teaching. Ruth is seventy-eight, lives alone, and moved from Cedar City to Kestrel Lake about two years before the interview. In Cedar City she ran into people she knew at the bank or the pharmacy. In Kestrel Lake she would come home from the store having spoken only to the cashier. She did not tell her daughter she was lonely, because she did not want to be a burden.

**Sarah:** And then her situation changes.

**Kiffer:** After cataract surgery she stopped driving, and with no buses she could go a week without leaving the house. Then the librarian started a Tuesday coffee morning and phoned Ruth herself to invite her. Ruth says she would not have gone without that call, because she could not walk into a room full of strangers on her own. Now she goes every week, and a man from the group drives her to her eye appointments.

**Sarah:** How did the first coding pass handle it?

**Kiffer:** The deductive codes cover a lot. Her old street gets moving. The Sunday phone calls from her daughter get family. The winter without driving gets transport and health. The coffee morning gets community. The man who drives her to appointments carries three codes at once: community, transport and health.

**Sarah:** And the inductive codes?

**Kiffer:** The pass proposes four. Chance encounters, for the unplanned contact she lost. Concealing loneliness, for hiding her feelings from her daughter. Personal invitation, for the librarian's call. And small-town visibility, for her care about what she says at coffee. Each needs a definition and rules before it goes into version one point one.

**Sarah:** Couldn't you just stretch an existing code? Personal invitation sounds like community.

**Kiffer:** You could, and you would lose something. The coffee morning had a notice on the board, and Ruth had seen it. What made the difference was a person asking her. If that recurs in other interviews, it matters for how the health authority runs its programs, and folding it into community would hide it. So the coder writes a memo to watch for it in the other twenty-three interviews.

**Sarah:** What changes when a team codes together?

**Kiffer:** The team agrees on the codebook first. Two coders code the same few transcripts independently, compare them segment by segment, and discuss every disagreement, which usually reveals an unclear definition. Some teams calculate an agreement statistic, which Health Sciences eight forty-one covers. For tools, you can use paper, a spreadsheet, or software such as Taguette, which is free. Transcripts are de-identified first, and a cloud tool needs the permission of your consent form and data management plan.

**Sarah:** Section four is about analytic approaches. What is an approach?

**Kiffer:** It is a recognized way of moving from data to findings, with its own aims, steps and standards. Coding is common to most of them, but what you do with the codes afterwards depends on the approach. Ideally you choose it when you design the study, because it affects how many people you interview and how.

**Sarah:** The lesson covers seven. Start with the ones that work across many participants.

**Kiffer:** Thematic analysis looks for patterns of meaning across a dataset, in six phases set out by Virginia Braun and Victoria Clarke. Qualitative content analysis sorts text into categories and sometimes counts them, which suits open-ended survey answers. Framework analysis, from Jane Ritchie and Liz Spencer, summarizes each participant's coded data in a matrix with one row per participant, so cases and groups can be compared. Rapid qualitative analysis summarizes each interview in a structured template for findings needed within weeks.

**Sarah:** And the three that go deeper.

**Kiffer:** Grounded theory, from Barney Glaser and Anselm Strauss and later Kathy Charmaz, builds a theory of a social process through constant comparison and theoretical sampling, where later participants are chosen to test ideas from earlier ones. Interpretative phenomenological analysis, developed by Jonathan Smith and colleagues, examines how a few people make sense of a significant experience, case by case. Narrative analysis, described by Catherine Kohler Riessman, keeps stories whole.

**Sarah:** The reading shows what each approach would do with Ruth's interview. Which one stood out to you?

**Kiffer:** Narrative analysis, because Ruth tells her move as a story. There is a setting, the move to Kestrel Lake, a complication, the first winter, a turning point, the librarian's call, and a resolution, the weekly coffee morning. A narrative analyst keeps that account whole and asks how she structures it. A framework analyst would instead summarize her coded data in her row of the matrix, alongside the other Kestrel Lake participants.

**Sarah:** How do the approaches map onto the data collection methods from Lesson ten?

**Kiffer:** In-depth interviews suit almost every approach, and they are essential for phenomenological and narrative analysis, which need long personal accounts. Focus groups fit those two poorly, because group conversation interrupts each person's account, but they suit thematic, framework and rapid analysis. Short open-ended survey answers from many people suit content analysis and thematic analysis. The table in the reading is a guide drawn from common practice, and Health Sciences eight forty-one takes up the debates.

**Sarah:** How should someone actually choose?

**Kiffer:** Ask four questions. What does the research question ask for: patterns, counts, comparison, a process, individual experience or stories? What data will you have? When are findings needed? And who will do the analysis?

**Sarah:** So what did the Cedar Valley team decide?

**Kiffer:** Framework analysis for the twenty-four interviews, because the health authority wants to compare Cedar City with the smaller communities and four people will share the coding. For the focus group with clinic staff and community connectors, rapid analysis, because the health authority asked for an interim brief within six weeks.

**Sarah:** And the approaches they set aside?

**Kiffer:** Narrative analysis of the moving stories, and interpretative phenomenological analysis for a small group of recently widowed participants. They recorded both as possible follow-up studies, because each would need longer, less structured interviews and training the team does not have. Writing down that reasoning belongs in the protocol.

**Sarah:** Before we finish, how would an analyst put both parts together on a new study?

**Kiffer:** For the quantitative part, define the analytic sample, convert the missing codes, build the Table one row by row, check every number against its source, and write a footnote stating the analytic sample, the denominators and the exclusions.

**Sarah:** And for the qualitative part?

**Kiffer:** Read the first transcript in full, write a starter codebook of five to eight deductive codes with all six columns, and complete a first coding pass with memos, adding inductive codes where nothing fits. Then write a short paragraph naming the approach for the full set of transcripts, justified with the four questions. Lesson twelve shows how these results become a report.

**Sarah:** Any final advice?

**Kiffer:** Write your rules down before you apply them, whether they are rules for missing codes or rules for a code called personal invitation. That habit makes both halves of this lesson trustworthy, and it is what readers of your report will rely on.

**Sarah:** Thanks, Kiffer. That's it for this week's Office Hours.

**Kiffer:** Thanks, Sarah. See you for Lesson twelve.
