HSCI 207 · Lesson 6

Choosing an Approach and Managing a Research Project

Research Methods in Health Sciences

Learning objectives for this lesson:

  • Classify a research question by type and use a five-step decision map to select a suitable design family.
  • Explain how ethics, the frequency of the outcome or exposure, time, money and existing data narrow the choice of design.
  • Distinguish the convergent, explanatory sequential and exploratory sequential mixed-methods designs by purpose, timing, priority and point of integration.
  • Describe how mixed-methods studies integrate their strands through connecting, building, merging and embedding, and interpret a joint display.
  • Identify the standard sections of a research protocol and explain how the protocol supplies the ethics application, data access requests, the data management plan and the methods section.
  • Manage protocol versions and amendments, and explain why protocols are registered and published.
  • Build a Gantt chart with dependencies, milestones and a critical path, and assign roles with a RACI matrix.
  • Prepare a justified line-item budget and organize project files with a folder structure, a naming convention, version control and a data management plan.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University. It is the applied research methods course of the Public Health Assessment and Analysis series.

Lesson 6 · HSCI 207

Choosing an Approach and Managing a Research Project

This short walkthrough orients you before you work through the lesson at your own pace.

Research Methods in Health Sciences
Why this lesson

From a good question to a workable study

A question states what you need to know. A design states how you will find out.

Project management keeps the study on schedule, within budget and in control of its files and data.

The road map

Four sections

1 · Question to design

A five-step decision map leads from question type to design family.

2 · Mixed methods

Three core designs combine quantitative and qualitative strands.

3 · The protocol

Twelve sections record the plan, with versions and amendments.

4 · Managing the project

Gantt charts, roles, budgets, files and data management keep it on track.

Running case

The Cedar Valley Social Connection Study

This fictional mixed-methods study of loneliness among adults aged 65 and older is led by Dr. Maya Hart with the Cedar Valley Health Authority.

1,600survey responses
24interviews
4focus groups
24months of planned work
How to use this lesson

Read, try, then apply

  • Each section opens with a narrated walkthrough like this one.
  • Try-it tasks let you practise each skill on a short example.
  • Section 4 brings the pieces together in a project plan.
Section 1 of 5

From Question to Design: A Decision Map

⏱ Estimated reading time: 35 minutes
Section 1 of 5

From Question to Design

A five-step decision map leads from a structured question to a design family.

35 minutes
The principle

The question chooses the design

PECO or PICO

A comparison signals that groups will be compared, and an intervention signals that someone delivers it.

SPIDER

The sample, phenomenon of interest, design, evaluation and research type point toward qualitative methods.

A trial gives the strongest evidence about an intervention that can be assigned. It cannot tell you how common loneliness is.

Decision map

Five questions in order

  1. Step 1 asks whether existing studies already answer the question.
  2. Step 2 asks whether the question is about numbers, meanings, or both.
  3. Step 3 asks whether the exposure is assigned on purpose, and whether at random.
  4. Step 4 asks whether groups are compared to estimate an association.
  5. Step 5 asks what the unit is and when exposure and outcome are measured.

The map is adapted from Grimes and Schulz (2002).

Analytic designs

Timing separates the designs

Ecological

Groups such as regions are the units.

Cross-sectional

Exposure and outcome are measured at once.

Cohort

People are followed forward from exposure.

Case-control

The study looks back from the outcome.

HSCI 230 Lessons 3 to 6 teach these designs in depth.

Constraints

What narrows the choice

Ethics
A rare outcome
A rare exposure
Time
Money and staff
Existing data

No one can be randomized to loneliness, so questions about its effects use observational designs.

Cedar Valley (fictional)

Each question finds its family

Descriptive

How common is loneliness? In the survey, 392 of 1,600 respondents (24.5 percent) scored 6 or higher.

Analytic: cohort

Is loneliness associated with emergency department visits in the following year? Linked records count the later visits.

Qualitative

How do older adults living alone experience a move to a smaller town? Interviews explore it.

Carry forward

Several families in one study

The Cedar Valley questions fall in three design families, so the study needs a way to join its strands.

Section 2 introduces the three core mixed-methods designs and the ways their strands are integrated.

Learning Objectives for this section

  • Explain what a study design is and why the research question should determine the design.
  • Use a five-step decision map to move from a structured research question to a design family.
  • Describe the descriptive, analytic observational, experimental, qualitative and evidence synthesis families, and name the later course that teaches each in depth.
  • Explain how ethics, the frequency of the outcome or exposure, time, money and existing data narrow the choice of design.
  • Match the questions of the Cedar Valley Social Connection Study to design families.

1.1 What a Study Design Is

A study design is the overall plan for how a study gathers and compares information to answer its question. It specifies who will be studied, what will be measured, when each measurement will be taken, and whether the researchers will change anything about the participants' circumstances. Designs that share the same basic logic form a design family. Cohort studies, case-control studies and cross-sectional studies, for example, belong to the analytic observational family, because each compares groups of people without the researchers assigning anything to them.

Lessons 1 and 2 gave you the two inputs this section needs. Lesson 1 sorted research questions into four kinds: descriptive (what is happening and how often), explanatory (why something happens or whether one factor affects another), predictive (who is likely to experience something in the future) and exploratory (how people experience something that is not yet well understood). Lesson 2 gave the question a structure. A PECO question (population, exposure, comparison, outcome) or a PICO question (population, intervention, comparison, outcome) already contains much of its design, because the comparison signals that groups will be compared and an intervention signals that someone will deliver something. A SPIDER question (sample, phenomenon of interest, design, evaluation, research type) points toward qualitative methods.

The principle that organizes this section is that the question chooses the design. Beginning researchers often start with a familiar method, such as a survey, and then look for a question it can answer, which tends to produce findings that do not match what the team or its partners needed to know. A second common error is to treat designs as a ranked list headed by the randomized controlled trial. A well-conducted trial provides the strongest evidence about the effect of an intervention that can ethically be assigned, and it cannot tell you how common loneliness is or how older adults experience a move to a new town. The most suitable design fits the question and can be carried out well with the resources available.

Where the individual designs are taught

This lesson maps the design families and the decisions that lead to each. HSCI 230 teaches the designs in depth: Lesson 3 introduces observational studies, Lesson 4 covers case-control studies, Lesson 5 covers cohort studies and randomized trials, and Lesson 6 covers ecological studies. HSCI 230 Lesson 10 Section 1 adds the biases specific to randomized trials, and HSCI 826 Lessons 6 to 8 extend randomized and quasi-experimental designs at the graduate level.

1.2 A Decision Map in Five Steps

A decision map is a sequence of questions that leads from a research question to a design family. The version in the figure adapts the classification of study types published by Grimes and Schulz (2002) and adds steps for evidence synthesis, qualitative questions and mixed methods. Work through the steps in order and stop at the first exit that fits your question.

Start: your structured question 1. Do existing studies already answer the question well enough? Yes Evidence synthesis Review (HSCI 230 L2; HSCI 241) No 2. Does it ask about numbers, about meanings and experiences, or both? Meanings: qualitative (HSCI 841) Both: mixed methods (Section 2) Numbers 3. Is the exposure assigned on purpose, and is the assignment random? Yes Experimental family Random allocation: randomized trial Other allocation: quasi-experimental (HSCI 230 L5; HSCI 826 L6 to L8) No: observational 4. Does the study compare groups to estimate an association? No Descriptive family Prevalence survey, case series, surveillance (HSCI 230 L3) Yes: analytic 5. What is the unit, and when are exposure and outcome measured? Ecological Groups such as regions are the units (HSCI 230 L6) Cross-sectional Exposure and outcome measured at once (HSCI 230 L3) Cohort Exposure first, then follow forward (HSCI 230 L5) Case-control Outcome first, then look back at exposure (HSCI 230 L4)
The decision map leads from a structured research question to a design family through five questions, adapted from the classification of study types by Grimes and Schulz (2002); the four red boxes at the bottom are the main analytic observational designs.

Step 1: Do existing studies already answer the question?

Before planning new data collection, check whether the answer already exists. If many studies have addressed the question, a systematic or scoping review may be the most useful project. The three to five key papers you learned to find in Lesson 2 are usually enough to tell whether a new study is needed.

Step 2: Does the question ask about numbers, about meanings, or about both?

Questions about how many, how much, how often, or how strongly two factors are related call for quantitative data. Questions about how people experience something, what it means to them, or how a process unfolds call for qualitative data. Questions with parts of both kinds call for mixed methods (Section 2). Two Cedar Valley questions show the split. "What proportion of older adults are lonely?" asks for a number, while "How do older adults living alone experience social connection after a move to a smaller town?" asks about meaning and experience.

Step 3: Is the exposure or intervention assigned on purpose, and is the assignment random?

Some exposures and interventions are assigned on purpose, either by the research team or by another agent, such as a government, a health authority or a program. Studies of deliberately assigned interventions form the experimental family. When the assignment uses a random process, such as a computer-generated allocation sequence, the study is a randomized controlled trial. When the assignment is deliberate and not random, the study is quasi-experimental. This includes a team offering a program at clinics it selects and comparing them with other clinics, and it also includes natural experiments and policy evaluations, in which a government or program decides who is exposed and the researchers compare groups or time periods, for example before and after a new policy. If no one assigns the exposure and the team only records exposures that people already have, such as loneliness or income, the study is observational, and you move to Step 4.

Step 4: Does the study compare groups to estimate an association?

An observational study with no comparison group is descriptive. It describes how often a condition occurs and how it is distributed by person, place and time. Prevalence surveys, case reports, case series and surveillance systems belong here. An observational study that compares groups to estimate whether an exposure is associated with an outcome is analytic, and you move to Step 5.

Step 5: What is the unit, and when are exposure and outcome measured?

If the units of analysis are groups, such as health regions or neighbourhoods, the study is ecological. If the units are individuals, the timing of measurement separates the three main designs. A cross-sectional study measures exposure and outcome at the same time. A cohort study starts with people who differ in their exposure and follows them forward to see who develops the outcome. A case-control study starts with people who have the outcome (cases) and people who do not (controls), and looks back to compare their earlier exposures.

Predictive questions use the same map with a different purpose: a prediction study estimates who is at risk of an outcome without asking whether each predictor causes it. Such studies usually draw on cohort data, because the predictors must be measured before the outcome occurs.

1.3 The Design Families at a Glance

The tabs summarize each family, the questions it answers, an example, and the course in which you will study it in depth.

What it answers. How common is a condition, and how is it distributed by person, place and time?

Typical designs. Cross-sectional prevalence surveys, case reports and case series, and surveillance systems that count cases over time.

Example. The Cedar Valley regional survey estimates the proportion of adults aged 65 and older who are lonely.

Strengths and limits. Descriptive studies are often the fastest and least costly design. Because they make no comparison, they cannot show whether one factor affects another.

Taught in depth. HSCI 230 Lesson 3, with the measures of frequency in HSCI 341 Lesson 4.

What it answers. Is an exposure associated with an outcome, and how strongly?

Typical designs. Analytic cross-sectional studies, cohort studies, case-control studies and ecological studies.

Example. The Cedar Valley PECO question asks whether loneliness is associated with emergency department visits.

Strengths and limits. These designs can study exposures that cannot be assigned, such as loneliness or income. Exposed and unexposed groups may differ in other ways, which is the problem of confounding that HSCI 341 Lesson 7 addresses.

Taught in depth. HSCI 230 Lessons 3 to 6.

What it answers. Does an intervention or policy change an outcome when its assignment is deliberate, whether the research team or another agent such as a government decides who receives it?

Typical designs. Individually randomized controlled trials, cluster randomized trials in which clinics or communities are allocated, and quasi-experimental designs such as non-randomized comparison groups, interrupted time series, natural experiments and policy evaluations.

Example. A future Cedar Valley study could allocate clinics to offer a community connector program and compare loneliness scores after one year.

Strengths and limits. Random allocation balances known and unknown confounders between groups on average. Trials are costly, and they can only test interventions that can ethically be assigned.

Taught in depth. HSCI 230 Lesson 5 Sections 5 to 7 and HSCI 230 Lesson 10 Section 1 teach randomized trials, and HSCI 826 Lessons 6 to 8 are the graduate extension.

What it answers. How do people experience, understand and act in a situation, and how does a process unfold?

Typical approaches. Qualitative description, phenomenology, grounded theory, ethnography, case study and narrative inquiry.

Example. The Cedar Valley SPIDER question asks how older adults living alone experience social connection after a move to a smaller town.

Strengths and limits. Qualitative studies capture meaning, context and process in participants' own words. They do not estimate how common an experience is in a population.

Taught in depth. Lessons 10 and 11 of this course introduce interviews, focus groups and coding, and HSCI 841 develops them.

What it answers. What does the body of existing research show about a question?

Typical designs. Systematic reviews, meta-analyses and scoping reviews.

Example. A scoping review could map published programs that address loneliness among rural older adults before the Cedar Valley team designs one.

Strengths and limits. A synthesis uses existing studies efficiently and shows where evidence is missing. Its conclusions can be no stronger than the studies it includes.

Taught in depth. HSCI 230 Lesson 2 and HSCI 241.

1.4 Constraints That Narrow the Choice

The decision map tells you which designs could answer a question. Ethical and practical constraints then tell you which of those designs you can carry out. Select each card to read about one of six constraints that researchers weigh at this stage.

EthicsClick to explore
A rare outcomeClick to explore
A rare exposureClick to explore
TimeClick to explore
Money and staffClick to explore
Existing dataClick to explore

Earlier lessons add two further considerations. Interest holders may value some questions more than others (Lesson 4), and where a study involves First Nations, Inuit or Métis communities, the design must respect the community's authority over its data and the agreements reached with it.

1.5 Applying the Map to the Cedar Valley Study

Case: The Cedar Valley Social Connection Study (fictional)

The Cedar Valley Social Connection Study is a fictional mixed-methods study planned by Dr. Maya Hart's team at a British Columbia university with the fictional Cedar Valley Health Authority, which serves about 210,000 residents, of whom about 46,000 are aged 65 and older. By the end of Lesson 3, the team had a PECO question, a SPIDER question, a causal web and a first DAG. It now has to decide which designs will answer its questions. The team runs each question through the decision map and records the result in a table.

Cedar Valley questionQuestion typeDesign family and designTaught in depth
What proportion of adults aged 65 and older in Cedar Valley are lonely?DescriptiveDescriptive: a cross-sectional regional survey (1,600 responses, of whom 392, or 24.5 percent, scored 6 or higher on the three-item UCLA Loneliness Scale)HSCI 230 Lesson 3
Among adults aged 65 and older, is loneliness, compared with lower loneliness scores, associated with emergency department visits in the following twelve months?ExplanatoryAnalytic observational: a cohort design in which the survey measures loneliness and linked records count later emergency department visitsHSCI 230 Lesson 5
How do older adults living alone experience social connection after a move to a smaller town?ExploratoryQualitative: 24 semi-structured interviews analyzed with framework analysisLessons 10 and 11; HSCI 841
How do clinic staff and community connectors view referral of isolated older adults to community supports?ExploratoryQualitative: a focus group with clinic staff and community connectorsLesson 10; HSCI 841
Which older adults are most likely to visit an emergency department in the next year?PredictivePrediction study using the same cohort data (a possible later analysis)HSCI 410
Does a community connector program reduce loneliness?Explanatory (an intervention)Experimental: a cluster randomized or quasi-experimental study (a possible future study)HSCI 230 Lesson 5 Sections 5 to 7 and Lesson 10 Section 1; HSCI 826 Lessons 6 to 8

Two features of the table deserve attention. First, the main questions fall in different families, which is why the Cedar Valley study uses mixed methods (Section 2). Second, the PECO question becomes a cohort design because the team links each survey response to emergency department visits that occur after the survey date. If the team compared loneliness with visits in the year before the survey, the same data would support only a cross-sectional comparison, and the team could not tell whether loneliness came first. Decisions about timing of this kind are written into the protocol (Section 3).

Try it: classify four questions

Run each question through the decision map, name the design family and the most likely design, and then compare your reasoning with the answers below.

(1) Among adults aged 50 and older in British Columbia, is regular volunteering associated with a lower rate of depression diagnosed over the next five years? (2) How do community pharmacists decide whether to raise social isolation with older patients? (3) What proportion of long-term care residents in a health region received an influenza vaccine this season? (4) Does a peer-led walking group, offered at six community centres chosen at random from twelve, reduce loneliness compared with the other six centres?

Answer to question 1v

The question is explanatory and asks for numbers. No one assigns volunteering, so the study is observational, and it compares volunteers with non-volunteers, so it is analytic. Volunteering is measured before depression is diagnosed, so the design is a cohort study.

Answer to question 2v

The question asks how pharmacists make a decision, which is a question about a process and its meaning to the people involved. It belongs to the qualitative family, and semi-structured interviews with pharmacists would suit it.

Answer to question 3v

The question asks how common something is at one point in time and makes no comparison. It is descriptive, and a cross-sectional count from immunization records or a survey of facilities would answer it.

Answer to question 4v

The investigators assign the program and use a random process to choose the centres, so the study is experimental. Because whole centres are allocated, it is a cluster randomized trial, which HSCI 230 Lesson 5 introduces and HSCI 826 Lesson 6 develops at the graduate level.

Reflection

A public health team in a mid-sized health region wants to study falls among adults aged 75 and older who live at home. It has drafted four questions.

(a) What proportion of adults aged 75 and older living at home fell in the past twelve months? (b) Is the use of sleeping medication associated with hip fracture? Hip fracture is relatively rare, and the region holds linked pharmacy and hospital records going back ten years. (c) How do older adults who have fallen describe their fear of falling again? (d) Does a home safety visit, offered to half of the eligible households chosen by computer-generated random allocation, reduce falls?

The decision map asks five questions in order: whether existing studies already answer the question; whether the question asks about numbers, about meanings and experiences, or both; whether the exposure is assigned on purpose, by the investigator or by another agent, and whether at random; whether the study compares groups to estimate an association; and what the unit is and when exposure and outcome are measured (cross-sectional: at the same time; cohort: exposure first, then people are followed forward; case-control: the study starts from people with and without the outcome and looks back at exposure). The four question types are descriptive, explanatory, predictive and exploratory.

For each question, name the question type and the design family and explain your reasoning in one or two sentences. For question (b), explain whether a cohort or a case-control approach would be more efficient and why.

Model answer

(a) This question is descriptive. It asks for a number, no one assigns anything and there is no comparison, so it belongs to the descriptive family; a cross-sectional survey of older adults living at home that asks about falls in the past twelve months would answer it.

(b) This question is explanatory. It asks about numbers, the exposure is not assigned and users are compared with non-users, so it is analytic observational. Because hip fracture is relatively rare, a case-control approach is efficient: the team identifies people with a hip fracture in the hospital records and a sample of controls without one, and compares their earlier pharmacy records. Because the linked records already cover ten years, a retrospective cohort comparing users and non-users would also be feasible without new data collection, and the choice may depend on the cost of extracting records.

(c) This question is exploratory. It asks how people describe an experience, so it belongs to the qualitative family, and semi-structured interviews with older adults who have fallen would suit it.

(d) This question is explanatory and concerns an intervention. The team assigns the visits by random allocation, so the study is experimental, specifically a randomized controlled trial. If the team had offered visits only in neighbourhoods it chose, the study would be quasi-experimental, and the same would be true if the health region had chosen the neighbourhoods.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A researcher asks, "What proportion of adults aged 65 and older in a health region report feeling lonely?" Which design family fits this question?

The question asks for a proportion at one point in time and makes no comparison, so it is descriptive, and a cross-sectional survey would answer it. Choosing who receives a survey is sampling, which differs from assigning an exposure, so the study is not experimental.

Question 2: In the decision map, which question separates the experimental family from the observational family?

A study belongs to the experimental family when the exposure or intervention is assigned on purpose, by the investigator or by another agent such as a government, and it is observational when no one assigns the exposure and the team only records exposures people already have. Comparison groups and the timing of measurement distinguish designs within the observational family.

Question 3: A team wants to study risk factors for a rare childhood cancer with limited time and funding. Which design is usually most efficient?

A case-control study starts from the cases that already exist and compares them with controls, which is efficient for a rare outcome. A prospective cohort would need to follow a very large number of children for a long time to observe enough cases.

Question 4: Why can the Cedar Valley team not use a randomized controlled trial to test whether loneliness causes emergency department visits?

A trial requires the team to assign the exposure, and no one can ethically or practically be made lonely. Questions about the effects of exposures that cannot be assigned use observational designs, such as the cohort design the team uses with linked records.
Section 2 of 5

Mixed-Methods Designs: Convergent, Explanatory Sequential and Exploratory Sequential

⏱ Estimated reading time: 35 minutes
Section 2 of 5

Mixed-Methods Designs

Convergent, explanatory sequential and exploratory sequential designs, and how their strands are integrated.

35 minutes
Integration

What makes a study mixed methods

Integration separates a mixed-methods study from a study that happens to contain two kinds of data.

Triangulation
Complementarity
Development
Initiation
Expansion

These are the five purposes for combining methods described by Greene, Caracelli and Graham (1989).

Notation

Timing and priority in a few symbols

QUAN + QUAL

Both strands run together with equal weight.

QUAN → qual

A main quantitative strand is followed by a supporting qualitative strand.

QUAL → quan

A main qualitative strand is followed by a supporting quantitative strand.

In Morse’s (1991) notation, capitals mark priority, a plus sign means concurrent strands, and an arrow means sequential strands.

Three core designs

Order and purpose

Convergent

Both strands are collected in one phase and merged to compare results.

Explanatory sequential

Numbers come first, and qualitative data help explain them.

Exploratory sequential

Qualitative findings come first and build a quantitative strand.

Creswell and Plano Clark (2018) describe these three core designs.

Integration methods

Four ways the strands meet

Connecting

One strand links to the other through sampling.

Building

One strand’s results shape the other’s data collection.

Merging

Results are brought together after analysis.

Embedding

Strands are linked at several points in a larger design.

These procedures follow Fetters, Curry and Creswell (2013).

Joint display

Reading the strands side by side

34.9%of respondents living alone scored 6 or higher
18.1%of respondents living with others scored 6 or higher

Interview participants living alone described evenings, weekends and meals as the hardest times. The two strands agree, which is confirmation.

Carry forward

Cedar Valley chooses QUAN → QUAL

  • The survey asks for consent to be contacted again and for years at the current address.
  • The final interview guide goes to the REB as an amendment.
  • The qualitative strand begins only after the survey closes.

Learning Objectives for this section

  • Define mixed-methods research and explain the five purposes for combining quantitative and qualitative strands.
  • Read and write the notation that shows the timing and priority of the strands in a mixed-methods study.
  • Distinguish the convergent, explanatory sequential and exploratory sequential designs by purpose, sequence and point of integration.
  • Describe the four methods of integration (connecting, building, merging and embedding) and read a joint display.
  • Justify a mixed-methods design for a study, using the Cedar Valley study as a model.

2.1 What Makes a Study Mixed Methods

Mixed-methods research combines quantitative and qualitative data within one study and integrates them to answer the research questions. Creswell and Plano Clark (2018) describe its core characteristics: the researcher collects and analyzes both kinds of data rigorously, integrates the two forms of data and their results, organizes these procedures into a recognizable design, and frames the work within theory and a stated philosophical position. Each part of the study that uses one kind of data is called a strand. The Cedar Valley study has a quantitative strand (the survey, the linked administrative data and the chart review) and a qualitative strand (the interviews and focus groups).

Integration is the feature that separates a mixed-methods study from a study that happens to contain two kinds of data. A study that runs a survey and a set of interviews and reports them in separate chapters, without relating one to the other, has collected two kinds of data and has not yet produced a mixed-methods analysis. Integration means that the results of one strand shape the other strand, or that the two sets of results are brought together and compared, so that the study as a whole says more than either strand could say alone.

Why combine the two kinds of data?

Greene, Caracelli and Graham (1989) reviewed published mixed-methods evaluations and identified five purposes for combining methods. A clear statement of purpose helps a team choose its design, because each purpose implies a different sequence. Select each card for a definition and a Cedar Valley example.

TriangulationClick to explore
ComplementarityClick to explore
DevelopmentClick to explore
InitiationClick to explore
ExpansionClick to explore

2.2 Timing, Priority and Notation

Two decisions describe the structure of any mixed-methods study. The first is timing: the strands may run at the same time (concurrently) or one after the other (sequentially). The second is priority, sometimes called weighting: the strands may carry equal weight in answering the study's questions, or one strand may be the main source of evidence while the other plays a supporting role. Morse (1991) introduced a short notation for these decisions that is still widely used.

Mixed-methods notation (Morse, 1991)

QUAN and QUAL in capital letters mark a strand with priority, and quan and qual in lower case mark a supporting strand. A plus sign (+) means the strands run at the same time, and an arrow (→) means one strand follows the other. Parentheses mark a strand embedded within a larger design.

NotationHow to read itExample
QUAN + QUALBoth strands run at the same time and carry equal weight.A survey of clinic patients and interviews with clinic staff on the same topics, run in the same months.
QUAN → qualA main quantitative strand is followed by a supporting qualitative strand.A regional survey followed by a small number of interviews to explain one unexpected result.
QUAL → quanA main qualitative strand is followed by a supporting quantitative strand.Focus groups that generate survey items, followed by a survey that checks how common the reported views are.
QUAN(qual)A qualitative strand is embedded within a larger quantitative design.Interviews with participants inside a randomized trial to learn how they experienced the intervention.

2.3 Three Core Designs

Creswell and Plano Clark (2018) describe three core designs that combine timing, priority and integration in recognizable ways: the convergent design, the explanatory sequential design and the exploratory sequential design. Most mixed-methods studies in the health sciences use one of them, either alone or as a building block within a larger study. The figure shows the order of steps in each, and the tabs describe when to choose each design and what it demands of the team.

Convergent design (QUAN + QUAL) Quantitative data collection and analysis Qualitative data collection and analysis Merge: compare the two sets of results Interpret where results agree, expand or differ Explanatory sequential design (QUAN → qual) Quantitative collection and analysis Connect: choose participants and build the guide Qualitative collection and analysis Interpret how the qualitative results explain the numbers Exploratory sequential design (QUAL → quan) Qualitative collection and analysis Build: items, an intervention or hypotheses Quantitative collection and analysis Interpret how far the qualitative findings generalize Quantitative strand Qualitative strand Integration step Interpretation Capital letters mark the strand with priority; lower case marks a supporting strand.
The three core mixed-methods designs differ in the order of their strands and in the step at which the strands are integrated, following Creswell and Plano Clark (2018).

Purpose. The convergent design collects quantitative and qualitative data in the same phase, analyzes each set separately, and then merges the results to compare them. It suits triangulation and complementarity, when a team wants a fuller picture of one topic from two angles.

Procedure. The team plans parallel questions in both strands so that the results address the same topics, collects both kinds of data, analyzes each with its own methods, and compares the results topic by topic.

Sampling. The strands usually have different sample sizes, which is expected because they serve different purposes. A survey might have several hundred respondents while the interviews have twenty participants.

Demands on the team. The design is the fastest of the three, but the team needs quantitative and qualitative skills at the same time, and it must decide in advance how it will handle results that disagree.

Purpose. The explanatory sequential design starts with a quantitative strand and follows it with a qualitative strand that helps explain the quantitative results. It suits questions in which numbers show a pattern, a surprise or a subgroup difference that the team needs to understand.

Procedure. The team collects and analyzes the quantitative data, identifies the results that need explanation, selects participants for the qualitative strand from the first sample, writes an interview or focus group guide that probes those results, and interprets how the qualitative findings explain the numbers.

Sampling. Because qualitative participants come from the quantitative sample, the survey or first data collection must ask for consent to be contacted again.

Demands on the team. The design takes longer than a convergent design, and the final interview guide depends on results that do not yet exist when the ethics application is written. Teams usually submit a draft guide with the application and then submit the final guide to the research ethics board (REB) as an amendment.

Purpose. The exploratory sequential design starts with a qualitative strand that explores a topic, uses the findings to build something, and then tests or extends it with a quantitative strand. It suits topics where no suitable instrument exists, where an intervention must be adapted to a community, or where a team needs to know how widely views heard in interviews are shared.

Procedure. The team collects and analyzes qualitative data, builds a product from the findings (survey items, an intervention component or hypotheses), and then collects quantitative data with that product.

Sampling. The quantitative sample is usually different from the qualitative sample and much larger.

Demands on the team. The design usually takes the longest of the three. When the product is a new instrument, its measurement properties must be tested before the results can be trusted, which HSCI 410 Lesson 6 teaches.

Many health studies embed these core designs in larger projects. A randomized trial may include a qualitative process evaluation, a program evaluation may move through several sequential phases, and a community-based participatory project may combine strands at each stage of a partnership. HSCI 826 Lesson 5 applies the three core designs to program evaluation.

2.4 Integration and Joint Displays

Fetters, Curry and Creswell (2013) describe integration at three levels. At the design level, the choice among the three core designs sets when and how the strands meet. At the methods level, the strands are linked through four procedures, described in the items below. At the interpretation and reporting level, the team presents the integrated results in a narrative that weaves the two kinds of findings together, through data transformation (for example, counting how many participants raised a theme), or in a joint display.

Connectingv

One strand links to the other through sampling. In an explanatory sequential design, the survey results determine who is invited to an interview. Connecting requires that the first strand record who agreed to be contacted again and the characteristics used to select them.

Buildingv

The results of one strand inform the data collection of the other. Interview themes can become survey items, and survey results can become prompts in an interview guide. Building is the main form of integration in exploratory sequential designs.

Mergingv

The two sets of results are brought together for comparison after each has been analyzed. Merging is the main form of integration in convergent designs, and a joint display is its usual tool.

Embeddingv

Data collection and analysis are linked at several points, often when a supporting strand sits inside a larger design. Interviews conducted at several stages of a trial to inform recruitment, delivery and interpretation are an example.

A joint display is a table or figure that places quantitative and qualitative results side by side, organized by topic, together with the conclusion drawn from reading them together. That conclusion is called a meta-inference, an overall inference that integrates the inferences from both strands (Teddlie & Tashakkori, 2009). Fetters and colleagues (2013) describe three ways the strands can fit: confirmation, when both lead to the same conclusion; expansion, when they agree in part and one adds something the other lacks; and discordance, when they disagree. Discordance is a result to be reported and examined, because it often points to a measurement problem, a subgroup difference or a new question. The table below is an illustrative joint display for the Cedar Valley study; Lesson 12 returns to the design of joint displays for a report.

TopicSurvey result (illustrative)Interview and focus group finding (illustrative)Meta-inference and fit
Living aloneOf 608 respondents living alone, 212 (34.9 percent) scored 6 or higher, compared with 180 of 992 (18.1 percent) living with others.Participants living alone described evenings, weekends and meals as the times when they felt most alone.Confirmation. Both strands link living alone with loneliness, and the interviews show when support would matter most.
Community sizeLoneliness was similar in Cedar City (116 of 480, 24.2 percent) and in the smaller communities and rural areas (276 of 1,120, 24.6 percent).Interview participants who had moved to a smaller town described feeling like outsiders for several years, while focus group members who had lived in small towns for decades described close networks of neighbours.Expansion. The similar averages may combine two groups with different experiences, so the team adds years at the current address, which the survey recorded, to its quantitative analysis.
Interest in programsOf the 392 lonely respondents, 220 (56.1 percent) said they would take part in a weekly social program.Several participants said they would avoid a program described as being for lonely people, because the label felt embarrassing.Discordance. Stated interest may overestimate attendance, and the way a program is described may affect who comes.

2.5 Choosing a Design for the Cedar Valley Study

Case: How the Cedar Valley team chose its design

Dr. Hart's team considered all three core designs. An exploratory sequential design would have made sense if the team had needed to develop its own loneliness measure, but the three-item UCLA Loneliness Scale is a validated instrument, and starting with interviews would have delayed the survey by most of a year. A convergent design would have been faster, because the interviews could have run during the survey period. The team rejected it because it wanted to choose interview participants using survey answers and to write interview prompts that probed the survey's patterns.

The team chose an explanatory sequential design and wrote it as QUAN → QUAL, with capital letters for both strands because its SPIDER question carries equal weight with its PECO question. The survey (1,600 completed responses) and the linked records come first. The qualitative strand follows: 24 semi-structured interviews with older adults living alone who moved to a smaller community in the past five years, and four focus groups (two with older adults, one with family caregivers, and one with clinic staff and community connectors).

The design integrates the strands at three points. The team connects them through two survey items, one asking for consent to be contacted again and one asking how long the respondent has lived at the current address, which together let the team select interview participants with a range of loneliness scores. It builds the interview guide from survey results by adding prompts about the patterns the survey shows. It merges the results at the interpretation stage in a joint display.

The choice has practical consequences that later sections take up. The ethics application must describe both phases, include a draft interview guide, and state that the final guide will be submitted as an amendment (Section 3). The survey consent form must include optional consent to be contacted again and to have survey answers linked to health records (Lesson 5). The timeline must place the interviews after the survey closes, which makes the qualitative strand part of the project's longest chain of tasks (Section 4).

Try it: choose a design and write its notation

For each scenario, name the core design, write its notation, and state the main point of integration. (1) A team must design a questionnaire on how newcomers to Canada experience access to primary care, and no suitable instrument exists. (2) A health authority survey found that uptake of a free shingles vaccine was much lower in two communities than in the rest of the region, and the team wants to understand why. (3) A team has six months to evaluate a new clinic intake process and wants both patient ratings and staff perspectives on the same aspects of the process.

Suggested answersv

(1) An exploratory sequential design, written QUAL → quan, in which interviews or focus groups with newcomers generate the questionnaire items. The main integration is building. (2) An explanatory sequential design, written QUAN → qual, in which interviews in the two communities explain the survey result. The main integration is connecting, because interview participants are chosen in the communities with low uptake. (3) A convergent design, written QUAN + QUAL, in which patient ratings and staff interviews are collected in the same period on the same aspects of intake. The main integration is merging in a joint display.

Reflection

A rural health region offers a free community transportation service that takes older adults to medical appointments, but few eligible residents use it. The program manager asks a research team to find out why use is low and how common each reason is. No existing survey measures reasons for not using community transportation. The team has funding for about 20 interviews and a survey of about 400 eligible residents, and it has eighteen months.

The three core mixed-methods designs are the convergent design (quantitative and qualitative data collected in the same phase, analyzed separately and merged for comparison), the explanatory sequential design (quantitative data first, followed by qualitative data that explain the results, integrated mainly by connecting the samples) and the exploratory sequential design (qualitative data first, used to build an instrument or intervention for a later quantitative strand, integrated mainly by building). In mixed-methods notation, capital letters mark a strand with priority, lower case marks a supporting strand, a plus sign means the strands run at the same time, and an arrow means one strand follows the other.

(a) Choose a design and justify it. (b) Write its notation. (c) Describe the main point of integration and what the team would do there. (d) Name one practical challenge of your design and how the team could manage it.

Model answer

(a) An exploratory sequential design fits best. The manager wants to know how common each reason for non-use is, but no instrument measures those reasons, so the team must first learn what the reasons are. Interviews with about 20 eligible residents, including people who have never used the service and people who stopped using it, would identify the reasons in residents' own words.

(b) QUAL → QUAN, because the interviews and the survey carry equal weight in answering the manager's question.

(c) The main integration is building. After coding the interviews, the team turns each theme, such as difficulty booking a ride by telephone, worry about long waits after appointments, or reluctance to be seen as needing help, into one or two survey items, pilots them with a few residents, and includes them in the survey of 400 residents. The survey then estimates how common each reason is.

(d) Time is the main challenge, because the survey cannot start until the interviews have been analyzed. The team could set up the survey platform and its recruitment plan while the interviews are under way and reserve a fixed window for writing and piloting items. A different strong answer might argue for a convergent design with open-ended survey questions if eighteen months proved too short, while accepting a less developed instrument.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In an explanatory sequential design, what is the main purpose of the qualitative strand?

In an explanatory sequential design, the quantitative strand comes first and the qualitative strand helps explain its results. Developing survey items from qualitative data describes the exploratory sequential design, and collecting both at once describes the convergent design.

Question 2: What does the notation QUAL → quan describe?

Capital letters mark the strand with priority and lower case marks a supporting strand, while the arrow shows that one strand follows the other. Equal priority in the same period would be written QUAN + QUAL.

Question 3: In a convergent study, survey results and interview themes disagree on one topic. What should the team do?

Discordance is a result in its own right. It may point to a measurement problem, a subgroup difference or a new question, so the team reports it and examines possible explanations. Dropping either strand would discard information the design was built to collect.

Question 4: Cedar Valley interview participants are selected from survey respondents who agreed to be contacted again. Which method of integration does this illustrate?

Connecting links one strand to the other through sampling, here by drawing interview participants from survey respondents. Merging compares results after analysis, and building uses one strand's results to write the other strand's questions.
Section 3 of 5

The Research Protocol and Its Sections

⏱ Estimated reading time: 30 minutes
Section 3 of 5

The Research Protocol

Its readers, its twelve sections, a one-page summary, and how it is versioned, amended and registered.

30 minutes
Protocol and proposal

Two documents with different jobs

Grant proposal

It argues that the study deserves funding.

Protocol

It tells the team, the REB and partners exactly what will happen.

Research team
Ethics board
Funders and partners
Data stewards
Advisory groups
Future readers
Source document

Write each decision once

Grant proposal
Ethics application
Data access request
Data management plan
Registration
Methods section

When each document draws on the protocol, the documents stay consistent with one another.

Twelve sections

What a protocol contains

Administrative information
Background and rationale
Questions and hypotheses
Design and framework
Setting and participants
Data and measures
Analysis plan
Ethics and engagement
Data management
Timeline
Dissemination
References and appendices

For trials, the SPIRIT 2025 statement (Chan et al., 2025) lists the required items.

Worked example

A one-page Cedar Valley summary

Design

The study uses an explanatory sequential design, with a cohort design for the emergency department question.

Data

The team has 1,600 survey responses, linked records, 300 charts, 24 interviews and 4 focus groups.

Timeline

The study runs 24 months, with REB approval by month 5 and the final report at month 24.

The test

Any row you cannot fill in marks a decision not yet made.

Versions and amendments

Keeping track of change

Drafts: 0.1, 0.2

Working drafts carry decimal numbers.

Submitted: 1.0, 2.0

Each version sent for approval takes a whole number.

Amendment log

Each approved change is recorded with its reason and the REB decision.

Under TCPS 2, changes to approved research need REB approval before they are made. Registration and publication let readers compare a report with its plan.

Carry forward

From plan to schedule

The protocol’s timeline and data management sections need tools of their own.

Section 4 builds the Cedar Valley Gantt chart and the roles, budget, files and data plan behind it.

Learning Objectives for this section

  • Explain what a research protocol is, how it differs from a grant proposal, and who reads it.
  • Identify the standard sections of a protocol and the content each section holds.
  • Explain how the protocol supplies content for the ethics application, the data access request, the data management plan and the methods section of a report.
  • Draft a one-page protocol summary for a study.
  • Describe how protocols are versioned, amended, registered and published.

3.1 What a Protocol Is and Who Reads It

A research protocol is the written plan for a study. It records the decisions the team has made about the question, the design, the participants, the data, the analysis and the safeguards, in enough detail that another qualified researcher could carry out the study as intended. Sections 1 and 2 produced two of those decisions for the Cedar Valley study: the design family for each question and the mixed-methods design that joins them. The protocol is where such decisions are written down, dated and kept together.

A grant proposal and a protocol share much of their content, and teams often adapt one from the other. They serve different purposes. A proposal argues that a study deserves funding, so it emphasizes the importance of the problem, the strength of the team and the value of the expected results. A protocol tells the people who will carry out, approve and oversee the study exactly what will happen, so it emphasizes procedures: who is eligible, how they are recruited, what is measured and when, how data are stored, and how they will be analyzed. A funded proposal usually needs to be expanded into a full protocol before data collection begins.

Several groups read a protocol, and each reads it for a different reason. Select each card to see what one group looks for.

The research teamClick to explore
The research ethics boardClick to explore
Funders and partnersClick to explore
Data stewardsClick to explore
Communities and advisory groupsClick to explore
Future readersClick to explore

The protocol also serves as the source document for the other papers a project produces. Each of the documents in the figure draws on the same decisions about design, participants, data and analysis. When those decisions are written once in the protocol and copied from there, the documents stay consistent with one another. When each document is written from memory, small differences appear, and an REB or a data steward may ask the team to explain them.

Research protocol one plan, version-controlled Grant proposal argues for funding Ethics application REB review (Lesson 5) Data access request linked data (Lesson 9) Data management plan files and data (Section 4) Registration public record (3.4) Methods section final report (Lesson 12)
The protocol is the source document for the other documents a project produces, so a change to a decision is made once in the protocol and then carried into each document that depends on it.

3.2 The Sections of a Protocol

No single format is required for every protocol. Funders, REBs and journals each publish templates, and the template you are given should take precedence. For randomized trials, the SPIRIT 2025 statement (Chan et al., 2025), which updated the 2013 statement, lists the items a trial protocol should contain. For systematic reviews, PRISMA-P (Moher et al., 2015) serves the same function, and HSCI 241 Lesson 2 teaches it. Protocols for observational, qualitative and mixed-methods studies usually contain the twelve sections described below. Open each item to see what the section holds, an example from the Cedar Valley protocol, and the lesson of this course that prepares you to write it.

1. Administrative informationv

This section gives the title, the version number and date, the names and roles of the team, the funder, the partner organizations and any registration number. Every page of the protocol carries the version number and date in its footer. Cedar Valley: the title page lists Dr. Maya Hart as principal investigator, the graduate research assistant, the community research associate, the advisory group of six older adults, the Cedar Valley Health Authority and the Cedar Valley First Nations Health Centre.

2. Background and rationalev

This section summarizes what is known, identifies the gap, and explains why the study is needed now. It is usually two to four pages with references. Cedar Valley: the section summarizes research on loneliness and health service use and notes that little is known about rural older adults in British Columbia. Lesson 12 teaches how to build this argument.

3. Questions, aims and hypothesesv

This section states the research questions in structured form, the aims and objectives, and any hypotheses. Cedar Valley: the PECO and SPIDER questions from Lesson 2, with the hypothesis that respondents who score 6 or higher on the UCLA scale have more emergency department visits in the following twelve months.

4. Design and conceptual frameworkv

This section names the design family for each question, the mixed-methods design and its notation, and the causal diagram that guided variable selection. Cedar Valley: an explanatory sequential design (QUAN → QUAL) with a cohort design for the PECO question, together with the DAG from Lesson 3.

What the design and conceptual framework section contains

Card 4 deserves a closer look, because it is where a protocol explains how theory shaped the study. The section usually has three linked parts. The first is the guiding theory or perspective, a general account of why the outcome occurs that the team draws on to choose its questions and variables. The second is a conceptual framework, which names the constructs the study will measure or explore and states how the team expects them to relate to one another. The third is the causal diagram derived from the framework, which turns the expected relations into explicit assumptions about which variable affects which, and therefore which variables must be measured.

In the Cedar Valley protocol, the guiding perspective is the web of causation together with Krieger’s (1994) critique of it, which led the team to treat income, transport and rural residence as causes in their own right (Lesson 3). The conceptual framework is the causal web, which names loneliness, its individual, relationship and structural causes, and the pathways through depression and physical inactivity to emergency department visits. The causal diagram is the DAG from Lesson 3, with its justification table. For further reading, HSCI 230 Lesson 7 Section 1 (Theory Before Instruments) shows how frameworks shape what a study measures, and HSCI 841 Lesson 2 Section 2.6 describes how theory enters a qualitative study through sensitizing concepts and conceptual frameworks.

5. Setting and participantsv

This section describes the setting, the target and study populations, eligibility criteria, the sampling approach and recruitment procedures. Cedar Valley: adults aged 65 and older living in the health authority's region, with separate criteria for interview participants. Lesson 7 teaches these decisions.

6. Data collection and measuresv

This section lists every instrument and data source, the variables each provides, and when each is collected. Cedar Valley: the REDCap survey with the three-item UCLA Loneliness Scale, linked administrative records obtained through Population Data BC, a chart abstraction form for 300 records at six clinics, and interview and focus group guides. Lessons 7 to 10 prepare these materials.

7. Analysis planv

This section states, before the data are collected, how each question will be analyzed: the descriptive statistics, the main association and the variables to be adjusted for, the qualitative approach, and how the strands will be integrated. Writing the plan in advance protects the study from choosing analyses after seeing which ones give striking results. Lesson 11 introduces the first analytic steps.

8. Ethics and engagementv

This section summarizes risks and how they are reduced, the consent process, privacy protections, the engagement plan and any research agreement with an Indigenous partner. It points to the full ethics application rather than repeating it. Lessons 4 and 5 prepared this content.

9. Data managementv

This section summarizes how data will be stored, protected, documented, shared and retained, and refers to the full data management plan, which Section 4 of this lesson teaches.

10. Timeline and milestonesv

This section gives the schedule of the main tasks and the dates by which key milestones should be reached, usually as a Gantt chart. Section 4 shows how to build one.

11. Dissemination and knowledge sharingv

This section states how results will reach each audience, including participants, partners, decision-makers and researchers. Cedar Valley: community events in four communities, a report to the health authority, a plain-language summary reviewed by the advisory group, and journal articles.

12. References and appendicesv

The appendices hold the instruments, consent forms, interview guides, recruitment materials and data dictionary. Each appendix carries its own version number, because it can change separately from the main text.

3.3 Worked Example: A One-Page Protocol Summary

A full protocol for a study like Cedar Valley may run to twenty or thirty pages with appendices. Most teams also keep a one-page summary for partners, new team members and committee meetings. The summary below follows the twelve sections in compressed form. Writing a summary of this kind early is a useful test: any row that you cannot fill in identifies a decision that has not yet been made.

ElementCedar Valley Social Connection Study (fictional), protocol version 2.0
TitleLoneliness, social isolation and health service use among adults aged 65 and older in Cedar Valley: a mixed-methods study
Team and partnersDr. Maya Hart (principal investigator), a graduate research assistant, a community research associate and an advisory group of six older adults, with the Cedar Valley Health Authority and the Cedar Valley First Nations Health Centre
QuestionsHow common is loneliness among adults aged 65 and older? Is loneliness associated with emergency department visits in the following twelve months? How do older adults living alone experience social connection after a move to a smaller town?
DesignExplanatory sequential mixed-methods design (QUAN → QUAL); cohort design for the association between loneliness and emergency department visits
ParticipantsAdults aged 65 and older in the health authority's region; for interviews, adults living alone who moved to a smaller community in the past five years and agreed to be contacted again
DataRegional survey (1,600 completed responses) with the three-item UCLA Loneliness Scale; linked records of physician visits, hospital discharges and emergency department visits (with consent); 300 chart reviews at six clinics; 24 interviews; four focus groups
AnalysisDescriptive statistics and a Table 1; comparison of emergency department visits by loneliness group, adjusted for the confounders identified in the DAG; first-cycle coding and qualitative description; integration in a joint display
Ethics and dataREB approval; separate optional consents for linkage and for contact about an interview; research agreement with the First Nations health partner; data management plan with identifiers held separately from responses
TimelineTwenty-four months, with REB approval by month 5, survey close at month 10 and final report at month 24
DisseminationCommunity events, a report to the health authority, a plain-language summary and journal articles

3.4 Versions, Amendments, Registration and Publication

Version control for protocols

A protocol changes as a study develops, and every reader needs to know which version is current. A simple convention works for most projects. Working drafts carry decimal numbers (0.1, 0.2 and so on). The first version submitted to the REB becomes version 1.0. Drafts prepared after that carry numbers such as 1.1 and 1.2, and the next version submitted for approval becomes 2.0. Each version is saved as a separate file, and the protocol keeps an amendment log that records every approved change.

A protocol amendment is a change to research that the REB has already approved. Under TCPS 2, researchers submit proposed changes to the REB and wait for approval before making them, with an exception for changes needed to remove an immediate risk to participants. Adding a survey question, changing recruitment materials, adding a site and changing how data are stored all require an amendment. The amendment log for the Cedar Valley protocol shows how changes are recorded.

VersionSubmittedChangeReasonREB decision
1.0Month 3First version submitted with the ethics application, including a draft interview guideInitial reviewApproved in month 5
2.0Month 9Final interview guide added, with prompts based on survey resultsThe explanatory sequential design builds the guide from the surveyApproved in month 10
3.0Month 12Telephone interviews added as an optionThe advisory group noted that some rural participants lack reliable internet accessApproved in month 13

Registration and publication

Registering a protocol creates a public, dated record of what a study planned to do. Clinical trials are expected to be registered in a public registry, such as ClinicalTrials.gov or a registry in the World Health Organization's network, and journals that follow the recommendations of the International Committee of Medical Journal Editors consider a trial for publication only if it was registered at or before the enrolment of its first participant. Registration is optional for most observational, qualitative and mixed-methods studies, but many teams preregister them, for example on the Open Science Framework, so that readers can compare the final report with the plan. Some journals, including BMJ Open and JMIR Research Protocols, publish full protocols.

Both practices respond to a documented problem. Chan and colleagues (2004) compared published trial reports with their protocols and found that outcomes were often reported selectively, with outcomes that showed clear differences more likely to appear in the published report. A public protocol lets readers see whether all planned outcomes were reported and whether any change to the plan was explained.

Students who have taken HSCI 241 met registration for reviews, in PROSPERO and on the Open Science Framework, in its Lesson 1 Section 4.3, and HSCI 230 Lesson 1 (Research Integrity and Reform) explains the problems that registration addresses. Neither course is required before this one, so this section stands on its own.

Try it: draft a protocol table of contents

Choose a research question, either one of the Cedar Valley questions or another question that interests you, and write the twelve section headings of a protocol for it. Under each heading, write one sentence stating a decision already made or a decision still needed. Mark the rows that cannot yet be completed, and note which lesson of this course covers each one.

Reflection

A student team plans a study of sleep among undergraduate students who work part-time. It will run an online cross-sectional survey of about 300 students on sleep duration and weekly hours of paid work, followed by 10 interviews with students who report working more than 20 hours a week.

A protocol for such a study usually has twelve sections: administrative information; background and rationale; questions, aims and hypotheses; design and conceptual framework; setting and participants; data collection and measures; analysis plan; ethics and engagement; data management; timeline and milestones; dissemination; and references and appendices. A common versioning convention numbers working drafts 0.1, 0.2 and so on, makes the first version submitted to the research ethics board (REB) version 1.0, numbers later drafts 1.1, 1.2 and so on, and makes the next version submitted for approval version 2.0. Under TCPS 2, changes to approved research must be approved by the REB before they are made, except to remove an immediate risk to participants.

(a) Write one or two sentences of content for four sections: design and conceptual framework; setting and participants; analysis plan; and data management. (b) After the REB approves version 1.0, the team wants to add a question about caffeine use to the survey. Describe what the team must do, what the new protocol version number will be, and what it should record in the amendment log.

Model answer

(a) Design: The study uses an explanatory sequential design (QUAN → qual), in which a cross-sectional survey estimates the association between weekly hours of paid work and sleep duration, and interviews explain how work schedules affect sleep. Setting and participants: About 300 undergraduates at one university who work for pay, recruited through course announcements, and 10 interview participants chosen from respondents who work more than 20 hours a week and agreed to be contacted again. Analysis plan: The team will report descriptive statistics, compare mean sleep duration across groups of weekly work hours with adjustment for year of study and living situation, code the interviews with a first-cycle codebook, and compare the strands in a joint display. Data management: Survey data will be kept on university-approved storage, contact details will be stored apart from responses with a linking key, and audio will be deleted once transcripts are checked.

(b) The team must submit an amendment describing the new question and its purpose, and it must not use the revised survey until the REB approves it. The amended protocol becomes version 2.0. The amendment log records the version number, the date submitted, the change (one survey item on caffeine use), the reason (caffeine may confound the association between work hours and sleep) and the date of the REB decision.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: How does a research protocol differ from a grant proposal?

A proposal persuades a funder that a study deserves money, while a protocol is the detailed operating plan that the team, the REB and partners follow. Both are written before data collection, and protocols are used for all kinds of studies.

Question 2: After the REB approves the Cedar Valley protocol, the team wants to add a survey question about pet ownership. What should it do?

Under TCPS 2, changes to approved research are submitted to the REB and approved before they are made, unless they are needed to remove an immediate risk to participants. A new application is unnecessary, because an amendment covers changes to an approved study.

Question 3: Which protocol section states, before data are collected, how the main association will be estimated and which variables will be adjusted for?

The analysis plan specifies the analyses in advance, which protects the study from choosing analyses after seeing which ones give striking results. The background explains why the study is needed, and dissemination describes how results will be shared.

Question 4: Which reporting guideline lists the items that a protocol for a randomized trial should contain?

The SPIRIT 2025 statement (Chan et al., 2025) lists the items a trial protocol should contain. STROBE, COREQ and PRISMA guide the reporting of completed observational studies, qualitative studies and systematic reviews.
Section 4 of 5

Managing a Research Project: Timelines, Roles, Budgets, Files and Data

⏱ Estimated reading time: 40 minutes
Section 4 of 5

Managing a Research Project

Timelines, roles, budgets, files, version control and the data management plan.

40 minutes
Timelines

From deliverables to a schedule

  • List the deliverables and break each into tasks.
  • Estimate durations and add time, because the planning fallacy makes estimates too short.
  • Set the dependencies, most of which are finish-to-start.
  • Mark the milestones and find the critical path.
Cedar Valley Gantt chart

The critical path runs through both strands

Protocol Ethics Survey Interviews Coding Integrate Report Data access and linkage (slack)

REB approval is due by month 5, the survey closes at month 10, and the final report is due at month 24.

Roles

One accountable person per task

Responsible

This person does the work.

Accountable

This person answers for completion.

Consulted

These people give input first.

Informed

These people are kept up to date.

The team charter records how authorship will be decided, using the ICMJE criteria.

Budget

Every line shows its calculation

Focus group transcription line
\[ 360 \text{ minutes} \times \$2.00 = \$720 \]
$153,480illustrative total, before data access costs

In-kind contributions

The university hosts the survey platform, and clinics provide meeting rooms.

Files and versions

Find any file, recover any version

2026-05-04_CVSCS_interview-guide_v03.docx

  • Folders are numbered in the order of the work.
  • Raw data stay read-only, and scripts write the cleaned data.
  • Identifiers, consent forms and audio live in a restricted folder.
  • Git tracks analysis code, and data never go into a code repository.
Data management plan

Seven headings, two sets of principles

Data collection
Documentation and metadata
Storage and backup
Security and privacy
Sharing and reuse
Preservation and retention
Responsibilities and resources

FAIR

Data should be findable, accessible, interoperable and reusable (Wilkinson et al., 2016).

CARE

Indigenous data governance rests on collective benefit, authority to control, responsibility and ethics (Carroll et al., 2020).

Putting it together

The Cedar Valley project plan

  • Design: explanatory sequential, with a cohort design for the PECO question.
  • Timeline: 24 months, with the critical path through both strands.
  • Roles and budget: one accountable person per task, and a calculation for every line.
  • Files and data: numbered folders, a naming convention and a data management plan.

Then complete the reflection, the knowledge check and the final assessment.

Learning Objectives for this section

  • Build a project timeline by breaking the work into tasks, estimating durations, setting dependencies and milestones, and identifying the critical path.
  • Assign roles with a RACI matrix and record team agreements, including authorship.
  • Prepare a line-item budget with a written justification for each line.
  • Set up a folder structure, a file naming convention and a version control routine for a project.
  • Draft a one-page data management plan that applies the FAIR and CARE principles.

4.1 Timelines and Gantt Charts

Project management is the set of practices that keeps a study on schedule, within budget and in control of its files and data. This section applies five of them to the Cedar Valley study: timelines, roles, budgets, file organization and data management.

A timeline begins with a work breakdown structure, a list of everything the project must produce, broken into tasks small enough that someone can estimate how long each will take. The team then works through five steps.

  1. List the deliverables, such as the approved ethics application, the cleaned survey file and the final report, and break each into tasks.
  2. Estimate how long each task will take, and add time to the estimate. People routinely underestimate how long their own tasks will take, even when similar tasks have run late before, a pattern known as the planning fallacy (Buehler et al., 1994).
  3. Identify the dependencies between tasks. The most common is a finish-to-start dependency, in which one task cannot begin until another ends: the survey pilot cannot start until the REB has approved the study.
  4. Place the tasks on a calendar and mark the milestones, the dated points at which a key result is reached, such as REB approval or the close of the survey.
  5. Find the critical path and protect it with buffer time.

A Gantt chart displays the result as horizontal bars on a calendar, one bar per task. The format is named after Henry Gantt, who developed bar charts for scheduling work in the 1910s. The critical path is the longest chain of dependent tasks from the start of the project to its end. Its length sets the shortest time in which the project can be finished, so any delay to a task on the critical path delays the whole project. Tasks off the critical path have slack, an amount of time by which they can slip without delaying the finish.

Year 1 Year 2 Task 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 Protocol and advisory set-up Ethics application and review REB approval Data access request Access approved Survey build and pilot Survey fieldwork Survey closes Chart review (six clinics) Linkage and data release Linked data received Quantitative analysis Interviews and focus groups Transcription and coding Integration (joint display) Reporting and sharing Final report Advisory group meetings Critical path Task with slack Milestone Advisory meeting
The Cedar Valley timeline over twenty-four months; red bars form the critical path from the protocol through the survey, the qualitative strand and integration to the final report, and teal bars have slack.

The chart shows the consequences of the design chosen in Section 2. Because interview participants are selected from survey respondents, the qualitative strand cannot begin until the survey closes in month 10, so both strands sit on the critical path. The data access request and the linkage run alongside them with some slack: a delay of two months could be absorbed, and a longer one would push back integration and the final report.

Three tasks depend on decisions made by others: the REB review, the data access request and the agreements with the six partner clinics. The team starts them early, builds buffer time after them, and plans work it can do while it waits, such as building the survey during the ethics review. It reviews the chart at each monthly meeting.

4.2 Roles and Responsibilities

Many problems in research teams come from unclear responsibility. A RACI matrix assigns one of four roles to each person for each task. The person responsible (R) does the work. The person accountable (A) answers for the task being completed and makes the final decision, and each task has exactly one. People who are consulted (C) give input before decisions are made, and people who are informed (I) are kept up to date.

TaskDr. Hart (lead)Graduate research assistantCommunity research associateAdvisory groupHealth authorityFirst Nations health centre
Ethics applicationARCCIC
Survey build and pilotARRCIC
Data access requestARIICC
Chart review at six clinicsARIICI
Interviews and focus groupsARRCIC
Coding and integrationARRCIC
Final report and knowledge sharingA/RRRCCC

The matrix sits inside a short team charter that records how often the team meets, how decisions are made, how members communicate, who may access which data, and how authorship will be decided. Under the criteria of the International Committee of Medical Journal Editors, an author has made a substantial contribution to the design of the work or to the collection, analysis or interpretation of its data; has drafted the work or revised it critically; has approved the final version; and has agreed to be accountable for the work. Advisory group members and community partners who meet these criteria are authors, and those who do not are named in the acknowledgements. A research agreement with a First Nations partner may also set out how the partner reviews publications before submission (Lesson 4). Agreeing on these matters at the start is usually easier than settling them when a manuscript is ready.

4.3 Budgets

A research budget lists the expected costs of the project, line by line, and a budget justification explains how each amount was calculated and why the expense is needed. Reviewers use it to check that the budget matches the protocol: 24 interviews in the protocol should mean 24 honoraria in the budget and enough staff time to transcribe 24 recordings. Each line shows its calculation as a quantity multiplied by a unit cost. The budget below is illustrative; real rates come from salary scales, vendor quotes and the funder's rules.

LineCalculationAmount
Personnel
Graduate research assistant15 hours a week × 48 weeks × 2 years at $30 an hour$43,200
Community research associate10 hours a week × 48 weeks × 2 years at $35 an hour$33,600
Chart abstraction300 charts × 45 minutes at $30 an hour$6,750
Participants and partners
Advisory group honoraria6 members × 12 meetings × $50$3,600
Advisory group travel12 meetings × $100 for mileage and parking$1,200
Interview honoraria24 participants × $50$1,200
Focus group honoraria24 participants in the three community groups × $40$960
Survey thank-you cards1,600 expected completed surveys × $10 grocery card$16,000
First Nations health centre partnership fundsAmount set in the research agreement$6,000
Data collection and analysis
Survey mailings5,250 people × 3 mailings (pre-notice letter, invitation and reminder postcard) × $1.00$15,750
Replacement paper questionnaires3,600 packages with return postage × $4.00$14,400
Focus group transcription360 audio minutes (4 focus groups of 90 minutes) × $2.00$720
Focus group interpreters2 sessions × $400$800
Travel to smaller communities30 trips × 150 km × $0.60 per km$2,700
Qualitative analysis software2 licences × $600$1,200
Data access and linkageQuote requested from the data providerTo be added
Knowledge sharing
Community presentations4 presentations × $600 for venue, refreshments and large-print summaries$2,400
Open-access publication fee1 article$3,000
TotalExcluding data access and linkage, which will be added when the quote arrives$153,480

A justification turns each line into a short argument. For transcription, the Cedar Valley justification reads: "Four focus groups of about 90 minutes will produce about 360 minutes of audio. Overlapping voices defeat most automated tools, so professional transcription at $2.00 per audio minute, costing $720, is needed for the groups. The interviewers will transcribe the 24 interviews from drafts produced by speech recognition software on an encrypted university laptop, so that work is costed under personnel. Clean verbatim transcripts are needed for the first-cycle coding described in the analysis plan." The justification also notes in-kind contributions, resources that partners provide without charge: the university hosts the REDCap survey platform, the partner clinics provide meeting rooms, and a health authority analyst advises on the administrative data.

Three checks prevent common problems. Read the funder's rules on eligible expenses (for the federal granting agencies, the Tri-Agency Guide on Financial Administration). Ask your research office whether salaries must include a percentage for employee benefits. Confirm that honoraria match the amounts approved by the REB, which checks that payment does not pressure people to take part (Lesson 5).

4.4 File Organization and Naming

A two-year project with several staff produces thousands of files. A folder structure agreed at the start lets anyone find a file without asking, and it keeps identifiable information away from everyday working files. The Cedar Valley structure numbers its folders in the order the work happens, keeps raw data read-only so the original export can always be recovered, and holds the linking key, consent forms and audio in a separate restricted location. Lesson 5 explained why identifiers are stored apart from responses.

CVSCS_project/ shared folder on approved institutional storage 00_admin approvals, agreements, budget and team charter 01_protocol protocol versions and the amendment log 02_instruments survey, abstraction form and interview guides 03_data_raw original exports, read-only and never edited 04_data_clean cleaned files, written only by scripts 05_scripts analysis scripts, numbered in the order they run 06_qualitative de-identified transcripts, codebook and memos 07_outputs tables, figures and reports README.txt what is where, naming rules and contacts CVSCS_restricted/ a separate location with restricted access Linking key (held by one person), signed consent forms and interview audio.
The Cedar Valley folder structure numbers folders in the order of the work and keeps identifiable material in a separate restricted location.

A file naming convention is a fixed pattern for file names that every team member follows. The Cedar Valley convention is date, project, content and version, separated by the underline character (_), as in 2026-05-04_CVSCS_interview-guide_v03.docx. A name such as Interview guide FINAL revised (2).docx gives no date, sorts unpredictably and leaves readers unsure which file is current. HSCI 410 Lesson 1 numbers analysis datasets with a two-digit suffix, such as bp01.csv, and records each version in a file log, which is compatible with this convention. Select each card for one of the rules behind the convention.

Dates first, in ISO formatClick to explore
No spaces or special charactersClick to explore
The same order every timeClick to explore
Versions with leading zerosClick to explore
No identifying informationClick to explore
A README fileClick to explore

4.5 Version Control

Version control is any system that records the changes made to a file over time, so that the team can see what changed, who changed it and why, and can return to an earlier version. Most projects use three systems, one for each kind of file.

Protocols, instruments and consent forms use named versions. Save a new file with a higher version number each time a version is shared or submitted, keep earlier files in an archive subfolder, and record the changes in a change log or, for the protocol, the amendment log.

Institutional cloud storage usually keeps an automatic version history, which lets a team restore a file that was overwritten by mistake. Automatic versions carry no description of what changed or why, so they supplement named versions at milestones without replacing them.

Analysis scripts are best managed with Git, a version control program that records each saved change as a commit with a short message, such as "Recode living arrangement into two groups". Services such as GitHub host the record so that a team can share it. Code repositories hold scripts and documentation only, and never data, because a file committed to a shared repository is hard to remove completely. Lesson 11 returns to scripts as the record of every change made to a dataset.

4.6 The Data Management Plan

A data management plan (DMP) is a living document that describes how a project's data will be collected, documented, stored, protected, shared and preserved, during the project and after it ends. Lesson 5 introduced the DMP with a focus on privacy; here it becomes an operating document. The Tri-Agency Research Data Management Policy (2021) asks Canadian institutions to develop data management strategies and requires DMPs for some funding opportunities. The DMP Assistant, a free bilingual online tool supported by the Digital Research Alliance of Canada, provides templates.

Two sets of principles guide the plan. The FAIR principles ask that data be findable, accessible, interoperable and reusable, so that they can be found and used again with appropriate permissions (Wilkinson et al., 2016). The CARE principles for Indigenous data governance (collective benefit, authority to control, responsibility and ethics) add attention to the people and purposes behind the data (Carroll et al., 2020). In Canada, the CARE principles sit alongside the First Nations principles of OCAP®, which Lesson 4 turned into agreements. Open each heading below to see what the section of a DMP covers and how the Cedar Valley one-page plan completes it.

1. Data collectionv

The plan lists each type of data, its format and its approximate volume. Cedar Valley: survey responses and chart abstraction forms in REDCap, exported as CSV files; linked administrative data held by the data provider; audio files; transcripts; and field notes.

2. Documentation and metadatav

The plan names the documents that let someone else understand the data. Cedar Valley: a data dictionary for each dataset, the qualitative codebook, the README file, and the protocol version under which each dataset was collected.

3. Storage and backupv

The plan states where data are stored and how they are backed up. A common rule of thumb is the 3-2-1 rule: keep three copies of important data, on two different types of storage, with one copy in a different location. Cedar Valley: working files on university-approved storage with automatic backup, and no project data on personal laptops or USB drives.

4. Security and privacyv

The plan describes who can access the data and how identifiable information is protected. Cedar Valley: identifiers are held apart from responses and joined only through a linking key in the restricted folder; transcripts are de-identified before coding; audio files are deleted once transcripts are checked, on the schedule approved by the REB; and the team plans to analyze the linked data inside a secure research environment and remove only aggregate results (Lesson 9).

5. Sharing and reusev

The plan states which data can be shared, with whom and under what conditions. Cedar Valley: a de-identified survey file may be deposited in a repository such as Borealis if participants consented to sharing; the linked data cannot be shared by the team; transcripts are not shared publicly; and data from the First Nations health partner's community are shared only as the research agreement allows.

6. Preservation and retentionv

The plan states how long data are kept and how they are destroyed. Cedar Valley: retention periods follow the REB approval and university policy, and the destruction of each restricted file is recorded in the administrative folder.

7. Responsibilities and resourcesv

The plan names who carries out each part and what it costs. Cedar Valley: Dr. Hart is accountable for the plan, the graduate research assistant maintains the data dictionaries and backups, and the transcription and software costs appear in the budget.

The team reviews the plan at each milestone and whenever the protocol changes. When telephone interviews were added in protocol version 3.0, it added a rule that they are recorded only on university-approved devices.

Reflection

A student-led study must be completed in 12 months. Its tasks, durations and dependencies are: A, write the protocol (2 months, no predecessor); B, ethics review (3 months, after A); C, build the survey (1 month, after A); D, pilot the survey (1 month, after both B and C); E, survey fieldwork (2 months, after D); F, interviews (3 months, after E); G, survey analysis (2 months, after E); H, write the report (2 months, after both F and G). The critical path is the longest chain of dependent tasks from start to finish, and tasks off the critical path have slack, the time by which they can slip without delaying the finish.

The team's file naming convention is date (YYYY-MM-DD), project code, content and a two-digit version number, separated by the underline character, for example 2027-01-15_SLEEP_protocol_v02.docx.

(a) Identify the critical path and the total duration, and state how much slack tasks C and G have. (b) The project must finish in 12 months. Propose one realistic change that would shorten the critical path and explain its trade-off. (c) Write the file name for the third version of the interview guide, saved on 4 March 2027. (d) Name two safeguards the team's data management plan should include for interview audio files.

Model answer

(a) The critical path is A, B, D, E, F and H, which takes 2 + 3 + 1 + 2 + 3 + 2 = 13 months, one month longer than allowed. Task C can start at the end of month 2 and need not finish until the pilot begins at the end of month 5, so it has 2 months of slack. Task G runs from month 8 to month 10, while H cannot start until F ends at month 11, so G has 1 month of slack.

(b) The team could shorten the interviews from 3 months to 2 by training a second interviewer and inviting interview participants as soon as the survey closes. The project would then take 12 months. The trade-off is the cost of training and the need for both interviewers to follow the guide in the same way. Another option is to draft the report's background and methods during the interviews, so that the report needs only one month after F ends, at the risk of revising the methods if procedures change.

(c) 2027-03-04_SLEEP_interview-guide_v03.docx

(d) Audio files should be stored only in an encrypted, access-restricted location approved by the institution and kept apart from the transcripts. They should be deleted on the schedule approved by the REB once the transcripts have been checked, and the transcripts should be de-identified, with participant codes in file names.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: What is the critical path of a project?

The critical path is the longest chain of dependent tasks, and its length sets the shortest time in which the project can finish. A delay to any task on it delays the whole project, while tasks off it have slack.

Question 2: In a RACI matrix, what does the letter A stand for?

The accountable person answers for the task being completed and makes the final decision, and each task has exactly one. The other letters stand for responsible (does the work), consulted (gives input) and informed (kept up to date).

Question 3: Which file name follows the naming convention recommended in this lesson?

The recommended name puts an ISO date first, then the project, content and a two-digit version, joined without spaces. The other names contain spaces or punctuation, an ambiguous date, a person's name, or the word final.

Question 4: Which practice belongs in a data management plan for interview audio files that contain participants' voices?

Audio files identify participants, so they belong in restricted, institution-approved storage apart from working files and are deleted on the schedule the REB approved. Personal laptops, ordinary email and shared folders expose them to people who do not need access.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson has moved from a research question to a study that a team can carry out. The decision map in Section 1 asks five questions in order: whether existing studies already answer the question, whether it asks about numbers or meanings, whether the exposure is assigned on purpose and at random, whether groups are compared, and what the unit and timing of measurement are. Ethical and practical constraints then narrow the choice. For the fictional Cedar Valley Social Connection Study, the map placed the prevalence question in the descriptive family, the PECO question in a cohort design built on linked records, and the SPIDER question in the qualitative family.

Because the Cedar Valley questions fall in different families, the study needs a mixed-methods design. Section 2 compared the convergent, explanatory sequential and exploratory sequential designs and showed that integration, through connecting, building, merging and embedding, is what makes a study mixed methods. The team chose an explanatory sequential design (QUAN → QUAL), and that choice shaped everything that followed: the consent to be contacted again, the amendment for the final interview guide, and the critical path of the timeline.

Sections 3 and 4 turned the design into a plan. The protocol records the decisions in one version-controlled document that supplies the ethics application, the data access request, the data management plan and the methods section. The Gantt chart, the RACI matrix, the justified budget, the folder structure and naming convention, version control and the data management plan keep the project on schedule, within budget and in control of its data.

Key Takeaways from this lesson

  • The research question determines the design, and the decision map reaches a design family by asking about existing evidence, the kind of answer needed, assignment of the exposure, comparison and timing.
  • Ethics, the rarity of the outcome or exposure, time, money and existing data narrow the designs that a team can carry out, and community agreements may shape them further.
  • This course maps the design families. HSCI 230 Lessons 3 to 6 teach the observational designs in depth, and HSCI 230 Lesson 5 also teaches randomized trials.
  • A mixed-methods study integrates its quantitative and qualitative strands, and collecting both kinds of data without relating them does not produce a mixed-methods analysis.
  • The convergent design merges strands collected together, the explanatory sequential design uses qualitative data to explain quantitative results, and the exploratory sequential design uses qualitative findings to build a quantitative strand.
  • A joint display places both strands' results side by side with a meta-inference, and discordance between strands is a finding to report and examine.
  • The protocol records a study's decisions in one version-controlled document, from which the ethics application, data access requests, the data management plan and the methods section are drawn.
  • Changes to approved research require an REB-approved amendment before they are made, and the amendment log records each change and its reason.
  • A Gantt chart built from a work breakdown structure shows dependencies, milestones and the critical path, and buffer time belongs after steps that others control.
  • Clear roles, a justified budget, a consistent folder structure and naming convention, version control and a data management plan guided by the FAIR and CARE principles keep a project organized and its data protected.

Core Concepts Reviewed

Section 1: study design, design family, the five-step decision map, the descriptive, analytic observational, experimental, qualitative and evidence synthesis families, and the constraints that narrow the choice of design.

Section 2: mixed-methods research, strands, timing and priority, mixed-methods notation, the convergent, explanatory sequential and exploratory sequential designs, connecting, building, merging and embedding, joint displays and meta-inferences.

Section 3: the research protocol, its twelve sections, the protocol as a source document, the one-page summary, version numbering, protocol amendments and the amendment log, and protocol registration and publication.

Section 4: work breakdown structures, Gantt charts, dependencies, milestones, the critical path, RACI matrices and team charters, budget justifications, folder structures, file naming conventions, version control and data management plans.

The final reflection asks you to apply the whole lesson to a new study, from the choice of design to the data management plan.

Reflection

A community health centre in a mid-sized British Columbia city asks a research team to study food insecurity among post-secondary students. It has three questions: (1) How common is food insecurity among students at the city's college? (2) Is food insecurity associated with lower self-rated mental health? (3) How do students who experience food insecurity manage their food and their studies from week to week? The team has 18 months and $40,000, a faculty lead, one research assistant and a student advisory group of five. A validated food insecurity questionnaire is available.

Useful definitions: the convergent design collects quantitative and qualitative data in the same phase and merges them; the explanatory sequential design collects quantitative data first and then qualitative data to explain the results; the exploratory sequential design collects qualitative data first and uses them to build an instrument or intervention for a quantitative strand. The critical path is the longest chain of dependent tasks in a project. A budget justification shows each line as a quantity multiplied by a unit cost.

Write a short plan that (a) names the design family for each question, (b) names a mixed-methods design with its notation and explains where the strands are integrated, (c) lists four milestones in order and identifies which tasks you expect to be on the critical path, (d) gives two budget lines with their calculations, and (e) states two decisions for the data management plan.

Model answer

(a) Question 1 is descriptive and fits a cross-sectional survey using the validated questionnaire. Question 2 is explanatory; because food insecurity cannot be assigned, it is analytic observational, and with a single survey it is a cross-sectional analysis that cannot show whether food insecurity came before poorer mental health. Question 3 is exploratory and fits qualitative interviews.

(b) An explanatory sequential design, QUAN → QUAL, fits. A validated instrument exists, so no exploratory phase is needed, and the survey can identify food-insecure students who agree to be contacted for an interview, which connects the strands. Survey results also shape the interview prompts, and a joint display merges the findings at the end.

(c) The milestones are REB approval (month 3), survey close (month 7), interviews complete (month 11) and the final report (month 18). Ethics review, survey fieldwork, the interviews and coding lie on the critical path, because the interviews depend on the survey.

(d) Interview honoraria: 15 participants × $40 = $600. Research assistant: 10 hours a week × 60 weeks at $28 an hour = $16,800.

(e) Contact details will be stored apart from survey responses, joined only by a linking key in a restricted folder. Interview audio will be deleted once transcripts are checked, and transcripts will be de-identified before coding.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Choosing an Approach and Managing a Research Project (15 Questions)

Question 1: A question asks whether adults who live alone have more emergency department visits in the following year than adults who live with others, using survey data linked to later health records. Which family and design fit?

The exposure (living alone) is not assigned and groups are compared, so the study is analytic observational. Because living arrangement is measured before the visits that the linked records count, the design is a cohort.

Question 2: A SPIDER-structured question asks how family caregivers experience arranging home care after a hospital discharge. Which design family fits?

The question asks about experience and meaning, which places it in the qualitative family. The other options are quantitative designs that estimate frequencies, associations or effects.

Question 3: A health authority offers a falls-prevention program at three clinics it selects and compares falls with three other clinics. Where does this study fall on the decision map?

The program is assigned on purpose, here by the health authority, and the clinics were chosen without a random process, so the study is quasi-experimental. Deliberate assignment by an agent other than the investigator, such as a health authority or a government, places a study in the experimental family on the decision map. A comparison group alone does not make a study a randomized trial.

Question 4: A team holds focus groups on how older adults describe barriers to transportation, turns the themes into survey items, and then surveys 800 older adults. Which design is this?

Qualitative data come first and are used to build the survey, which is the exploratory sequential design with integration by building. An explanatory sequential design would run the survey first.

Question 5: Which situation best suits a convergent design?

A convergent design collects both kinds of data in one phase on the same topics and merges them, which is the fastest of the core designs. Building an instrument suits an exploratory sequential design, and explaining survey results suits an explanatory sequential design.

Question 6: What is a joint display?

A joint display organizes quantitative and qualitative results by topic and states the conclusion drawn from reading them together, which may show confirmation, expansion or discordance.

Question 7: Why is the protocol described as the source document for the ethics application, the data access request and the methods section?

Each of these documents draws on the same decisions about design, participants, data and analysis. Writing those decisions once in the protocol and carrying them into each document keeps the documents consistent.

Question 8: Version 1.0 of a protocol was approved. The REB then approves an amendment that adds telephone interviews. Under the convention in this lesson, what is the amended protocol called?

Versions submitted for approval take whole numbers, so the amended protocol becomes version 2.0, and the amendment log records the change, its reason and the REB decision. Decimal numbers mark working drafts between submissions.

Question 9: The survey pilot cannot start until the REB approves the study. In project management terms, what is this relationship?

One task (the pilot) cannot start until another (ethics review) finishes, which is a finish-to-start dependency. REB approval is also a milestone, but the relationship between the two tasks is the dependency.

Question 10: Which practice best protects a timeline from delays in ethics review and data access?

The team cannot control how long these reviews take, so it starts them as early as possible, adds buffer time after them, and plans work for the waiting period. Optimistic estimates reflect the planning fallacy.

Question 11: A budget line reads "Focus group transcription: $720". What would a budget justification add?

A justification shows each amount as a quantity multiplied by a unit cost and explains why the expense is needed, so reviewers can check it against the protocol. Listing participants would also breach confidentiality.

Question 12: Why does the Cedar Valley folder structure keep raw data read-only and write cleaned data only through scripts?

Keeping the raw export unchanged means the team can always return to the original data, and writing cleaned files through scripts records every change so that it can be checked and repeated.

Question 13: What do the FAIR principles ask of research data?

Wilkinson and colleagues (2016) proposed that data be findable, accessible, interoperable and reusable. Accessible means available under clear conditions, which may include restrictions for sensitive data.

Question 14: How do the CARE principles relate to the FAIR principles?

The CARE principles (collective benefit, authority to control, responsibility and ethics) complement the FAIR principles by attending to Indigenous peoples' rights and the purposes data serve (Carroll et al., 2020).

Question 15: When should a research team agree on roles and authorship?

Roles and authorship belong in the team charter at the start of the project and should be revisited as roles change. Settling them when a manuscript is ready is usually harder.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Choosing an Approach and Managing a Research Project.

You can now take a structured research question through a decision map to a design family, choose and justify a mixed-methods design, write a protocol outline with a version history, and plan a project with a Gantt chart, a RACI matrix, a justified budget, an organized folder structure and a data management plan.

Lesson 7 turns to sampling, recruitment and measurement. It defines target, source and study populations, distinguishes probability from non-probability samples, introduces the participant flow diagram, and shows how to move from conceptual to operational definitions and choose validated instruments.

Continue to Lesson 7 →
Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

These terms, frameworks and people appear in this lesson on choosing a design and managing a research project.

Core Concepts
Study design The overall plan for how a study gathers and compares information to answer its question, including who is studied, what is measured, when, and whether the researchers intervene.
Design family A group of study designs that share the same basic logic, such as the descriptive, analytic observational, experimental, qualitative and evidence synthesis families.
Decision map A sequence of questions that leads from a research question to a design family; this lesson's map adapts the classification of Grimes and Schulz (2002).
Descriptive study An observational study without a comparison group that describes how often a condition occurs and how it is distributed by person, place and time.
Analytic observational study A study that compares groups to estimate whether an exposure is associated with an outcome, where no one assigns the exposure on purpose.
Experimental study A study of an exposure or intervention that is assigned on purpose, by the investigator or by another agent; with random allocation it is a randomized controlled trial, and without it a quasi-experimental study.
Quasi-experimental study A study of an intervention whose assignment, by the investigator or by another agent, is not random. Examples include studies with non-randomized comparison groups, interrupted time series and natural experiments.
Cross-sectional study A study that measures exposure and outcome at the same time in each participant.
Cohort study A study that starts with people who differ in exposure and follows them forward to see who develops the outcome.
Case-control study A study that starts with people who have the outcome (cases) and people who do not (controls) and compares their earlier exposures.
Mixed-methods research Research that collects and analyzes both quantitative and qualitative data and integrates them to answer its questions.
Convergent design A mixed-methods design that collects quantitative and qualitative data in the same phase, analyzes them separately and merges the results for comparison.
Explanatory sequential design A mixed-methods design in which a quantitative strand comes first and a qualitative strand follows to help explain its results.
Exploratory sequential design A mixed-methods design in which a qualitative strand comes first and its findings are used to build an instrument, intervention or hypotheses for a quantitative strand.
Integration The linking of quantitative and qualitative strands through connecting, building, merging or embedding, so that the study says more than either strand alone.
Meta-inference An overall conclusion drawn by integrating the inferences from the quantitative and qualitative strands of a study.
Frameworks & Tools
Mixed-methods notation Morse's (1991) shorthand in which capital letters mark a strand with priority, lower case a supporting strand, a plus sign concurrent strands and an arrow sequential strands.
Joint display A table or figure that places quantitative and qualitative results side by side by topic, together with the meta-inference drawn from them.
Research protocol The written plan for a study that records its decisions about questions, design, participants, data, analysis and safeguards in enough detail to be carried out as intended.
Protocol amendment A change to research that an REB has already approved, which must be approved before it is made unless needed to remove an immediate risk to participants.
SPIRIT statement A guideline listing the items a randomized trial protocol should contain; the current version is SPIRIT 2025 (Chan et al., 2025), which updated the 2013 statement.
Work breakdown structure A list of everything a project must produce, broken into tasks small enough that their duration can be estimated.
Gantt chart A chart that shows each task of a project as a horizontal bar on a calendar, often with milestones and dependencies marked.
Milestone A dated point in a project at which a key result is reached, such as REB approval or the close of a survey.
Critical path The longest chain of dependent tasks from the start of a project to its end, which sets the shortest possible duration of the project.
RACI matrix A table that assigns each person a role for each task: responsible, accountable, consulted or informed.
Budget justification A written explanation of how each budget amount was calculated, as a quantity multiplied by a unit cost, and why the expense is needed.
File naming convention A fixed pattern for file names, such as date, project, content and version, that every member of a team follows.
Version control Any system that records changes to files over time so that a team can see what changed, who changed it and why, and return to earlier versions.
Data management plan A living document describing how a project's data will be collected, documented, stored, protected, shared and preserved.
FAIR principles Principles stating that research data should be findable, accessible, interoperable and reusable (Wilkinson et al., 2016).
CARE principles Principles for Indigenous data governance: collective benefit, authority to control, responsibility and ethics (Carroll et al., 2020).
Key People
John W. Creswell American methodologist whose books on research design and mixed methods, including Designing and Conducting Mixed Methods Research with Vicki Plano Clark, describe the core designs used in this lesson.
Vicki L. Plano Clark Mixed-methods methodologist and co-author with John Creswell of Designing and Conducting Mixed Methods Research, which sets out the convergent, explanatory sequential and exploratory sequential designs.
Janice M. Morse Nurse researcher and qualitative methodologist who in 1991 proposed the notation used to describe the priority and timing of mixed-methods strands.
Jennifer C. Greene Evaluation methodologist who, with Caracelli and Graham (1989), identified five purposes for combining methods: triangulation, complementarity, development, initiation and expansion.
Michael D. Fetters Family physician and mixed-methods researcher who, with Curry and Creswell (2013), described integration at the design, methods and interpretation levels and promoted the joint display.
David A. Grimes Physician and epidemiologist who, with Kenneth Schulz, published a 2002 series in The Lancet on clinical research methods that included the classification of study designs adapted in this lesson's decision map.
Henry L. Gantt American engineer and management consultant who developed bar charts for scheduling work in the 1910s; the Gantt chart is named after him.
An-Wen Chan Physician and researcher who led the development of the SPIRIT guideline for trial protocols and published evidence of selective outcome reporting in randomized trials.
No matching entries. Try a different search term.