HSCI 207 · Lesson 9

Data Sources and Data Linkage

Research Methods in Health Sciences

Learning objectives for this lesson:

  • Distinguish primary data from secondary data and describe the roles of data custodians, data stewards, data platforms and researchers.
  • Describe the data available from Statistics Canada, including the confidential files held in Research Data Centres, and from the Canadian Institute for Health Information.
  • Describe Population Data BC, its counterparts in other provinces such as ICES in Ontario, and the role of Health Data Research Network Canada in research across provinces.
  • Outline the stages and contents of a data access request and plan realistically for its timelines and costs.
  • Carry out a simple deterministic and probabilistic record linkage and interpret the decisions it produces.
  • Explain missed matches and false matches and how uneven linkage error can bias the findings of a study.
  • Explain de-identification techniques, the separation principle and the Five Safes framework, and describe how secure research environments and output checking protect privacy.
  • Write a section on data sources, linkage and safeguards for a study’s data management plan, including a Five Safes table.

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University. It is the applied research methods course of the Public Health Assessment and Analysis series.

Lesson 9 · HSCI 207

Data Sources and Data Linkage

This lesson covers Canadian data sources, data access requests, record linkage and privacy safeguards.

About 3 hours
The running case

Cedar Valley needs data it did not collect

What the survey has

The survey measures loneliness among 1,600 adults aged 65 and older.

What the question needs

Emergency department visits are recorded in provincial health records.

The Cedar Valley Social Connection Study is fictional.

Roadmap

Four sections

1. Data sources

Section 1 maps where Canadian health data are held.

2. Access

Section 2 explains data access requests, timelines and costs.

3. Linkage

Section 3 teaches deterministic and probabilistic linkage.

4. Safeguards

Section 4 covers de-identification and the Five Safes.

Building on earlier lessons

What this lesson adds to a data management plan

  • A section on data sources, linkage and safeguards says where each data set comes from and how it is protected.
  • A five-row Five Safes table summarizes the safeguards.
  • The section applies even when a study collects all of its own data.
How to use this lesson

Working through the lesson

  • Each section ends with a reflection and a four-question knowledge check.
  • The glossary lists the terms, organizations and people in the lesson.
  • Real organizations change their processes, so their current guidance is the authority for any real project.
Section 1 of 5

The Canadian Research Data Environment

⏱ Estimated reading time: 35 minutes
Section 1 of 5

The Canadian Research Data Environment

This section maps who holds health data in Canada and how researchers reach them.

Two kinds of data

Primary and secondary data

Primary data

A research team collects primary data for its own study.

Secondary data

Secondary data were collected for another purpose, such as care or billing, and are reused for research.

Secondary data offer large populations and long follow-up, at the cost of limited variables and slow access.

Roles

Who does what

Custodian

The custodian is legally responsible for holding the data.

Steward

The steward decides whether a request is approved.

Platform

The platform prepares, links and hosts data for analysis.

Researcher

The researcher uses the data only as approved.

Statistics Canada

Three levels of access

Published tables

Published tables are free on the agency website.

Public use files

Public use microdata files have identifying details coarsened.

Master files

Master files are used only in Research Data Centres by approved researchers.

CIHI

Standardized pan-Canadian data

Discharge Abstract DatabaseNational Ambulatory Care Reporting SystemContinuing Care Reporting SystemHome Care Reporting SystemPrescription drug data

CIHI codes records to common national standards so that provinces can be compared.

Provincial centres

Linked provincial data

Population Data BC

PopData supports linkage in British Columbia, now with the Health Data Platform BC.

ICES, MCHP, Health Data NS

Similar centres serve Ontario, Manitoba and Nova Scotia.

HDRN Canada and distributed analysis support research across provinces.

Carry forward

From sources to access

  • Different questions need different sources.
  • The Cedar Valley question needs provincial records linked to the survey.
  • First Nations partners take part in decisions about data on their members.

Learning Objectives for this section

  • Distinguish primary data collection from the secondary use of existing data, and describe the roles of data custodians, data stewards and data platforms.
  • Describe the kinds of data Statistics Canada makes available and the three levels at which researchers can reach them: published tables, public use microdata files and confidential files in Research Data Centres.
  • Describe the role of the Canadian Institute for Health Information and the main databases it holds on hospital stays, emergency department visits, continuing care and prescriptions.
  • Describe Population Data BC and its counterparts in other provinces, including ICES in Ontario, and explain why research across provinces relies on networks such as Health Data Research Network Canada.
  • Match a research question to the Canadian data source most likely to answer it.

Introduction

Lesson 8 showed how researchers collect quantitative data directly and introduced the main kinds of administrative health records: physician billing claims, hospital discharge abstracts, prescription records and vital statistics. This lesson asks where existing data are held in Canada, how a researcher obtains permission to use them, how records from different sources are joined, and what safeguards protect the people whose information is used.

Primary data are data that a research team collects for its own study, such as the answers to a survey it designed. Secondary data were collected for another purpose and are reused for research. A hospital records a diagnosis so that the patient can be treated and the hospital funded; a researcher who later counts hospital stays for heart failure is making secondary use of that record. Secondary data offer large numbers of people, long follow-up and information about people who never answer surveys, with no extra burden on the people described. Their limitations follow from their origin: they contain only what the original purpose required, their coding reflects administrative rules, and access takes time.

Case: the Cedar Valley Social Connection Study (fictional)

The Cedar Valley Social Connection Study is a fictional mixed-methods study used throughout this course. A team led by Dr. Maya Hart at a British Columbia university is working with the fictional Cedar Valley Health Authority, which serves about 210,000 residents, about 46,000 of them aged 65 and older. The team’s survey of older adults received 1,600 completed responses, and 392 respondents (24.5 percent) scored 6 or higher on the three-item UCLA Loneliness Scale. One question asks whether loneliness is associated with emergency department visits. The survey measures loneliness, but the visits are recorded in health system data, so the team must find where those records are held, apply for access, link them to the responses of people who consented, and protect the linked file. Each section of this lesson follows one of those steps.

1.1 Who Holds Health Data in Canada

Provinces and territories deliver most health care, so most records of physician visits, hospital stays, prescriptions and insurance registrations are created by provincial and territorial systems and held by ministries of health, health authorities and agencies. National organizations hold two other kinds of data. Statistics Canada, the national statistical agency, conducts the census and a large program of surveys. The Canadian Institute for Health Information (CIHI) receives records from the provinces and territories and organizes them to common national standards so that they can be compared.

Several roles recur in every access process, although organizations name them differently. A data custodian is the organization legally responsible for holding and protecting a data set, such as a ministry of health. A data steward is the person or body that decides, on the custodian’s behalf, whether a request to use the data should be approved. A data platform or data centre prepares data for research, links records from different sources and provides a secure place to analyze them. The researcher requests access, analyzes the data and is accountable for using them as approved. In British Columbia, the Ministry of Health is the custodian of physician billing records, data stewards review requests to use them, and Population Data BC has long been one of the platforms through which researchers request and analyze them.

National organizations Provincial data centres Networks and hubs Statistics Canada Surveys, census and linked files CIHI Pan-Canadian hospital and drug data PopData BC British Columbia ICES Ontario MCHP Manitoba Health Data NS Nova Scotia CRDCN Research Data Centres on campuses HDRN Canada Support for multi-province requests Ministries, hospitals, registries and vital statistics agencies supply the records.
Figure 1.1. A simplified view of where Canadian health research data are held. Provincial centres hold most linked administrative data, national organizations hold survey and standardized pan-Canadian data, and networks connect researchers to them. Other provinces and territories have their own arrangements.

1.2 Statistics Canada Surveys and Research Data Centres

Statistics Canada collects information under the Statistics Act, which requires it to keep information about individual people confidential. Its surveys use probability samples (Lesson 7) and weights to describe provinces and, for larger surveys, health regions.

Canadian Community Health Survey (CCHS)v

The CCHS is a large cross-sectional survey of health status, health care use and the determinants of health among people aged 12 and older living in private households. It excludes some groups, including people living on First Nations reserves, people living in institutions such as long-term care homes, and full-time members of the Canadian Armed Forces. It asks about sense of belonging to the local community, which makes it useful in social connection research, but because it excludes long-term care residents it leaves out many of the oldest and frailest adults.

Canadian Health Measures Survey (CHMS)v

The CHMS combines an interview with direct physical measures, such as blood pressure and blood and urine samples, which provide measured values where other surveys rely on self-report, at the cost of a much smaller sample.

Census of Populationv

The census is conducted every five years, with a longer questionnaire for a sample of households (about one in four in recent censuses). Its tables describe the living arrangements of small areas, including the number of older adults who live alone.

General Social Survey and other social surveysv

The General Social Survey has covered a different theme in each cycle, such as caregiving, families and social identity, and in recent years Statistics Canada has asked about loneliness in some of its social surveys.

Three levels of access

LevelWhat it containsHow researchers obtain it
Published tablesCounts, percentages and estimates already calculated by Statistics CanadaFree on the agency’s website
Public use microdata files (PUMFs)One record per respondent, with identifying details removed or coarsened (for example, age in groups and broad geography)Through university libraries under the Data Liberation Initiative, a partnership between Statistics Canada and post-secondary institutions
Confidential master filesFull record-level files with detailed variables and fine geography, and some linked filesOnly for approved projects, inside a Research Data Centre or an approved virtual environment

A public use microdata file describes real respondents, but identifying details have been removed or coarsened so that no one can be recognized. Students and researchers can usually analyze a PUMF obtained through a university library, although a file that reports only the province cannot answer a question about one health region.

Research Data Centres and the CRDCN

A research data centre (RDC) is a secure facility where approved researchers analyze confidential Statistics Canada files. The Canadian Research Data Centre Network (CRDCN), a partnership between Statistics Canada and Canadian universities, operates RDCs on more than thirty campuses, including campuses in British Columbia. RDCs hold survey master files, some administrative files such as tax and hospitalization records, and files from the agency’s Social Data Linkage Environment, which links survey, census and administrative records. Statistics Canada has, for example, linked responses from some health surveys to hospital and mortality records for respondents who agreed to linkage.

The researcher submits a proposal explaining why the confidential data are needed. Approved researchers undergo security screening and are sworn in as deemed employees of Statistics Canada, which places them under the confidentiality obligations of the Statistics Act. They work on computers that cannot send files out, and every result they wish to take away is reviewed by Statistics Canada staff to confirm that it cannot reveal information about an individual. The network has introduced a virtual RDC for remote work under similar controls, and fees depend on the institution and type of project, so students should ask local RDC staff about current arrangements.

What this means for Cedar Valley

The Cedar Valley team could use census tables to describe how many older adults in the region live alone and the CCHS to compare community belonging among older adults in British Columbia with other provinces. Neither can answer the team’s main question, because Statistics Canada files cannot be joined to the Cedar Valley survey.

1.3 The Canadian Institute for Health Information

CIHI is an independent, not-for-profit organization, created in 1994, that collects health system data from the provinces, territories and health care organizations and publishes information about Canada’s health systems. Its value to researchers lies in standardization. Hospitals code diagnoses with the Canadian version of the International Classification of Diseases (ICD-10-CA) and procedures with the Canadian Classification of Health Interventions (CCI), so a hospital stay for pneumonia is recorded the same way in Nova Scotia and British Columbia.

Discharge Abstract Database (DAD)Click to explore
National Ambulatory Care Reporting System (NACRS)Click to explore
Continuing Care Reporting System (CCRS)Click to explore
Home Care Reporting System (HCRS)Click to explore
National Prescription Drug Utilization Information System (NPDUIS)Click to explore

CIHI’s public reports and interactive tools give aggregate results, such as hospitalization rates by province, without any application. Researchers who need custom tables or record-level files submit a data request, which CIHI reviews for privacy and which can involve cost-recovery fees. Record-level CIHI files are de-identified and are not usually linked to a researcher’s own survey. Studies that need that kind of linkage more often use the provincial version of the same records through a provincial data centre, because the province holds the health insurance numbers that make linkage possible.

1.4 Population Data BC and Its Counterparts in Other Provinces

Population Data BC (PopData) is a data and education resource hosted at the University of British Columbia. It has long helped researchers request, link and analyze British Columbia administrative data. These include the Ministry of Health, whose data sets cover physician services billed to the Medical Services Plan, hospital discharges, prescriptions recorded in PharmaNet and registration files, along with sources such as vital statistics, the cancer registry and education records. Researchers submit a data access request through PopData’s Data Access Unit, data stewards review it, and approved researchers analyze the data remotely in a secure environment. PopData can also link data that researchers collected themselves, such as survey responses, to administrative records, provided that participants explicitly agreed to the linkage on a consent form that meets the data providers’ requirements and has been approved by a research ethics board.

Arrangements in British Columbia have been changing. In 2024 the Ministry of Health began accepting academic data requests through the Health Data Platform BC, and since August 2025 new research projects that request Ministry of Health data have been handled under that program through a shared PopData and Health Data Platform BC request process. Analysis is moving from PopData’s Secure Research Environment to a cloud-based Trusted Analysis Environment. Students should check the current process before planning a project; the principles in this lesson apply in either arrangement.

OrganizationProvinceWhat it does
Population Data BC, with the Health Data Platform BCBritish ColumbiaSupports requests for and linkage of provincial administrative data, including researcher-collected data with consent
ICES (originally the Institute for Clinical Evaluative Sciences)OntarioAn independent, not-for-profit research institute that holds linked, coded health data for Ontario residents under the province’s health privacy law; researchers outside ICES reach the data through its Data and Analytic Services
Manitoba Centre for Health Policy (MCHP)ManitobaA research centre at the University of Manitoba that maintains the Manitoba Population Research Data Repository
Health Data Nova ScotiaNova ScotiaA data centre at Dalhousie University that supports access to linked provincial health data

Other provinces and territories have their own arrangements, which also change over time.

Research across provinces

Provincial privacy laws and data agreements usually require record-level health data to stay in the province, so a study comparing provinces cannot simply pool their files. One solution is distributed analysis: the team writes a common protocol and analysis code, analysts in each province run the code on their own data, and only summary results are combined. The Canadian Network for Observational Drug Effect Studies (CNODES) has used this approach to study the safety of medications across several provinces. Health Data Research Network Canada (HDRN Canada) is a non-profit network of provincial, territorial and pan-Canadian data organizations that works to make research across regions easier. Its Data Access Support Hub offers researchers a single entry point and a common request form for data from more than one region, an inventory of available data sets, and an inventory of algorithms (tested definitions of conditions written in terms of administrative codes).

1.5 Cohort Studies, Registries and Indigenous-Governed Data

Some sources were created for research from the start. The Canadian Longitudinal Study on Aging (CLSA) follows more than 50,000 adults who were aged 45 to 85 when recruited and collects information on health, function and social participation over many years; researchers apply through its own access process. Disease registries, such as provincial cancer registries, record every diagnosed case of a condition in a defined population.

Data about First Nations, Inuit and Métis people raise questions of governance that go beyond privacy law. The First Nations Information Governance Centre conducts the First Nations Regional Health Survey and holds it under the OCAP® principles of ownership, control, access and possession described in Lesson 4. Provincial records also include First Nations people, and the use of data that identify them can require approval under agreements with First Nations governance bodies. For the Cedar Valley team, the fictional Cedar Valley First Nations Health Centre takes part in decisions about whether and how data about its members are requested, analyzed and reported.

Try it: match the question to the source

For each Cedar Valley question, name the source you would try first and one limitation of that source. (1) How many adults aged 65 and older in the region live alone? (2) Are respondents who score 6 or higher on the loneliness scale more likely than others to visit an emergency department in the following year? (3) How does the rate of hospital stays among older adults in British Columbia compare with other provinces? (4) Is social participation in mid-life associated with later changes in health among Canadian adults?

Suggested answersv

(1) Census tables, which count people who live alone; living alone is a different concept from loneliness. (2) The team’s survey linked to provincial records through Population Data BC; only respondents who consented can be included, and emergency department coverage must be checked for the years needed. (3) CIHI reports or a CIHI data request using the Discharge Abstract Database; differences between provinces can reflect how services are organized as well as differences in health. (4) The Canadian Longitudinal Study on Aging, which follows the same adults over time; its volunteers may be healthier than the population as a whole.

Section 2 turns from where the data are held to how a researcher obtains permission to use them.

Reflection

A health authority planner asks you which data source to use for each of three questions. (a) What proportion of adults aged 65 and older in each community of the region live alone? (b) Do older adults in British Columbia who report a weak sense of belonging to their local community have more hospital stays than those who report a strong sense of belonging? (c) How do emergency department visit rates among older adults in British Columbia compare with those in Ontario? Four sources are available. Census tables give counts of living arrangements for small areas and are free online. The Canadian Community Health Survey asks people aged 12 and older living in private households about their sense of community belonging; it excludes people living in long-term care, and its confidential files, including some files linked to hospital records for respondents who agreed to linkage, are available only in Research Data Centres after a proposal is approved. CIHI’s National Ambulatory Care Reporting System holds emergency department visits in a standard national format, but coverage of emergency departments differs between provinces. Population Data BC supports access to linked British Columbia administrative data, which contain no measure of community belonging. For each question, choose a source, explain why it fits, and name one limitation you would report.

Model answer

(a) Census tables are the best source, because they count living arrangements for small areas, including each community, and they are free. One limitation is that living alone is a different concept from loneliness, so the tables describe a risk factor and say nothing about how people feel. Counts for very small communities may also be rounded.

(b) The Canadian Community Health Survey linked to hospital records, analyzed in a Research Data Centre, fits because it is the only source that combines a measure of community belonging with hospital stays. Population Data BC holds the hospital records but no belonging measure. Limitations include the exclusion of people in long-term care, who are among the heaviest users of hospitals, and the restriction to respondents who agreed to linkage. Access requires an approved proposal, so the work would take months.

(c) CIHI’s National Ambulatory Care Reporting System fits because it records visits in the same format in both provinces. Before comparing, I would check that emergency departments in both provinces report for the years needed, because coverage differs. I would also note that differences in rates may reflect how emergency and primary care are organized in each province as well as differences in health.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A researcher counts hospital stays for heart failure using discharge abstracts that hospitals created for patient care and funding. Which statement describes these data?

Secondary data were collected for another purpose, here patient care and hospital funding, and are reused for research. Choosing which records to count does not make data primary; primary data are collected by the research team for its own study.

Question 2: A researcher needs detailed geography to study loneliness among older adults in a single health region using Statistics Canada survey data. Which level of access does the project require?

Fine geography and detailed variables are kept only in the confidential master files, which approved researchers analyze in a Research Data Centre. Public use microdata files coarsen geography to protect respondents, which is why they are tempting but unsuitable here. CIHI does not hold Statistics Canada survey files.

Question 3: Why do studies that link a researcher’s own survey to health records in British Columbia usually work through a provincial data platform instead of CIHI?

Provincial custodians hold the records with personal health numbers, which allow a consenting respondent’s survey answers to be joined to his or her health records. CIHI does provide some record-level data, but those files are de-identified and are not usually linked to a researcher’s own survey. Ethics approval is required in either route.

Question 4: A team wants to compare hospital use among older adults in British Columbia, Ontario and Manitoba, and each province requires its record-level data to stay in the province. Which approach fits?

Distributed analysis keeps record-level data in each province: analysts run the same code locally and only summary results are combined, as networks such as CNODES have done. Pooling the records would breach the requirement, and assuming similarity would answer a different question.
Section 2 of 5

Applying for Access: Data Access Requests, Timelines and Costs

⏱ Estimated reading time: 30 minutes
Section 2 of 5

Applying for Access

This section covers the stages, contents, review, timelines and costs of a data access request.

The process

Typical stages of a request

  • The researcher makes an enquiry, checks feasibility and obtains a cost estimate.
  • The research ethics board approves the study and the researcher submits the request.
  • Data stewards review the request, and the team signs agreements and completes training.
  • The platform links and releases the data, and outputs are checked before release.
Contents

What a request contains

Purpose and public benefitStudy populationData sets, years and variablesLinkage and consentPeople and securityAnalysis and outputs

Cedar Valley requested records for the 1,312 respondents who consented to linkage.

Data minimization

Request only what the question needs

Requested

The team requested age in years, health service area, visit dates and main diagnoses.

Left out

The team left out full birth dates, full postal codes and physicians’ names.

Review

Who reviews, and on what grounds

Data stewards

Stewards consider lawful authority, public benefit, minimization and security.

Ethics board

TCPS 2 Article 5.7 requires ethics approval before data linkage.

Indigenous governance

First Nations partners review requests involving their members.

The institution

The university signs the research agreement.

Timelines and costs

Plan for months and budget for fees

  • Requests, approvals and data release commonly take months.
  • Fees can cover data preparation, linkage and the secure environment.
  • A cost estimate obtained early can go into a grant application.
Carry forward

From approval to linkage

  • The approved request defines who and what will be linked.
  • Linkage involves decisions that can produce errors.
  • Section 3 links six pairs of records by two methods.

Learning Objectives for this section

  • Describe the typical stages of a data access request, from the first enquiry to the closure of the project.
  • Prepare the main parts of a data access request, including a study population definition and a justified list of data sets and variables.
  • Explain who reviews a request and on what grounds, including data stewards, research ethics boards and Indigenous governance bodies.
  • Explain why access to administrative data takes months, identify the costs a project should budget for, and build data access into a project timeline.

Introduction

Administrative health data cannot be downloaded. A researcher who wants to use them must make a formal application, usually called a data access request, which describes the research, the people and records it needs, and the safeguards that will protect them. The request is reviewed by the people responsible for the data, and access is granted only for the approved purpose. This section describes that process in general terms. Each organization has its own forms and rules, which change over time, so the details of any real application come from the organization’s current guidance and its data access staff.

Planning for access begins well before the request is submitted. The Cedar Valley team (a fictional study) knew from the start that it wanted to link survey responses to health records. It therefore wrote the linkage into its consent form (Lesson 5), asked respondents for their personal health number in the survey (Lesson 8), and contacted Population Data BC while the survey was still being designed. Of the 1,600 adults aged 65 and older who completed the survey, 1,312 (82.0 percent) consented to linkage and provided a personal health number. Those 1,312 people are the population for which the team can request linked records.

2.1 The Stages of a Data Access Request

Processes differ between organizations, but most follow a similar sequence. Figure 2.1 shows eight typical stages. Population Data BC, for example, describes its process in stages that run from planning the request through data steward review, contracts and privacy training, data release, analysis and pre-publication review, to project closure.

1 Enquiry and intake meeting 2 Feasibility and cost estimate 3 Research ethics board approval 4 Submit the request 5 Data steward review 6 Agreements and privacy training 7 Linkage and data release 8 Analysis, output review, closure Planning and approval (teal) comes before access and use (red). Requests often return to earlier stages when reviewers ask for changes.
Figure 2.1. Typical stages of a data access request. The order and names differ between organizations, and some organizations accept the request and the ethics application in parallel.

In the enquiry and intake stage, the researcher contacts the organization’s data access staff, describes the question and learns which data sets might answer it. In the feasibility and cost stage, the researcher confirms that the data exist for the needed years and population, reads the data set documentation (often called metadata or a data dictionary), and asks for a cost estimate, which funders may want to see in a grant application. The ethics stage obtains approval from a research ethics board (REB). The submission stage completes the request form. In data steward review, the people responsible for each data set decide whether to approve the request, often after asking questions. The agreements and training stage covers the signed research agreement, confidentiality undertakings from every team member, privacy training and the creation of accounts. In the linkage and release stage, the platform links and extracts the records and places them in a secure environment. The final stage covers analysis, output review and closure: results are checked before they leave the environment, some organizations review publications before release, and at the end of the project the data are destroyed or retained as the agreement specifies.

2.2 What a Data Access Request Contains

A data access request is a short research proposal written for a particular audience. Data stewards want to know whether the use is appropriate, whether the amount of data requested is the minimum the question needs, and whether the people and settings involved will keep the data safe. The tabs below describe the usual parts of a request, with the Cedar Valley version of each.

What to write. State the research question, the reason it matters for the health of the population, and how the findings will be used. Reviewers look for a clear public benefit and a question that the requested data can answer.

Cedar Valley. The team asks whether adults aged 65 and older who report loneliness have more emergency department visits in the year after the survey than those who do not, after accounting for age, sex, prior health conditions and community. The findings will inform the health authority’s planning of social connection programs.

What to write. Define exactly who is included, using criteria that the data provider can apply. This definition is often called the cohort definition. A vague definition, such as “older adults in the region”, cannot be turned into an extract.

Cedar Valley. The population is the 1,312 survey respondents who consented to linkage and provided a personal health number. No other residents are requested.

What to write. List each data set, the years required and each variable, with a one-line justification for every variable. Request the least detailed version that answers the question, for example age in years rather than full date of birth.

Cedar Valley. The worked example below lists the team’s data sets and variables.

What to write. Explain which files will be linked, what identifiers will be used and who will handle them, and state whether participants consented to linkage or whether a waiver of consent is requested.

Cedar Valley. Respondents gave written consent to linkage. The team will send the identifiers of consenting respondents to the platform’s linkage staff, separately from their survey answers.

What to write. Name every person who will see the data and describe their role and training. Describe where the analysis will take place.

Cedar Valley. Dr. Hart and the graduate research assistant will analyze the data inside the secure environment. The community research associate and the advisory group will see only approved summary results.

What to write. Summarize the planned analysis and the kinds of results that will be released, and describe how findings will be shared with the communities involved.

Cedar Valley. The team will compare the proportion of lonely and non-lonely respondents with an emergency department visit and will report results by community only where counts are large enough to protect privacy.

Data minimization

The principle that runs through every part of a request is data minimization: a project should receive only the people, years and variables that its question requires, at the least identifying level of detail that will work. Population Data BC’s guidance puts the idea directly: the selection of data fields in an extract must be restrictive. Minimization reduces the harm that would follow a breach, and it makes approval easier, because every extra variable is another item a data steward must be persuaded to release.

Worked example: the Cedar Valley data request (fictional)

The table lists the data sets the team requested for the 1,312 consenting respondents. The years are expressed relative to each person’s survey date.

Data setPeriodMain variablesJustification
Registration and demographic fileTwo years before to one year afterAge in years, sex, health service area of residence, months of coverageConfirms residence and coverage during follow-up and supplies adjustment variables
Physician billing records (Medical Services Plan)Two years before to one year afterService date, diagnostic code, type of practitionerIdentifies prior chronic conditions and primary care contact before the survey
Hospital discharge abstractsTwo years before to one year afterAdmission and discharge dates, main diagnosisIdentifies prior hospital stays, a marker of poorer health
Emergency department visit recordsOne year afterVisit date, triage level, main diagnosisMeasures the outcome
Death records (vital statistics)One year afterDate of deathEnds follow-up for people who died

The team also recorded what it chose not to request. It did not ask for full dates of birth, because age in years was sufficient, or for full postal codes, because the health service area identified each community. It did not ask for physicians’ identities, which the question does not require. Before submitting, the team confirmed with data access staff that emergency department records were available for every community in the region for the years needed.

2.3 Who Reviews a Request, and on What Grounds

Several groups review a request, and each looks at it from a different angle.

Data stewards decide whether the use is permitted under the law and agreements that govern each data set and whether it is appropriate. In British Columbia, the Freedom of Information and Protection of Privacy Act sets the conditions under which public bodies may disclose personal information for research, and Ontario’s Personal Health Information Protection Act plays a similar role for ICES. Stewards consider the public benefit of the research, whether the data requested are the minimum needed, whether consent was obtained or its absence is justified, and whether the people and settings involved are secure.

Research ethics boards review the study under the Tri-Council Policy Statement (TCPS 2), whose rules on privacy, secondary use of information and consent Lesson 5 described. TCPS 2 contains a specific article on data linkage (Article 5.7). It requires REB approval before linkage takes place, asks researchers to describe the data to be linked and the likelihood that linkage will create identifiable information, and, where it will, asks researchers to show that the linkage is essential to the research and that appropriate security measures will protect the information. Data providers usually require proof of current ethics approval before they release data, and any change to the request later may require an amendment to both approvals.

Indigenous governance bodies have authority over data about their communities under agreements and under the principles described in Lesson 4. In the Cedar Valley study, the fictional Cedar Valley First Nations Health Centre reviewed the request, agreed with the team on how results about First Nations participants would be reported, and asked to review any such results before release.

The researcher’s institution signs the research agreement, because the university, as well as the researcher, takes on legal obligations for the data.

2.4 Timelines and Costs

Access to linked administrative data takes time. Population Data BC describes the request, approval and data provisioning process as one that can take some months, and requests involving several data providers, new linkages or researcher-collected data tend to take longer. Researchers should treat any study that depends on linked administrative data as a long-term undertaking, and a study that must finish within a few months is rarely feasible unless the data are already available to the team.

Why requests are delayedv

Requests are often delayed for predictable reasons. The study population may be defined in terms that the data provider cannot apply. The list of variables may be long and poorly justified, which prompts questions from stewards. Ethics approval may be missing, expired or written for a different version of the study. A request may involve several data providers, each with its own review. Agreements may wait for signatures from the university. Data preparation and linkage are done in a queue with other projects. Changes made after approval, such as adding a variable, usually require an amendment.

How to shorten the waitv

A researcher can shorten the wait by contacting data access staff early, attending an intake meeting with a clear question and population, reading the data documentation before choosing variables, keeping the request as small as the question allows, obtaining ethics approval that explicitly covers the linkage, and starting the request while primary data collection is still under way.

Access also costs money. Data platforms commonly charge fees to recover the cost of preparing, linking and extracting data and of providing a secure analysis environment, and the size of those fees depends on the complexity of the request. Statistics Canada RDC access can involve fees that depend on the institution and project, and CIHI custom requests can involve cost-recovery charges. A project budget should also include the time of an analyst who can work with large administrative files, the time all team members need for privacy training, and the cost of amendments. A cost estimate obtained during the feasibility stage allows these items to be included in a grant application.

ActivityCedar Valley planning allowance
Intake meeting, feasibility check and cost estimateMonths 1 to 2, while the survey is being built
Ethics approval covering linkage and submission of the requestMonths 2 to 4
Data steward review, questions and revisionsMonths 4 to 7
Agreements, confidentiality undertakings, privacy training and accountsMonths 7 to 9
Linkage, data preparation and releaseMonths 10 to 12
Analysis and output reviewFrom month 13

These allowances are the fictional team’s own planning assumptions. Actual times depend on the request and the organization, and the team confirmed its assumptions with data access staff at the intake meeting. The months are counted from the intake meeting, which the team held in month 3 of the project, so access approval falls in project month 11 and the linked data arrive in month 14, as the Gantt chart in Lesson 6 shows. Lesson 6 also showed how to place such allowances in a project timeline so that analysis and reporting are scheduled realistically.

Try it: justify three variables

Suppose you want to study whether older adults who live alone are less likely to fill their prescriptions after a hospital stay. Write one sentence of justification for each of three variables you would request (for example, the dispensing date from prescription records). For each one, state the least detailed version that would still answer the question. Then name one variable you would deliberately leave out and explain why.

Once access is approved, the platform must find each consenting respondent in the administrative files. Section 3 describes how that record linkage is done and what happens when it goes wrong.

Reflection

Two research assistants draft a request for linked administrative data for a small pilot study. They plan to link a survey of 400 family caregivers to provincial physician billing records. Their draft states that the study population is “caregivers in the region”; requests full dates of birth, full postal codes, every physician billing record from 2000 to the present, and the names of the physicians each person saw; does not mention consent to linkage; notes that ethics approval has not yet been sought; says the analysis will be done on one assistant’s personal laptop; and expects the data within six weeks so that the pilot can be finished within four months. A data access request usually describes its purpose, study population, data and variables, linkage and consent, people and security, and analysis and outputs, and it passes through enquiry, feasibility and cost, ethics approval, submission, data steward review, agreements and training, linkage and release, and output review. Identify at least four problems with the draft and propose a revised plan.

Model answer

The study population is too vague for a data provider to apply. It should be defined as the caregivers among the 400 survey respondents who consented to linkage and supplied a personal health number. The variable list breaks the principle of data minimization: age in years and a broad area of residence would replace full dates of birth and postal codes, the years should be limited to the period the question needs (for example, two years before and one year after the survey), and physicians’ names should be dropped because the question does not need them. Consent is missing: linkage of researcher-collected data to provincial health records requires each participant’s explicit agreement to the linkage on a consent form that meets the data providers’ requirements and has REB approval. Ethics approval must come first, because TCPS 2 requires REB approval before linkage and providers ask for proof of approval. A personal laptop is unacceptable; the data would stay in the platform’s secure environment. Finally, six weeks is unrealistic, because requests, approvals and data release commonly take months.

A revised plan would keep the pilot feasible by preparing the request for later submission while the team analyzes a public use microdata file now. If the linkage goes ahead, the team would contact data access staff first, obtain a cost estimate, and schedule the linked analysis for a later phase of the study.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: What is the main purpose of the feasibility stage of a data access request?

The feasibility stage confirms that the data exist for the population and years the question needs and produces a cost estimate, which can go into a grant application. Undertakings, output checking and ethics approval are separate stages.

Question 2: Which change to a draft data request best applies the principle of data minimization?

Data minimization means requesting only the people, years and variables the question requires, at the least identifying level of detail that works. Age in years and a health service area answer most questions that a full birth date and postal code would. The other options all enlarge the request beyond what the question needs.

Question 3: What does TCPS 2 Article 5.7 require of researchers who propose to link data?

Article 5.7 requires research ethics board approval before data linkage and asks researchers to describe the data to be linked and the likelihood that identifiable information will be created. Consent is one possible basis for linkage, but TCPS 2 also allows REBs to consider linkage without consent under defined conditions.

Question 4: A research assistant hopes to obtain linked provincial administrative data and finish the analysis within three months. What is the most accurate advice?

Requests, data steward review, agreements and data preparation commonly take months, so linked administrative data rarely fit a three-month timeline. Ethics review is required, and research assistants can use administrative data as members of approved projects.
Section 3 of 5

Record Linkage: Deterministic and Probabilistic Methods and Linkage Error

⏱ Estimated reading time: 40 minutes
Section 3 of 5

Record Linkage

This section covers deterministic and probabilistic linkage, linkage error and its consequences.

Vocabulary

Matches and links

Match

A match is a pair of records that truly belongs to one person.

Link

A link is the decision to treat two records as one person.

The theory of probabilistic linkage comes from Fellegi and Sunter (1969).

Deterministic linkage

Rule: link if the health numbers are identical

Linked

Pairs 1, 2 and 4 link correctly.

Kept apart

Pairs 5 and 6 are correctly left unlinked.

Missed

Pair 3 is missed because two digits were transposed.

Probabilistic linkage

Weights from two probabilities

Agreement weight
\[ w = \log_2\left(\frac{m}{u}\right) \]

Date of birth

Agreement adds 10.9 points, because u is very small.

Sex

Agreement adds only 1.0 point, because u is about one half.

Thresholds

Scoring the six pairs

Link (15 or more)

Pairs 1, 2, 3 and 5 link; pair 5 is a false match.

Clerical review

Pair 4 scores 12.9, and the reviewer accepts it.

Non-link (below 5)

Pair 6, the spouses, scores below zero.

Linkage error

Missed matches and false matches

Missed match

A missed match undercounts the person’s events.

False match

A false match adds another person’s events.

Uneven missed matches shrank the illustrative Cedar Valley gap from 10.5 to about 8 percentage points.

Carry forward

From linkage to protection

  • The Cedar Valley linkage joined 1,303 of 1,312 consenting respondents.
  • A linked file reveals more than any of its sources.
  • Section 4 describes the safeguards that protect it.

Learning Objectives for this section

  • Define record linkage and describe the identifiers used to link health records and the errors they commonly contain.
  • Carry out a simple deterministic linkage using an exact rule and explain how stepwise rules extend it.
  • Explain how probabilistic linkage adds agreement and disagreement weights into a total weight, and apply upper and lower thresholds with clerical review.
  • Distinguish missed matches from false matches and explain how each can bias the results of a study.
  • Describe what researchers can do to detect, reduce and report linkage error.

Introduction

After its request was approved, the Cedar Valley team (a fictional study) needed the platform to find each of its 1,312 consenting respondents in the provincial files. That task is record linkage: bringing together records that belong to the same person (or family, address or organization) from different sources, or from within one source. Linkage allows a survey answer about loneliness to sit beside a record of an emergency department visit, and it is also a source of error that is invisible in the final data set.

3.1 What Record Linkage Is

Halbert Dunn (1946), an American public health statistician, used the term “record linkage” to describe assembling the records of a person’s life, from birth to death, into a single “book of life”. Howard Newcombe and colleagues (1959), working in Canada, showed that a computer could link birth and marriage records automatically by weighing how strongly agreement on names and other details suggested that two records belonged together. Ivan Fellegi and Alan Sunter (1969), statisticians at what is now Statistics Canada, gave probabilistic linkage the mathematical form it still has. Fellegi later served as Chief Statistician of Canada.

A linkage key is the identifier or combination of identifiers used to decide whether two records belong together, such as a health number, or a surname combined with a date of birth. A match is a pair of records that truly belong to the same person. A link is the decision, made by a rule or a reviewer, to treat two records as belonging to the same person. Linkage error is the gap between them.

3.2 Identifiers and Why They Disagree

Health records in British Columbia carry a personal health number (PHN), a lifetime identifier assigned to each person registered for provincial health insurance. It is the strongest linkage key available, because each person has one and no two people should share one. Names, birth dates, sex and postal codes can all disagree between two records of the same person.

IdentifierCommon reasons two records of the same person disagree
Personal health numberDigits mistyped or transposed when copied into a survey or form; number missing
SurnameSpelling variants (MacLeod and McLeod); change of name on marriage; hyphenated names recorded in part; different transliterations from other alphabets
Given nameShort forms and nicknames (William and Bill); use of a middle name; names recorded in a different order
Date of birthDay and month reversed; typing errors; missing values replaced by a default date
SexRecording errors; changes in recorded sex or gender over time
Postal codeMoves between the dates of the two records; a mailing address in one record and a home address in the other

These errors are unevenly spread. People who move often, whose names are often misspelled or transliterated, or who have changed their names are more likely to have records that disagree. Older adults who have recently moved to a smaller town, the group at the centre of one Cedar Valley qualitative question, are one example.

3.3 Deterministic Linkage

Deterministic linkage links two records when they agree exactly on a specified linkage key or set of identifiers. The simplest rule uses one high-quality identifier: link two records if their personal health numbers are identical. A stepwise (or multi-pass) deterministic linkage applies a sequence of rules, starting with the strictest; records left unlinked by the first rule are tried against a second, such as agreement on surname, given name, date of birth and sex.

Worked example, part 1: six candidate pairs (fictional)

Each row pairs a survey record with a hospital record, showing the last four digits of each health number. The final column shows the truth, which the linkage team would not know.

PairSurvey recordHospital recordPHN (survey / hospital)Truth
1Helen Ostrowski, F, 14 Mar 1951, V0X 2B1Helen Ostrowski, F, 14 Mar 1951, V0X 2B14471 / 4471Same person
2Donald MacLeod, M, 2 Jun 1948, V0X 1C5Donald McLeod, M, 2 Jun 1948, V0X 1C52093 / 2093Same person
3Margaret Chen, F, 30 Sep 1955, V0X 3K2Margaret Chen, F, 30 Sep 1955, V0X 3K28812 / 8821Same person
4William Thomas, M, 11 Jan 1944, V0X 1C5Bill Thomas, M, 11 Jan 1944, V1A 4R95530 / 5530Same person
5Ruth Okafor, F, 22 Jul 1950, V0X 2B1Joan Okafor, F, 22 Jul 1950, V0X 2B16604 / 6648Different people (twin sisters)
6Helen Ostrowski, F, 14 Mar 1951, V0X 2B1Peter Ostrowski, M, 9 Nov 1949, V0X 2B14471 / 3307Different people (spouses)

Rule: link if the personal health numbers are identical. Pairs 1, 2 and 4 link, despite the spelling difference in pair 2 and the different given name and postal code in pair 4, because the rule looks only at the health number. Pairs 5 and 6 are correctly left unlinked. Pair 3 is a missed match: the respondent transposed two digits when writing her health number. A second pass requiring agreement on surname, given name, date of birth and sex would recover pair 3.

Deterministic linkage is transparent and works very well when a unique identifier is recorded accurately in both files. A single error in the linkage key prevents a link, however, because a strict rule cannot tell one transposed digit from a different person.

3.4 Probabilistic Linkage

Probabilistic linkage compares several identifiers and asks how strongly the pattern of agreement and disagreement supports the conclusion that two records belong to the same person. It is used when no unique identifier is available in both files, or when the identifier is incomplete. Following Fellegi and Sunter (1969), each identifier is described by two probabilities.

The two probabilities behind each weight

The m-probability is the probability that an identifier agrees when the two records truly belong to the same person; it is a little below 1 because of recording errors, nicknames and moves. The u-probability is the probability that the identifier agrees by chance when the records belong to different people; for sex it is about one half, and for a full date of birth among older adults it is very small.

An agreement weight is large when m is high and u is low. Software converts the ratio m ÷ u to a logarithmic scale so that weights can be added; each additional point means the observed pattern is twice as likely for a true match as for a non-match. A disagreement weight, built from (1 − m) ÷ (1 − u), is negative.

Total weight = the sum of the agreement or disagreement weight for each identifier.

Software usually estimates the probabilities from the data and applies blocking, comparing only pairs that agree on a simple variable such as year of birth, over several passes.

Worked example, part 2: scoring the same six pairs (fictional)

Suppose the hospital file had no health numbers, so the team must link on name, date of birth, sex and postal code. The first table shows illustrative probabilities and weights; the second scores the six pairs.

IdentifiermuWeight if it agreesWeight if it disagrees
Surname0.950.01+6.6−4.3
Given name0.900.01+6.5−3.3
Date of birth0.970.0005+10.9−5.1
Sex0.990.50+1.0−5.6
Postal code0.800.01+6.3−2.3
PairSurnameGiven nameBirth dateSexPostal codeTotalDecisionTruth
1+6.6+6.5+10.9+1.0+6.331.3LinkSame
2−4.3+6.5+10.9+1.0+6.320.4LinkSame
3+6.6+6.5+10.9+1.0+6.331.3LinkSame
4+6.6−3.3+10.9+1.0−2.312.9Clerical reviewSame
5+6.6−3.3+10.9+1.0+6.321.5LinkDifferent
6+6.6−3.3−5.1−5.6+6.3−1.1Non-linkDifferent

The team set an upper threshold of 15 (pairs at or above it are linked) and a lower threshold of 5 (pairs below it are not linked). Pairs in between go to clerical review, in which a trained person examines the records and decides. Agreement on a full birth date earns far more weight than agreement on sex, because different people rarely share a birth date and often share a sex. Pair 3 now links, because the transposed health number plays no part. The reviewer accepts pair 4, because Bill is a common short form of William and the postal code shows a move within the region. Pair 5 is a false match: twin sisters who live together agree on everything except their given names, and their total of 21.5 passes the upper threshold.

Non-link Clerical review Link Different people Same person missed matches false matches −10 0 10 20 30 Total weight (thresholds at 5 and 15)
Figure 3.1. A schematic sketch of total weights. Most candidate pairs involve different people and score low; true matches score high. The tails that cross the thresholds become missed matches and false matches, and the review zone catches some of each.

Moving the thresholds trades one error for another. Raising the upper threshold to 25 would send pairs 2 and 5 to clerical review, where the reviewer would probably reject the twins, at the cost of more manual review; lowering the thresholds would reduce missed matches and increase false matches.

3.5 Linkage Error and How It Is Measured

A missed match (a false negative) occurs when two records of the same person are not linked. A false match (a false positive) occurs when records of two different people are linked. When the truth is known for a sample of pairs, for example from a careful manual check, two measures summarize performance. Sensitivity is the proportion of true matches that were linked. Positive predictive value is the proportion of links that are true matches.

Method in the worked exampleTrue matches linkedLinks that were true matchesErrors
Deterministic, health number only3 of 4 (sensitivity 75 percent)3 of 3 (positive predictive value 100 percent)One missed match (pair 3)
Probabilistic, with clerical review4 of 4 (sensitivity 100 percent)4 of 5 (positive predictive value 80 percent)One false match (pair 5)

Six pairs are too few to judge a method, but they show the usual pattern: deterministic rules on a strong identifier tend to produce few false matches and some missed matches, while probabilistic methods recover more true matches at the risk of some false ones. Many linkage units combine the two, with a deterministic pass on the health number followed by probabilistic passes for the records that remain.

3.6 The Consequences of Linkage Error

The consequences of linkage error depend on how unlinked records are handled and on whether errors are spread evenly across the groups being compared (Harron et al., 2017; Doidge & Harron, 2019).

Missed matches undercount eventsClick to explore
False matches add the wrong eventsClick to explore
Uneven error creates biasClick to explore
Consent to linkage selects peopleClick to explore
Worked example, part 3: how uneven missed matches shrink a difference (illustrative)

Suppose that 90 of the 298 consenting lonely respondents (30.2 percent) and 200 of the 1,014 other consenting respondents (19.7 percent) truly had at least one emergency department visit in the following year, a gap of 10.5 percentage points. Now suppose a weaker linkage that relies on postal codes misses 10 percent of lonely respondents’ visits, because they moved more often, and 3 percent of other respondents’ visits. The analysis would count 81 lonely respondents with a visit (90 minus 9) and 194 others (200 minus 6). The observed proportions would be 81 ÷ 298 = 27.2 percent and 194 ÷ 1,014 = 19.1 percent, a gap of about 8 percentage points. The association would appear about one quarter smaller, entirely because of the linkage.

In the actual (fictional) Cedar Valley linkage, the platform linked 1,268 of the 1,312 consenting respondents in a deterministic pass on health number and date of birth and 30 more in probabilistic passes; 5 of 8 pairs sent to clerical review were accepted. In total, 1,303 respondents (99.3 percent) were linked. Five of the 9 unlinked respondents were lonely (1.7 percent of 298) and 4 were not (0.4 percent of 1,014), a small error in the same direction as the example, which the team reported.

3.7 What Researchers Can Do

Ask for information about linkage qualityv

Researchers usually receive linked data without identifiers, so they depend on the linkage unit for information about quality. The GUILD guidance (Gilbert et al., 2018) asks those who link data to share their methods, the proportion of records linked and how linkage rates differ between groups.

Compare linked and unlinked recordsv

Analysts can compare linked and unlinked people on every variable in the original file, such as age, community and loneliness score. A difference warns that missed matches may bias the results.

Look for implausible records and test the resultsv

Events after a recorded death can signal false matches. A sensitivity analysis repeats the main analysis under different assumptions, for example using only the most certain links; a conclusion that holds under each is less likely to be an artefact of the linkage.

Report the linkagev

The RECORD statement (Benchimol et al., 2015), an extension of the STROBE reporting guideline for studies using routinely collected health data, asks authors to describe the linkage methods and the quality of the linkage. Lesson 12 introduces reporting guidelines more fully.

Try it: score two new pairs

Use these weights (agreement: surname +6.6, given name +6.5, date of birth +10.9, sex +1.0, postal code +6.3; disagreement: surname −4.3, given name −3.3, date of birth −5.1, sex −5.6, postal code −2.3) and thresholds of 5 and 15. Pair A agrees on surname, given name, sex and postal code, but its birth dates are 04 Jul 1949 and 07 Apr 1949. Pair B agrees on given name, date of birth and sex, and disagrees on surname and postal code. Calculate each total, state the decision, and say what a clerical reviewer would look for.

Answerv

Pair A: 6.6 + 6.5 − 5.1 + 1.0 + 6.3 = 15.3, just above the upper threshold, so the pair is linked. The reversed day and month is a common recording error, so a link is plausible, although a cautious team might review pairs this close to the threshold. Pair B: −4.3 + 6.5 + 10.9 + 1.0 − 2.3 = 11.8, which goes to clerical review. The reviewer would check for a spelling variant or change of surname and for a move within the region.

Section 4 describes how a linked file, which reveals more than any of its sources, is protected.

Reflection

A linkage uses these weights. Agreement: surname +6.6, given name +6.5, date of birth +10.9, sex +1.0, postal code +6.3. Disagreement: surname −4.3, given name −3.3, date of birth −5.1, sex −5.6, postal code −2.3. Pairs scoring 15 or more are linked, pairs scoring below 5 are not linked, and pairs in between go to clerical review. Pair X agrees on surname, given name, date of birth and sex and disagrees on postal code. Pair Y agrees on surname, sex and postal code and disagrees on given name and date of birth. Pair Z agrees on date of birth and sex and disagrees on surname, given name and postal code. (1) Calculate each total and state the decision. (2) In the full study, which compares hospital visits between older adults who moved in the past year and those who did not, 12 percent of movers’ records and 2 percent of other participants’ records fail to link, and unlinked participants are counted as having no visits. Explain the likely effect on the comparison and describe two things the researchers could do about it.

Model answer

(1) Pair X: 6.6 + 6.5 + 10.9 + 1.0 − 2.3 = 22.7, which is above 15, so it is linked; a changed postal code is consistent with a move. Pair Y: 6.6 − 3.3 − 5.1 + 1.0 + 6.3 = 5.5, which falls just inside the review zone. A reviewer would notice that the two records share a surname and an address but have different given names and birth dates, which suggests two members of one household, and would probably reject the link. Pair Z: −4.3 − 3.3 + 10.9 + 1.0 − 2.3 = 2.0, which is below 5, so it is not linked.

(2) Movers’ visits would be undercounted far more than other participants’ visits, because 12 percent of movers would appear to have no visits whatever happened to them. If movers truly have more visits, the observed difference would shrink and could disappear; if they have fewer, it would be exaggerated. The bias comes from the linkage, so a reader could not detect it from the results. The researchers could ask the linkage unit for linkage rates by mover status and compare linked and unlinked participants on the survey variables. They could also repeat the analysis using only participants whose records linked, or add a deterministic pass on health numbers if those are available, and report the linkage methods and rates following the RECORD statement.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In record linkage, what is the difference between a match and a link?

A match is a true pair of records belonging to the same person, whereas a link is the decision, by a rule or a reviewer, to treat two records as belonging together. Linkage error is the gap between the two: a match that is not linked or a link that is not a match.

Question 2: Why does agreement on full date of birth receive a much larger weight than agreement on sex?

The agreement weight grows as the u-probability, the chance that two different people agree, falls. About half of all pairs of different people share a sex, while very few share a full birth date. Both identifiers have high m-probabilities, so the difference comes from u.

Question 3: Using a lower threshold of 5 and an upper threshold of 15, a pair of records scores 12.9. What happens to it?

Pairs at or above the upper threshold are linked and pairs below the lower threshold are rejected. Pairs in between, such as this one, go to clerical review, where a trained person examines the records and decides.

Question 4: In a study, 10 percent of lonely respondents’ records and 3 percent of other respondents’ records fail to link, and unlinked respondents are counted as having no emergency department visits. What is the likely effect?

Missed matches make people appear to have no visits, so they understate visit rates, and here they understate the lonely group’s rate more. If lonely respondents truly visit more often, the observed difference shrinks. The errors are uneven, so they do not cancel out.
Section 4 of 5

Privacy Safeguards: De-identification, the Five Safes and Secure Research Environments

⏱ Estimated reading time: 35 minutes
Section 4 of 5

Privacy Safeguards

This section covers de-identification, the Five Safes and secure research environments.

Identifiability

Direct and indirect identifiers

Direct

Names, health numbers and street addresses identify a person alone.

Indirect

Birth dates, postal codes and rare diagnoses identify people in combination.

Sweeney (2000) showed how few details can single a person out.

De-identification

Techniques and the separation principle

RemovalPseudonymizationGeneralizationSuppressionAggregation

Under the separation principle, linkage staff see identifiers and researchers see content.

Five Safes

Five questions, adjusted as a set

Safe projects

The use is appropriate and approved.

Safe people

The researchers are trained and accountable.

Safe settings

The setting prevents unauthorized use.

Safe data

The data are no more identifying than needed.

Safe outputs

Released results cannot identify anyone.

Output checking

Checking a Cedar Valley table

Before

Five cells fell between 1 and 4 across four communities.

After

Combining the smaller communities left every cell at 5 or more.

Results about First Nations participants were reviewed by the health centre before release.

Putting it together

A data sources, linkage and safeguards section

  • State whether the study uses existing data, and if so who holds them and how access is requested.
  • State whether data will be linked, and check that the consent form asks permission.
  • Complete a five-row Five Safes table for the study.
Next steps

Reflection and final assessment

  • The section reflection applies the Five Safes to a flawed data-sharing plan.
  • The knowledge check has four questions on this section.
  • The final assessment has a reflection and fifteen questions across the lesson.

Learning Objectives for this section

  • Distinguish direct identifiers from indirect identifiers and explain why combinations of indirect identifiers can identify people, especially in small communities.
  • Describe the main de-identification techniques, including removal, pseudonymization, generalization, suppression and aggregation, and explain the separation principle used in linkage.
  • Explain the Five Safes framework and use it to assess a data access arrangement.
  • Describe the features of a secure research environment and the checks applied to results before they are released.
  • Apply these safeguards in the data management plan for a study.

Introduction

A linked file can reveal more about a person than any of its sources. The Cedar Valley survey (a fictional study) records how lonely a respondent feels; the health records show when that person went to an emergency department and why. Joined together, they form a detailed account of a person’s life that the person shared on the understanding that it would be protected. Lesson 5 introduced privacy and confidentiality under TCPS 2, the categories of identifiable information, and data management plans. This section describes the specific safeguards that make the research use of linked administrative data possible: reducing how identifiable the data are, controlling who uses them and where, and checking every result before it is released.

4.1 How People Are Identified

A direct identifier identifies a person on its own: a name, a personal health number, a street address, a telephone number or an email address. An indirect identifier, also called a quasi-identifier, does not identify anyone on its own but can do so in combination with other information. Date of birth, postal code, sex, ethnicity, occupation, a rare diagnosis and the dates of hospital stays are all indirect identifiers.

The power of combinations is easy to underestimate. Latanya Sweeney (2000), a computer scientist, estimated that about 87 percent of the population of the United States could be uniquely identified by three items: five-digit ZIP code, sex and full date of birth. Removing names from a file therefore does little to protect people if the file keeps detailed dates and locations. TCPS 2 reflects this by distinguishing directly identifying, indirectly identifying, coded, anonymized and anonymous information (Lesson 5), and by treating information as identifiable whenever it could reasonably be expected to identify a person, alone or in combination with other available information.

Risk is higher in small populations. In a city of 90,000 people, a 91-year-old woman who visited an emergency department in March is one of many. In a small community such as the fictional Kestrel Lake in the Cedar Valley region, a neighbour who knows those three facts might recognize her. This is why rural communities, small First Nations communities and people with rare conditions need particular care in the release of results, and why the release rules described below often require small groups to be combined.

4.2 De-identification Techniques

De-identification is the set of techniques that reduce the chance that a person can be identified from a data set, while keeping the data useful for the research question. It reduces risk to a low level without removing it entirely, so it is always combined with the other safeguards in this section. Khaled El Emam (2013), a Canadian researcher in this field, describes de-identification as a process of measuring the risk of re-identification and applying techniques until the risk is acceptably low for the setting in which the data will be used.

TechniqueWhat it doesCedar Valley example
RemovalDeletes direct identifiers from the analysis fileNames, health numbers and street addresses never enter the analysis file
Pseudonymization (coding)Replaces identifiers with a code; the key that connects codes to people is held elsewhereEach respondent carries a project-specific study number that means nothing outside this project
GeneralizationReplaces precise values with broader categoriesAge in years replaces date of birth; the health service area replaces the postal code
SuppressionRemoves or groups rare valuesAges above 95 are recorded as “95 or older”
AggregationReleases counts or summaries for groups instead of recordsThe advisory group sees tables of counts, never individual records
Rounding and perturbationAlters released numbers slightly so that small counts cannot be read exactlyStatistics Canada, for example, randomly rounds many census counts

The separation principle

Linkage needs identifiers, and analysis needs content, but no single person needs both. The separation principle organizes linkage around that fact. Data providers send identifying information (names, health numbers, birth dates) to a linkage unit, which uses them to work out which records belong together and assigns each person a project-specific study number. The content (loneliness scores, visit dates, diagnoses) travels separately, without names or numbers, and is joined using the study numbers inside a secure environment. Linkage staff therefore see identifiers without health content, and researchers see health content without identifiers. Because the study numbers are specific to one project, files released to different projects cannot be joined to each other.

Survey file consenting respondents Health records data providers Linkage unit sees identifiers only assigns study numbers Content scores, dates, diagnoses no names or numbers Secure research environment study number + content Output checking before any release identifiers and study numbers content without identifiers
Figure 4.1. The separation principle. Identifiers (teal) and content (red) travel by different routes and meet only as study numbers and de-identified content inside the secure environment.

4.3 The Five Safes Framework

De-identification protects the data themselves. Other safeguards control the project, the people, the setting and the results. The Five Safes framework, developed by Felix Ritchie and colleagues from work on secure data access at the United Kingdom’s Office for National Statistics (Desai et al., 2016), organizes all of these safeguards as five questions. Statistical agencies and data centres in several countries use it to design and explain their access arrangements, and it is a practical tool for planning any study that uses sensitive data. Click each card to read the question it asks.

Safe projectsClick to explore
Safe peopleClick to explore
Safe settingsClick to explore
Safe dataClick to explore
Safe outputsClick to explore

The five dimensions work as a set. When one is strong, others can be lighter, and when one must be weak, others must compensate. A public use microdata file has very safe data, because detail has been removed, so students and researchers at participating institutions can analyze it on an ordinary computer. A confidential master file in a Research Data Centre has much less safe data, because it keeps detail that researchers need, so access requires screened people, a controlled setting and checked outputs. The table compares three arrangements described in Section 1.

DimensionPublic use microdata fileResearch Data CentreProvincial linked data in a secure environment
ProjectsAny lawful use under the licence termsProposal approved before accessData access request approved by data stewards, with ethics approval
PeopleUsers at participating institutionsScreened researchers sworn in as deemed employeesNamed team members with privacy training and confidentiality undertakings
SettingsThe user’s own computerSecure physical or virtual centreRemote secure research environment
DataDetail removed or coarsenedDetailed record-level filesDe-identified, linked records with study numbers
OutputsNo check requiredEvery output reviewed by Statistics Canada staffOutputs checked against the platform’s rules

The Five Safes is a planning framework. Privacy law, data agreements, TCPS 2 and Indigenous data governance under principles such as OCAP® apply alongside it, and a First Nations partner may set conditions on projects, people and outputs that go further than any of the five questions require.

4.4 Secure Research Environments and Output Checking

A secure research environment (also called a trusted research environment) is a controlled computing system in which approved researchers analyze sensitive data without being able to remove them. Population Data BC’s Secure Research Environment, the cloud-based Trusted Analysis Environment that is replacing it, the analysis environments at ICES and the virtual Research Data Centre all follow this model, although their details differ. Common features include remote access to a virtual desktop rather than a copy of the data, strong authentication when logging in, no internet, email or copy-and-paste out of the environment, a record of what each user does, and a single controlled route by which files enter and results leave.

That single route is where output checking happens. Before a table, figure or model result leaves the environment, it is checked for disclosure risk, either by the researcher against the platform’s rules, by platform staff, or both. The main risks are listed below.

Small cellsv

A table cell based on very few people can identify them, especially when combined with other information. Many organizations require counts below a set threshold (for example, below five) to be suppressed or combined with other categories. The exact rule differs between organizations, so researchers use the rule that applies to their data.

Differencingv

Two tables that are each safe can be unsafe together. If one table reports 40 people and a second, otherwise identical table excludes one small community and reports 38, subtraction reveals a count of 2. Checking outputs as a set prevents this.

Extreme values and individual pointsv

A maximum or minimum value often belongs to a single person. Scatter plots and lists of outliers can show individual records. These are usually replaced with percentiles or summary measures.

Model resultsv

Regression results are usually low risk, but a model fitted to a very small subgroup, or one that includes a category with very few people, can reveal information about those people.

Worked example: checking a Cedar Valley output (fictional)

The graduate research assistant prepared the following table of linked respondents with three or more emergency department visits in the year after the survey. The team’s rule, confirmed with the platform, is that no released cell may be between 1 and 4.

GroupCedar CityRiversideNorth BenchKestrel LakeTotal
Lonely (score 6 or higher)943218
Not lonely1254122

Five cells fall between 1 and 4, so the table cannot be released. Suppressing those five cells alone would leave a problem, because the totals and the remaining cells would allow some suppressed values to be worked out by subtraction. The team instead combined the three smaller communities.

GroupCedar CityOther communitiesTotal
Lonely (score 6 or higher)9918
Not lonely121022

Every cell is now 5 or more. The team also checked that no other output in the same report showed Kestrel Lake separately in a way that would allow differencing, and it sent results that concerned First Nations participants to the Cedar Valley First Nations Health Centre for review before release, as the partners had agreed.

4.5 Worked Example: The Cedar Valley Five Safes Plan

The Cedar Valley team summarized its safeguards in a Five Safes table that it attached to its data management plan and its data access request. A plan of this kind shows reviewers, partners and participants how each risk is handled.

DimensionCedar Valley safeguards (fictional)
Safe projectsThe research ethics board approved the study, including linkage; the data stewards approved the data access request; the Cedar Valley First Nations Health Centre agreed to the use of data about its members; the purpose is to inform the health authority’s social connection programs.
Safe peopleOnly Dr. Hart and the graduate research assistant can open the linked data; both completed the platform’s privacy training and signed confidentiality undertakings; the university signed the research agreement; the community research associate and the advisory group see only released results.
Safe settingsAll analysis of linked data takes place in the platform’s secure environment; the team stores the survey identifiers on the university’s secure server in a separate location from the survey answers, as its data management plan describes.
Safe dataLinked records carry project-specific study numbers; age in years replaces date of birth; the health service area replaces the postal code; ages above 95 are grouped.
Safe outputsNo released cell may be between 1 and 4; small communities are combined; the platform checks outputs; results about First Nations participants are reviewed by the health centre before release.

Writing the data sources, linkage and safeguards section

A data management plan can include a section titled “Data sources, linkage and safeguards” with three parts. First, it states whether the study uses any existing data source; if it does, it names the source, the organization that holds it and the route by which access is requested, using the descriptions in Sections 1 and 2. Second, it states whether the study links any data; if it does, it names the identifiers, explains who holds them, and confirms that the consent form asks participants for permission to link. Third, it includes a five-row Five Safes table modelled on the Cedar Valley example in Section 4.5. When a study uses only data the team collects itself, the table still applies: it describes how identifiers are separated from responses, where the data are stored, who sees them, and the rule for small numbers in reported results.

Reflection

A researcher has linked survey and hospital data for 250 older adults living in five communities, the smallest of which has about 300 residents. She plans two things. First, she will email a spreadsheet of the linked records to the six members of a community advisory group so that they can look for patterns; the spreadsheet contains a study number, full date of birth, full postal code, sex, community, diagnosis codes and the date of each emergency department visit. The advisory group members have not completed privacy training or signed confidentiality undertakings. Second, she will publish a table of emergency department visits by community and five-year age group in which several cells contain 1, 2 or 3 people. The Five Safes framework asks whether the project, the people, the setting, the data and the outputs are safe. Identify a problem under each of the five dimensions and propose a revised plan.

Model answer

Safe projects: sharing record-level data with the advisory group is probably outside the use approved by the data stewards and the ethics board, which named who would see the data. Safe people: the advisory group members have no training and no confidentiality undertakings, so they should not receive record-level data at all. Safe settings: email and personal computers offer no control over copying or onward sharing; linked data should stay inside the secure research environment. Safe data: full date of birth, full postal code and visit dates would allow neighbours in a community of 300 people to recognize individuals, even without names. Safe outputs: cells of 1 to 3 people could identify individuals, especially in the smallest community.

In a revised plan, the researcher would keep all record-level data in the secure environment and analyze it there herself. She would generalize ages to broader groups and combine the smaller communities, so that every cell in the published table met the platform’s threshold, and she would check that no other table allowed suppressed values to be recovered by subtraction. She would then meet the advisory group to interpret the checked, aggregate tables together, which keeps their knowledge in the analysis without exposing anyone’s records.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which of the following is an indirect identifier, also called a quasi-identifier?

An indirect identifier does not identify anyone alone but can do so in combination with other information, as date of birth does with postal code and sex. Health numbers, names, street addresses and email addresses identify a person on their own, so they are direct identifiers.

Question 2: Under the separation principle used in data linkage, who sees the names and health numbers used to link the records?

The separation principle sends identifiers to a linkage unit that sees no health content, and sends content without identifiers to the researchers. No one needs both, so no one receives both.

Question 3: A public use microdata file can be analyzed on an ordinary computer, while confidential master files require a Research Data Centre. How does the Five Safes framework explain the difference?

The five dimensions work as a set. Public use files have detail removed, which makes the data very safe, so the controls on people and settings can be lighter. Master files keep detail, so they require screened people, a controlled setting and checked outputs. Public use files still describe real respondents.

Question 4: A table shows 2 lonely respondents in Kestrel Lake with three or more emergency department visits. Under a rule that no released cell may be between 1 and 4, what is the best response?

Combining small categories removes the small cell while keeping the table truthful and useful. Names are not needed for neighbours to recognize people in a small community, changing values falsifies the results, and the advisory group would be the people most likely to recognize the respondents.
Section 5 of 5

Final Assessment

⏱ Estimated time: 25 minutes

Bringing It All Together

This lesson followed the path that existing data take from the systems that create them to a research result. In Canada, most administrative health records are created by provincial and territorial health systems, while Statistics Canada holds the census and national surveys and CIHI organizes records from the provinces to common national standards. Researchers reach these data through published tables, public use microdata files, Research Data Centres, CIHI data requests and provincial platforms such as Population Data BC and ICES, and networks such as HDRN Canada support research that crosses provincial borders.

Access depends on a data access request that defines the study population precisely, justifies each variable, explains the linkage and consent, names the people involved and describes the secure setting. Data stewards, research ethics boards and Indigenous governance bodies each review the request, and the process commonly takes months and carries costs that belong in the project budget. Once approved, records are joined by deterministic rules, probabilistic weights or both. Both methods make errors, and missed matches and false matches can bias a study’s results when they fall unevenly on the groups being compared.

The linked file is protected by layers of safeguards. De-identification and the separation principle reduce what any one person can see, and the Five Safes framework asks whether the project, the people, the setting, the data and the outputs are each safe enough. Secure research environments and output checking keep record-level data inside a controlled system and ensure that released results cannot identify anyone, which matters most in small communities such as those in the fictional Cedar Valley region.

Key Takeaways from this lesson

  • Secondary data were collected for another purpose, such as patient care or billing, and offer large populations and long follow-up at the cost of limited variables and slow access.
  • Provinces and territories hold most administrative health records, Statistics Canada holds census and survey data, and CIHI organizes provincial records to common national standards.
  • Statistics Canada data are available as published tables, as public use microdata files and, for approved projects, as confidential master files in Research Data Centres.
  • Population Data BC, ICES, the Manitoba Centre for Health Policy and Health Data Nova Scotia support access to linked provincial data, and HDRN Canada and distributed analysis support research across provinces.
  • A data access request defines the study population exactly, justifies each variable under the principle of data minimization, and describes linkage, consent, people, setting and outputs.
  • Data stewards, research ethics boards and Indigenous governance bodies review requests on different grounds, and TCPS 2 Article 5.7 requires ethics approval before data linkage.
  • Access to linked administrative data commonly takes months and carries fees, so it belongs in the project timeline and budget from the start.
  • Deterministic linkage applies exact rules to identifiers, while probabilistic linkage adds agreement and disagreement weights and uses thresholds with clerical review.
  • Missed matches undercount events and false matches attach the wrong events, and either can bias comparisons when it falls unevenly on the groups being compared.
  • De-identification, the separation principle, the Five Safes framework, secure research environments and output checking work together to protect the people whose records are used.

Core Concepts Reviewed

Section 1: primary and secondary data, data custodians and stewards, Statistics Canada surveys, public use microdata files, Research Data Centres and the CRDCN, CIHI databases, Population Data BC and the Health Data Platform BC, ICES and other provincial centres, HDRN Canada and distributed analysis.

Section 2: the stages of a data access request, the study population definition, data minimization, data steward and ethics review including TCPS 2 Article 5.7, Indigenous data governance, and timelines and costs.

Section 3: record linkage, linkage keys, matches and links, deterministic and stepwise linkage, m- and u-probabilities, agreement weights, thresholds and clerical review, missed and false matches, sensitivity and positive predictive value, and linkage bias.

Section 4: direct and indirect identifiers, de-identification techniques, the separation principle, the Five Safes framework, secure research environments, output checking and differencing.

The final reflection asks you to plan the data sources, access, linkage and safeguards for a new study from start to finish.

Reflection

A community organization in British Columbia runs a weekly meal program for older adults. It keeps attendance records with each participant’s name, date of birth, sex and postal code, but no health numbers. Of 600 participants, 480 signed a consent form allowing their attendance records to be linked to provincial hospital records for research; some participants attend with a spouse who shares their surname and address. A researcher wants to know whether regular attendance is associated with fewer hospital stays in the following year. Write a short data plan (about 200 words) that (1) names who holds each kind of data and the general route to access, (2) lists three things the data access request must contain, (3) states which linkage approach the identifiers allow and one linkage error to expect, with its likely effect, and (4) gives one safeguard for each of the Five Safes (projects, people, settings, data and outputs).

Model answer

(1) The community organization holds the attendance records, and the Ministry of Health is the custodian of the hospital discharge records, which the researcher would request through the shared Population Data BC and Health Data Platform BC process, after an intake meeting and a feasibility and cost check. (2) The request must define the population as the 480 consenting participants, list each variable with a justification (admission and discharge dates, main diagnosis, age in years and health service area), and include evidence of written consent and ethics approval that explicitly covers linkage, as TCPS 2 Article 5.7 requires.

(3) Without health numbers, the linkage must be probabilistic, or stepwise deterministic, on name, date of birth, sex and postal code. Spouses who attend together share a surname and postal code, so a false match between them is possible; it would attach one spouse’s hospital stays to the other and blur any real difference between regular and occasional attenders. Missed matches among people who moved are also likely and would undercount their stays.

(4) Projects: ethics and data steward approval with a clear public benefit. People: only named, trained analysts who have signed confidentiality undertakings. Settings: analysis only in the secure research environment. Data: study numbers, age in years and health service area replace identifiers. Outputs: small cells combined or suppressed and every table checked before release.

Minimum 30 characters required.

✓ Reflection saved

Final Knowledge Assessment

Final Assessment, this lesson: Data Sources and Data Linkage (15 Questions)

Question 1: The Cedar Valley team needs emergency department visits for survey respondents who consented to linkage. Which source and route fit best?

Only the provincial records can be linked to the Cedar Valley survey, because the province holds them with health numbers and the respondents consented. Census tables, survey microdata and CIHI reports describe other people or give aggregate results.

Question 2: Which statement about Statistics Canada Research Data Centres is accurate?

RDC users have approved projects, pass security screening, are sworn in as deemed employees under the Statistics Act, and have every output reviewed before it leaves. Record-level files never leave the centre, and RDCs hold Statistics Canada survey master files.

Question 3: A researcher wants to study emergency department visits using a CIHI database. Which database is it, and what should be checked first?

Emergency department visits are recorded in the National Ambulatory Care Reporting System, and its coverage of emergency departments differs between provinces. The Discharge Abstract Database holds inpatient discharges.

Question 4: What does Health Data Research Network Canada’s Data Access Support Hub provide?

The hub coordinates requests for multi-regional data through one entry point and a common form, with inventories of data sets and algorithms. The data stay with each provincial or territorial organization, and ethics review is still required.

Question 5: A data steward asks why a request includes full postal codes. Which reply applies data minimization?

Data minimization asks for the least identifying version of each variable that answers the question. Postal codes are indirect identifiers, so removing names does not make them harmless, and requests for possible future questions enlarge the request without justification.

Question 6: Which reviewers decide whether a use of data is permitted under the law and agreements governing each data set and whether the data requested are the minimum needed?

Data stewards review requests on behalf of the custodians, considering lawful authority, public benefit, minimization, consent and security. Research ethics boards review the study under TCPS 2, and output checkers review results after analysis.

Question 7: In the lesson’s worked example, pair 3 had a transposed digit in the survey’s health number. The deterministic rule missed it, but probabilistic linkage linked it. Why?

The probabilistic linkage in the example did not use the health number at all, and pair 3 agreed on every identifier it did use, giving the maximum total weight of 31.3. A strong identifier with an error defeats an exact rule, while other identifiers can still carry the link.

Question 8: In the worked example, twin sisters who live together scored 21.5 and were linked. What kind of error is this, and which change would most likely catch it?

Linking records of two different people is a false match. Raising the upper threshold above 21.5 would send the pair to clerical review, where a reviewer would see the different given names. Lowering the upper threshold would make automatic links more likely.

Question 9: A linkage correctly links all 4 true matches in a sample and makes 5 links in total. What are its sensitivity and positive predictive value?

Sensitivity is the proportion of true matches linked, 4 ÷ 4 = 100 percent. Positive predictive value is the proportion of links that are true matches, 4 ÷ 5 = 80 percent, because one of the five links is a false match.

Question 10: In the Cedar Valley survey, 76.0 percent of lonely respondents and 83.9 percent of other respondents consented to linkage. What does this imply for the linked file?

Only consenting respondents can be linked, and lonely respondents consented less often, so they are under-represented in the linked file. This is a selection issue that arises before any linkage algorithm runs, and the figures show that consent is related to loneliness.

Question 11: Latanya Sweeney estimated that most of the United States population could be uniquely identified by sex combined with which two items?

Sweeney (2000) estimated that about 87 percent of the United States population could be uniquely identified by five-digit ZIP code, sex and full date of birth, which shows that removing names does not by itself make data anonymous.

Question 12: Which safeguard belongs to the safe settings dimension of the Five Safes?

Safe settings concern the place where data are used, such as a secure environment that blocks downloads and records activity. Grouping ages is a safe data measure, training is a safe people measure, and checking tables is a safe outputs measure.

Question 13: Why might suppressing only the small cells in a table still fail to protect privacy?

This risk is called differencing. When totals and other cells are published, a suppressed value can often be recovered by subtraction, which is why the Cedar Valley team combined small communities. Tables of counts contain no direct identifiers.

Question 14: Which statement about the Five Safes framework is accurate?

The five dimensions are adjusted together: very safe data, as in a public use file, allow lighter controls on people and settings. The framework is a planning tool that sits alongside privacy law, TCPS 2 and Indigenous data governance, and it applies to any study using sensitive data.

Question 15: A study uses only interviews that the research team conducts. How does this lesson apply to its data management plan?

The Five Safes applies to any sensitive data, including interview transcripts: the plan should separate identifiers from transcripts, store data securely, limit who sees them and check that reported quotations and results cannot identify participants. No linkage or Research Data Centre is involved.
✦ Complete the final reflection above before submitting

Congratulations!

You have successfully completed this lesson: Data Sources and Data Linkage.

You can now describe where Canadian health data are held and how researchers reach them, prepare the main parts of a data access request, carry out and interpret a simple deterministic and probabilistic linkage, explain how linkage error can bias a study, and apply de-identification and the Five Safes framework to protect the people whose records are used. You can also write the data sources, linkage and safeguards section of a data management plan.

Lesson 10 turns to qualitative data collection. It covers in-depth interviews, key informant interviews and focus groups, writing and piloting an interview guide, conducting and moderating sessions, and recording and transcription, including the privacy questions that automated transcription raises. The Cedar Valley team’s 24 interviews and four focus groups will serve as the worked example.

Reference

Glossary: Key Terms, People & Frameworks

📚 Reference page, available throughout the lesson

These terms, organizations and people appear in this lesson on Canadian data sources, record linkage and privacy safeguards.

Core Concepts
Primary data Data that a research team collects for its own study, such as the answers to a survey it designed.
Secondary data Data collected for another purpose, such as patient care or billing, and reused for research.
Data custodian The organization legally responsible for holding and protecting a data set, such as a provincial ministry of health.
Data steward The person or body that decides, on a custodian’s behalf, whether a particular request to use the data should be approved.
Data access request A formal application to use data held by another organization, describing the research, the study population, the data and variables, the linkage, the people and the safeguards.
Data minimization The principle that a project should receive only the people, years and variables its question requires, at the least identifying level of detail that will work.
Distributed analysis An approach in which analysts run common code on data held separately in each jurisdiction and only summary results are combined, so that record-level data never leave their home province.
Record linkage The process of bringing together records that belong to the same person, family, address or organization from different sources or from within one source.
Personal health number The lifetime identifier assigned to each person registered for provincial health insurance in British Columbia, and the strongest linkage key in provincial health records.
Deterministic linkage Linkage that joins two records when they agree exactly on a specified identifier or set of identifiers; stepwise versions apply a sequence of rules from strictest to least strict.
Probabilistic linkage Linkage that adds agreement and disagreement weights across several identifiers into a total weight and compares it with thresholds, following Fellegi and Sunter (1969).
m-probability and u-probability The m-probability is the chance that an identifier agrees for records of the same person; the u-probability is the chance that it agrees by chance for records of different people.
Agreement weight The score added when an identifier agrees, larger when the m-probability is high and the u-probability is low; disagreement weights are negative.
Clerical review Examination by a trained person of record pairs whose total weight falls between the lower and upper thresholds, to decide whether they belong to the same person.
Missed match Two records of the same person that the linkage failed to join; also called a false negative.
False match Records of two different people that the linkage joined as if they belonged to one person; also called a false positive.
Indirect identifier Information such as date of birth, postal code or a rare diagnosis that does not identify a person alone but can do so in combination with other information; also called a quasi-identifier.
De-identification The set of techniques, such as removal, pseudonymization, generalization, suppression and aggregation, that reduce the chance that a person can be identified from a data set.
Pseudonymization Replacing identifiers with a code, such as a project-specific study number, while the key connecting codes to people is held separately.
Separation principle An arrangement in which identifiers go to a linkage unit that sees no health content, and content goes to researchers without identifiers, so that no one sees both.
Output checking Review of tables, figures and model results before they leave a secure environment, to confirm that they cannot identify individuals through small cells, differencing or extreme values.
Frameworks & Tools
Five Safes framework A framework that asks whether the project, the people, the setting, the data and the outputs of a data access arrangement are safe, with the five dimensions adjusted as a set (Desai et al., 2016).
Public use microdata file (PUMF) A Statistics Canada file with one record per respondent in which identifying details have been removed or coarsened, available to university users through the Data Liberation Initiative.
Research Data Centre (RDC) A secure facility, operated through the Canadian Research Data Centre Network, where approved researchers sworn in as deemed employees analyze confidential Statistics Canada files and have their outputs reviewed before release.
Canadian Institute for Health Information (CIHI) An independent, not-for-profit organization that collects health system data from the provinces and territories to common national standards, including the Discharge Abstract Database and the National Ambulatory Care Reporting System.
Population Data BC A data and education resource hosted at the University of British Columbia that supports requests for, and linkage and analysis of, British Columbia administrative data; since 2025 Ministry of Health data are requested through a shared process with the Health Data Platform BC.
ICES An independent, not-for-profit research institute in Ontario that holds linked, coded health data for Ontario residents and provides access to researchers outside ICES through its Data and Analytic Services.
Health Data Research Network Canada A non-profit network of provincial, territorial and pan-Canadian data organizations whose Data Access Support Hub helps researchers request data from more than one region.
Secure research environment A controlled computing system, also called a trusted research environment, in which approved researchers analyze sensitive data remotely without being able to download or copy them out.
RECORD statement A reporting guideline that extends STROBE to studies using routinely collected health data, including the reporting of linkage methods and quality (Benchimol et al., 2015).
Key People
Halbert L. Dunn An American physician and vital statistician who used the term record linkage in 1946 to describe assembling the records of a person’s life into a single “book of life”.
Howard B. Newcombe A Canadian scientist who, with colleagues, showed in 1959 that computers could link birth and marriage records automatically by weighing agreement on names and other details, laying the groundwork for probabilistic linkage.
Ivan P. Fellegi A statistician who, with Alan Sunter, published the 1969 theory of probabilistic record linkage and who later served as Chief Statistician of Canada.
Alan B. Sunter A statistician who co-authored with Ivan Fellegi the 1969 paper that set out the mathematical theory of probabilistic record linkage.
Felix Ritchie An economist who developed the Five Safes framework from work on secure access to confidential data at the United Kingdom’s Office for National Statistics.
Latanya Sweeney An American computer scientist whose research on re-identification showed that most of the United States population could be uniquely identified by ZIP code, sex and full date of birth.
Khaled El Emam A Canadian researcher whose work on measuring re-identification risk shaped risk-based approaches to the de-identification of health information.
Katie Harron A statistician whose research on data linkage has described how linkage error arises and how it can bias the results of studies that use linked data.
No matching entries. Try a different search term.