Chapter 5 How to Acquire Clinical Data

When designing a clinical research study, researchers may collect new clinical data, use data collected by collaborators, or access data already available through their institution. However, it is becoming increasingly common to conduct research using existing clinical data resources. For these studies, a major question is how to acquire these data.

When evaluating a source of clinical data, it is useful to understand why the data were collected, which individuals are included in the data, what variables are available, what requirements, applications, or costs are associated with access.

The major categories of data acquisition sources discussed in this chapter include registries, data commons, cohort studies and biobanks, electronic health record (EHR) platforms, and linked clinical and genomic data resources.

5.1 Learning Objectives

Learning Objectives: 1. List the major categories of clinical data sources, 2. Understand when a researcher may want to use one category of data sources over others, 3. Compare data sources in terms of where the data come from, what data are included, and ease of access

5.2 Registries

A registry is a collection of data about a group of patients with a shared condition or experience. Registries are typically sponsored by government agencies, professional organizations, research institutions, or private companies. Registries have specific guidelines about who will be included, which could be based on the condition of interest, geographic area, or other factors. The data from registries can be used to understand trends in a disease over time and geographic location.

Cancer registries include data about cancer patients. These data are typically collected by hospitals or other medical facilities that diagnose or treat cancer, aggregated by regional registries, and then sent to national cancer registries. These data often include information about patient demographics, tumor characteristics, cancer stage, and initial treatment. Some registries also include information about outcomes after initial diagnosis and treatment.

Cancer registries can be classified as either population-based or hospital-based. Population-based registries (including SEER and NPCR, which you will read about below) collect data about all reported cancer cases within a geographic area. These registries are useful for calculating cancer incidence across time, space, and demographic groups. Hospital-based registries (including NCDB, which you will also read about below) collect data from patients treated in a specific type of cancer program or healthcare system. These registries are useful for improving patient care and studying specific cancer types more closely.

https://www.nih.gov/health-information/nih-clinical-research-trials-you/list-registries https://seer.cancer.gov/registries/cancer_registry/ https://training.seer.cancer.gov/registration/types/hospital.html

5.2.1 US-based cancer registries

In the United States, there are two major organizations that fund cancer registries: the National Cancer Institute (NCI) and the Centers for Disease Control and Prevention (CDC). The NCI oversees the Surveillance, Epidemiology, and End Results (SEER) program, while the CDC oversees the National Program of Cancer Registries (NPCR). Both programs support smaller regional cancer registries across the United States, although typically SEER includes more detailed data for specific geographic regions, and NPCR includes less detailed data for almost the entire country. Data from both of these programs are combined into the U.S. Cancer Statistics (USCS) database. The reason that these data can be combined in a straightforward way is that the North American Association of Central Cancer Registries (NAACCR) has established standards for collecting and reporting cancer registry data in the United States and Canada.

https://training.seer.cancer.gov/operations/standards/setters/naaccr.html

note: add map of overlapping coverage of these two registries!

5.2.1.1 Surveillance, Epidemiology, and End Results (SEER) program

The SEER program has been funded by the NCI since 1973 to collect data about cancer within the US population. SEER combines data about cancer incidence and survival from regional registries. These regional registries cover approximately 45% of the US population, and include all reported cancers diagnosed among residents of the geographic area they cover. The major goal of the SEER program is to gather data to study patterns of cancer incidence and mortality over time and within geographic areas or across demographic subgroups.

https://seer.cancer.gov/about/goals.html

5.2.1.1.1 What population do the data represent?

The SEER program includes data about people diagnosed with cancer while residing in geographic areas covered by SEER registries.

Include this map here: https://seer.cancer.gov/i/seer-map.PNG

5.2.1.1.2 What types of data are included?

SEER registries include the following data:

  • Demographics
  • Cancer characteristics (primary site, morphology, stage)
  • Initial treatment
  • Outcomes (vital status, survival time, cause of death)

https://seer.cancer.gov/data/

5.2.1.1.3 How to access the data

There are two data products available from the SEER program: SEER Research Data and SEER Research Plus and National Childhood Cancer Registry (NCCR) Data.

SEER Research Data includes data from 1975 through the most recent year available, excluding geographic region, month of diagnosis, and other demographic fields. Any user with an email address can access this data by registering here.

SEER Research Plus and NCCR Data includes all SEER Research Data, with the addition of geographic region, month of diagnosis, and other demographic fields that are removed from the SEER Research Data. In order to access this data, the user must have an eRA Commons or Department of Health and Human Services (HHS) account. You will use this account to fill out an application form and sign data use and data limitations agreements. The SEER program will process your request within two business days, and send information about how to access the data and download the accompanying software. See here for more information about requesting access.

Additionally, SEER maintains specialized databases that provide additional variables and link the registry data to other administrative databases such as Medicare. To access these databases, you must already have access to SEER Research Plus and NCCR Data, and then complete an additional application process to request specialized data. Specialized databases are listed here. Two of these databases also require NCI Central IRB approval to access.

5.2.1.2 National Program of Cancer Registries (NPCR)

The NPCR was established by the CDC in 1992 to provide funding and assistance to state cancer registries and to establish standards for collecting and reporting this data. At the same time, states were authorized to create laws governing cancer case reporting and data collection, in order to make cancer registries as complete and comprehensive as possible. NPCR supports cancer registries in 46 states, the District of Columbia, Puerto Rico, and several US territories. Some states receive support from SEER rather than NPCR, and several registry systems contribute data to both programs. Data about new cancer cases are collected at medical facilities, sent to central cancer registries at the state level, and then submitted to NPCR.

https://www.cdc.gov/national-program-cancer-registries/about/index.html

5.2.1.2.1 What population do the data represent?

The NPCR includes data about people who have been diagnosed with cancer or received cancer care from medical providers within any of the geographic areas that are covered by NPCR registries. This includes 46 states, the District of Columbia, Puerto Rico, the US Pacific Island Jurisdictions, and the US Virgin Islands, and represents 97% of the US population.

https://www.cdc.gov/national-program-cancer-registries/about/index.html

5.2.1.2.2 What types of data are included?

Data in the NPCR include:

  • Demographics
  • Cancer characteristics (primary site, histology, grade, behavior, stage)
  • Initial treatment

https://www.cdc.gov/national-program-cancer-registries/data-modernization/index.html

5.2.1.2.3 How to access the data

Data from the NPCR are accessed through the SEER program. Cancer incidence data from the NPCR and SEER programs are combined as the U.S. Cancer Statistics (USCS). The USCS has a public use database with cancer incidence and population data. This data does not include geographic information at the county or zip code level. This data dictionary specifies what data are included in the public use database. To access this data, you must already have access to the SEER Research Plus data (steps to access this data are provided here). Once you have this access, then you can follow the steps outlined here, which include filling out a brief form. Access requests should be processed within 2 business days.

The USCS also has a restricted access database. This database includes more detailed geographic variables, such as county and census tract, along with additional demographic variables, and the data dictionary can be accessed here. This database requires more time and effort to access. The data access application requires the researcher to explain their need for the restricted-use variables and describe their research proposal and plan in detail. The application can be seen here. An important part of this process is identifying the minimal set of restricted use variables that you will need for your project, and explaining why these are necessary for the analysis. If your proposal is approved, you will need to complete a Confidentiality Training, and then make an appointment at a research data center (RDC) or a Federal Statistical Research Data Center (FSRDC) to access the data. An overview of the process for accessing the USCS restricted access database can be found here.

5.2.1.3 National Cancer Database (NCDB)

The NCDB is hospital-based cancer registry. It differs from SEER and NPCR registries because instead of sourcing data from a specific geographic region, it includes data from registries at more than 1,500 hospitals or other medical facilities that are accredited by the Commission on Cancer (CoC). The CoC is a program within the American College of Surgeons, and requires accredited programs to meet a set of standards for care. CoC-accredited programs treat more than 74% of newly diagnosed cancer patients. NCDB data contain more detailed information about the facility in which a patient is treated than SEER and NPCR data, and includes longitudinal follow-up and survival information not generally available in the public-use USCS database. Database access is only available to researchers associated with CoC-accredited cancer programs.

https://www.facs.org/quality-programs/cancer-programs/commission-on-cancer/coc-accreditation/

5.2.1.3.1 What population do the data represent?

The NCDB includes data about patients who have been treated by CoC accredited cancer programs, collected from hospital registries.

5.2.1.3.2 What types of data are included?

Data in the NCDB include:

  • Demographics (including education and income for patient’s zip code)
  • Cancer characteristics (primary site, histology, grade, stage)
  • Treatment
  • Outcomes (30 day mortality, 60 day mortality, last contact or death, vital status)

https://www.facs.org/media/ilqb5snq/2024-data-dictionary.pdf

5.2.1.3.3 How to access the data

The NCDB manages a dataset called the Participant User Data File (PUF) which includes information about cancer cases from CoC-accredited hospitals and cancer programs. The PUFs are available only to researchers associated with CoC-accredited cancer programs who submit an application. This application includes letters of support from principal investigators (PIs) in the CoC-accredited program and a description of your proposed research project and analysis plan.

5.2.2 International registries

Many countries outside of the United States also maintain comprehensive cancer registries.

The Nordic Cancer Registries include population-based cancer data from several Nordic countries, which have some of the oldest and most comprehensive cancer registries in the world. NORDCAN provides free and public aggregated incidence, mortality, prevalence, and survival data at the level of country, year, cancer entity, sex, and age group. Access to individual level data from any individual registry within the Nordic Cancer Registries requires a request to that specific registry, and typically requires a local collaborator or affiliation.

Another comprehensive cancer registry is the National Cancer Registration and Analysis Service (NCRAS) in England. These registry data are linked with information on treatments, outcomes, and healthcare utilization. These data can only be accessed for health care purposes, and can be requested through this process.

While many countries have cancer registries, only one in three countries worldwide currently report high quality cancer incidence data. The Global Initiative for Cancer Registry Development is an initiative led by the World Health Organization. Regional hubs help individual countries build up their capacity to collect, analyze, and report data from cancer registries. More information about these regional hubs and individual country’s registry programs can be found here.

Include this image here: https://gicr.iarc.fr/about-the-gicr/the-value-of-cancer-data/map_world_registrystatus.pdf

https://nordcan.iarc.fr/en https://digital.nhs.uk/ndrs/about/ncras https://gicr.iarc.fr

5.2.3 Summary

Provide summary table to compare these resources! (maybe just the US resources because these are the ones with detail)

5.3 Data commons

A data commons is a cloud-based platform that enables a research community to store, manage, analyze, and share data while controlling access through a common governance framework. In addition to hosting data, a data commons typically provides computational resources, software tools, and services that allow researchers to work with data within the same environment. Data commons also often harmonize data from multiple sources, allowing datasets collected by different organizations or studies to be analyzed together. A primary goal of a data commons is to promote data sharing and reuse by reducing barriers to accessing data and computation. Data commons often aim to reduce the gap between the large volume of data that has been collected and the smaller subset of data that can be readily accessed and reused for research. By centralizing data, infrastructure, and governance, data commons can help research communities collaborate more efficiently and share the costs associated with data storage and analysis.

The major features of a data commons are:

  1. A data commons stores data and computing tools and computation resources
  2. A data commons facilitates analysis within the platform
  3. A data commons contains multiple harmonized datasets
  4. A data commons is typically designed for collaboration and data reuse

Several cancer-focused data commons exist, and are excellent sources of existing cancer data for researchers.

https://www.nature.com/articles/s41597-023-02029-x

5.3.1 Cancer Research Data Commons (CRDC)

The CRDC is a data commons run by the National Cancer Institute (NCI) to accelerate cancer research. The CRDC consists of seven specialized data commons, as well as a cloud infrastructure and other computing resources. The seven data commons included in the CRDC are:

  • Genomic Data Commons (GDC): DNA methylation and whole genome, whole exome, RNA-seq, miRNA-seq, and ATAC-seq data
  • Proteomic Data Commons (PCD): mass-spectrometry-based proteomic data
  • Imaging Data Commons (IDC): de-identified radiology and pathology data
  • Integrated Canine Data Commons (ICDC): genomic and clinical data from canine patients with cancer
  • Clinical and Translational Data Commons (CTDC): clinical, biospecimen, and molecular characterization data from NCI-funded studies
  • Population Science Data Commons (PSDC): population studies data
  • General Commons (GC): data that do not fit into other CRDC data commons

As well as including these component data commons, the CRDC provides a cloud resource with access to NCI-funded data, hundreds of publicly available tools and workflows, and computational resources for working with large-scale data.

While each data commons within the CRDC is heavily utilized for cancer research, the most popular and heavily cited component is the GDC. The IDC is also heavily used for imaging data. More information about the other data commons, as well as instructions for data access can be found here.

https://www.cancer.gov/about-nci/organization/cbiit/projects/crdc

5.3.1.1 Genomic Data Commons (GDC)

The GDC contains genomic data related to cancer, as well as tools and software to analyze this data. This data commons is a central repository for data from several major NCI studies of genomic changes in cancer, including The Cancer Genome Atlas (TCGA) and the Therapeutically Applicable Research to Generate Effective Treatment (TARGET) program.

5.3.1.1.1 What population do the data represent?

The GDC includes genomic data contributed by many NIH-funded cancer studies and external studies. A partial list of data sources can be found here.

5.3.1.1.2 What types of data are included?

The GDC includes the following data:

  • Genomic and molecular data
    • DNA sequencing
    • Copy number variation
    • Structural variation
    • DNA methylation
    • RNA expression
    • Protein expression
  • Clinical metadata (diagnosis, age, sex, etc.)

https://gdc.cancer.gov/about-data

5.3.1.1.3 How to access the data

The GDC data portal allows researchers to download datasets or to analyze data within the GDC platform. Data within the GDC are classified as either open access or controlled access.

Open access data do not require any authentication or authorization to access. This includes genomic data that are not individually identifiable, and most clinical and all biospecimen data elements. The data can be accessed through the Genomic Data Commons Data Portal.

Controlled access data include individually identifiable data and some clinical data elements. Accessing controlled access data requires first obtaining an eRA Commons account and access to the database of Genotypes and Phenotypes (dbGaP), then obtaining access to the specific controlled access research project through dbGaP, and finally logging into the GDC Data Portal to access. The eRA Commons lets federal agencies manage research grants, and accounts can only be made for individuals at research organizations by officials from the organization (https://www.era.nih.gov/register-accounts/create-and-edit-an-account.htm). The dbGaP is commonly used to manage controlled-access data, enabling researchers to request datasets and data access committees to review requests (https://grants.nih.gov/policy-and-compliance/policy-topics/sharing-policies/accessing-data/dbgap). Only senior researchers can request data through dbGaP, but once they have obtained access to project data, they can provide access to other members of their lab with GDC accounts (https://gdc.cancer.gov/access-data/data-access-processes-and-tools). Senior researchers must be permanent employees of their institutions who are either academic professors or researchers with responsibilities that include laboratory or research program administration. Laboratory staff and trainees are not eligible to submit data access requests in dbGaP, but can take part in data analysis for projects overseen by a senior researcher. (https://grants.nih.gov/policy-and-compliance/policy-topics/sharing-policies/accessing-data/dbgap)

5.3.1.1.4 The Cancer Genome Atlas (TCGA)

TCGA was a major cancer genomics project led by the NCI and the National Human Genome Research Institute which characterized over 20,000 samples across 33 types of cancers.

5.3.1.1.4.1 What population do the data represent?

TCGA data come from samples collected at a network of tissue source sites, primarily located at cancer centers and hospitals in the United States, with additional contributions from a smaller number of international institutions (https://gdc.cancer.gov/resources-tcga-users/tcga-code-tables/tissue-source-site-codes). Samples were collected from patients with a wide range of cancer types.

5.3.1.1.4.2 What types of data are included?

The TCGA includes the following data:

  • Demographics
  • Genomic and molecular data
    • Genomic
    • Epigenomic
    • Transcriptomic
    • Proteomic
  • Outcomes
5.3.1.1.4.3 How to access the data

Most of the data from TCGA is part of the open access data stored in the GDC, and accessed through the GDC Data Portal. Individual-level genomic data can be accessed through the controlled access tier of the GDC, described above.

5.3.1.1.5 Therapeutically Applicable Research to Generate Effective Treatments (TARGET)

The TARGET initiative was another major genomic program led by the NIH and NCI to generate genomic data in order to understand the molecular basis of childhood cancers. The cancer types primarily studied in this initiative include Acute Lymphoblastic Leukemia, Acute Myeloid Leukemia, Wilms Tumor, Neuroblastoma, and Osteosarcoma (https://www.cancer.gov/ccg/research/genome-sequencing/target/studied-cancers).

5.3.1.1.5.1 What population do the data represent?

TARGET data were collected from children or young adults with one of the cancer types listed above, mainly gathered by clinicians at member institutes of the Children’s Oncology Group (COG) (https://www.cancer.gov/ccg/research/genome-sequencing/target/about).

5.3.1.1.5.2 What types of data are included?

TARGET includes the following data:

  • Demographics
  • Genomic and molecular data
    • Genomic
    • Epigenomic
    • Transcriptomic
    • Proteomic
  • Outcomes
5.3.1.1.5.3 How to access the data

Most of the data from TARGET is part of the open access data stored in the GDC, and accessed through the GDC Data Portal. Individual-level genomic data can be accessed through the controlled access tier of the GDC, described above.

5.3.1.2 Imaging Data Commons (IDC)

The IDC is a data commons for publicly available cancer imaging data. It provides more than 85 TB of freely accessible imaging data. The images and image-derived data are harmonized to use a standard digital image format. The IDC includes data from many sources, including The Cancer Genome Atlas (TCGA), The Cancer Imaging Archive (TCIA), the Human Tumor Atlas Network (HTAN), and many others. The data can be explored with the IDC portal, visualized with a browser-based viewer, and downloaded.

5.3.1.2.1 What population do the data represent?

The IDC includes data from many sources, including The Cancer Genome Atlas (TCGA), The Cancer Imaging Archive (TCIA), the Human Tumor Atlas Network (HTAN), and many others. A full list can be found here.

5.3.1.2.2 What types of data are included?

The IDC includes the following data:

  • Imaging
    • Radiology images
    • Digital pathology images
    • Multispectral microscopy images
  • Image-derived features
    • Annotations
    • Parametric maps
    • Measurements
    • Expert assessments

https://datacommons.cancer.gov/repository/imaging-data-commons

5.3.1.2.3 How to access the data

The data can be freely explored with the IDC portal, visualized with a browser-based viewer, and downloaded. The IDC also maintains a Python package for programatically querying metadata and downloading imaging data.

5.3.2 Other data commons

5.3.2.1 International Cancer Genome Consortium (ICGC)

Similarly to TCGA, the ICGC aimed to collect genomic data for all major tumor types. As an international consortium, they also aimed to facilitate global genomic data sharing. The group, which involved 86 teams and spanned nearly every continent, produced genomic data from more than 20,000 tumors across 26 cancer types.

https://www.icgc-argo.org/page/65/icgc-initiatives-

5.3.2.1.1 What population do the data represent?

ICGC data come from samples that were taken by consortium members.

5.3.2.1.2 What types of data are included?

The ICGC includes the following data:

  • Demographics
  • Genomic and molecular data
    • Genomic
    • Epigenomic
    • Transcriptomic
    • Proteomic
  • Clinical data
  • Outcomes

https://docs.cancergenomicscloud.org/docs/icgc-data

5.3.2.1.3 How to access the data

While the ICGC data portal has been shut down, the data can be accessed following the steps here.

The ICGC includes open access release data, which can be freely downloaded using these instructions.

The ICGC also includes controlled release data. Access to this data requires an application.

5.3.2.2 Pediatric Cancer Data Commons (PCDC)

The PCDC is a data commons run by the Data for the Common Good institution at the University of Chicago. The data commons includes pediatric, young adult, and adult cancer data, as well as an analysis platform with tools to explore available data and assess study feasibility. The data come from more than 40 countries and encompass most types of pediatric cancer. The PCDC standardizes and harmonizes the data, which supports defining cohorts based on diseases and demographics across studies.

https://commons.cri.uchicago.edu/pcdc/

5.3.2.2.1 What population do the data represent?

The data in the PCDC come from clinical trials related to the following diseases: rhabdomyosarcoma, non-rhabdomyosarcoma, neuroblastoma, germ cell tumors, Hodgkin lymphoma, and acute myeloid leukemia.

https://docs.pedscommons.org/DataAccessAndGovernance/

5.3.2.2.2 What types of data are included?

The PCDC includes the following data:

  • Demographics
  • Cancer characteristics
  • Treatment
  • Laboratory data
  • Genomic data
  • Outcomes
5.3.2.2.3 How to access the data

PCDC data can be accessed through their data portal. A data portal account requires an authenticated email and institutional affiliation. Open access data allow researchers to explore data dictionaries and build cohorts. They do not include participant level data. Researchers are expected to use this open access data for determining feasibility of a proposed study and for early hypothesis exploration, not for publication.

Scientists who would like to do research on PCDC data need to request access to a dataset associated with a specific disease, and their request will be reviewed. Project request forms can be found here.

5.3.2.3 Childhood Cancer Clinical Data Commons (C3DC)

The C3DC is a database of demographic and clinical data related to pediatric cancers, supported by the NCI. The data in the C3DC have been harmonized and include a standard set of data elements. This data commons allows researchers to search for harmonized participant-level clinical data across multiple studies and create custom cohorts to analyze.

5.3.2.3.1 What population do the data represent?

Data that are included in the C3DC come from forty studies, listed here, and include over 59,000 participants.

5.3.2.3.2 What types of data are included?

The C3DC includes the following data:

  • Demographics
  • Cancer characteristics
  • Treatment
  • Laboratory data
  • Genomic data
  • Outcomes

https://clinicalcommons.ccdi.cancer.gov/data_model

5.3.2.3.3 How to access the data

The data can be explored with this web portal, and cohorts can be compared with this web portal. The C3DC only includes deidentified open-access clinical data.

5.3.3 Summary

Provide summary table to compare these resources!

5.4 Cohorts and Biobanks

Cohort studies are observational studies in which a group of participants is enrolled and followed over time. Researchers collect information about participants at baseline and then track outcomes such as disease diagnoses, health behaviors, or mortality. By observing how health changes over time, cohort studies can be used to identify factors associated with disease risk, progression, and outcomes.

Biobanks are organizations that collect, store, manage, and distribute biological samples, such as blood, saliva, tissue, or DNA, for biomedical research. These samples are typically linked to information about the participants who provided them, including demographic characteristics, health records, environmental exposures, and survey responses. Biobanks may be operated by hospitals, research institutions, private companies, patient advocacy organizations, or government agencies.

Many modern research initiatives combine these two approaches. Participants are enrolled into a cohort study, contribute biological samples to a biobank, and are then followed over time through surveys, physical examinations, electronic health records, or other data sources. This combination allows researchers to connect biological measurements with health outcomes and disease development. Two well-known examples are the UK Biobank and the NIH All of Us Research Program.

Biobanks may be population-based, meaning they recruit participants from the general population, or disease-oriented, meaning they focus on individuals with a particular disease. Although the resources discussed in this section are not cancer-specific, they contain rich clinical, genomic, and lifestyle data that are frequently used in cancer research.

https://allofus.nih.gov/article/biobank https://www.broadinstitute.org/what-is-a-biobank https://pmc.ncbi.nlm.nih.gov/articles/PMC8275637/#CR15

5.4.1 UK Biobank

The UK Biobank is a large population-based cohort study and biobank that includes data from more than 500,000 middle-aged participants in the United Kingdom. Participants were recruited between 2006 and 2010 and contributed biological samples, physical measurements, and detailed information about their health and lifestyle. The study links these data with healthcare records and continues to follow participants over time.

5.4.1.1 What population do the data represent?

Over 500,000 volunteers were recruited for the UK Biobank between 2006 and 2010. Participants lived within 25 miles of the Biobank’s 22 research centers in Scotland, England, and Wales, and had ages between 40 and 69 when recruited. The participants on average were slightly healthier and wealthier than the general UK population. Approximately 95% of the participants were white, and there were relatively low numbers of participants from different ethnic groups.

5.4.1.2 What types of data are included?

The UK Biobank includes the following data:

  • Demographics and lifestyle
  • Survey/questionnaire data
  • Physical measurements
  • Biomarker data
  • Genomic data
  • Imaging data
  • EHR data
  • Environmental exposures

5.4.1.3 How to access the data

Access to the data from the UK Biobank requires researchers to fill out an application and pay an access fee. Eligible researchers must have a track record of health-related research and be affiliated with a recognized research organization. The application requires a summary of the intended research, and the process is outlined here. The fee structure for accessing this data is provided here.

5.4.2 All of Us

All of Us is a nationwide research program launched by the NIH in 2015 with the goal of building a large and diverse resource for precision medicine research. Like the UK Biobank, All of Us combines a longitudinal cohort study with a biobank. Participants contribute surveys, physical measurements, electronic health record data, and biological samples, and many participants consent to long-term follow-up. The program places particular emphasis on recruiting populations that have historically been underrepresented in biomedical research. Biological samples are processed and stored through a national biobank managed by the Mayo Clinic. Although All of Us is not focused specifically on cancer, its scale and breadth of clinical and biological data make it a valuable resource for many types of cancer research.

https://www.joinallofus.org/about https://www.joinallofus.org/faq https://pmc.ncbi.nlm.nih.gov/articles/PMC9436122/

5.4.2.1 What population do the data represent?

All of Us includes over 800,000 participants across the US. The program aims to include populations that are historically under-represented in medical research.

5.4.2.2 What types of data are included?

All of Us includes the following data. While the survey responses are collected on all participants, the other types of data are collected only on a subset of participants:

  • Demographics and lifestyle
  • Survey/questionnaire data
  • Physical measurements
  • Biomarker data
  • Genomic data
  • Imaging data
  • EHR data
  • Unstructured clinical notes
  • Wearable device data
  • Cognitive assessment data

https://support.researchallofus.org/hc/en-us/articles/4619151535508-Data-Types-and-Organization

5.4.2.3 How to access the data

The All of Us dataset has three access tiers. The Public Tier includes deidentified and aggregated data. These data are available through the data browser. The Registered Tier includes individual-level data, including data from EHR, wearables, surveys, and physical measurements. The Controlled Tier includes genomic data, demographic fields in EHR that are suppressed in other tiers, and accurate dates of events that are shifted in other tiers.

In order to access the Registered and Controlled tiers, your institution must have signed a Data Use and Registration Agreement with All of Us. You can check here whether your institution has an agreement in place for the tier that you would like to access, or submit a request if not. If your institution does have an agreement, then you will need to create an account and complete a training.

https://www.researchallofus.org/data-tools/data-access/

5.4.3 Summary

Provide summary of this section (may not need table because we just have the two cohort studies/biobanks and they are quite similar)!

5.5 EHR platforms

Electronic health records (EHR) contain digital information generated through a patient’s interactions with the healthcare system over time, including demographics, diagnoses, medications, laboratory tests, clinical notes, images, and other clinical observations. Unlike registries and cohort studies, in which data are collected primarily for research purposes, EHR are primarily collected and stored to support clinical care, healthcare billing, and health system operations. Although research is not their primary purpose, they provide a very useful longitudinal real world data source for clinical researchers. However, these records may be incomplete, missing relevant variables, or recorded differently across health care systems and institutions. Instead of accessing EHR directly from healthcare systems, researchers typically access data through EHR platforms. These platforms aggregate EHR data from healthcare systems, standardize and harmonize the data, and create research databases. Part of this process involves mapping original data into common data models, which provide a standardized way to represent diagnoses, medications, laboratory results, healthcare encounters, and other clinical information. EHR platforms differ in their funding, major goals, cost to access, and data scope. EHR platforms are increasingly used for cancer research because they can capture detailed treatment histories, lab results, medications, and outcomes that may not be available in traditional cancer registries.

5.5.1 Research EHR platforms

5.5.1.1 PCORnet

PCORnet is a health data resource funded by the Patient Centered Outcomes Research Institute (PCORI). PCORnet receives electronic health data from a large network of health systems and standardizes data under a shared data model. PCORnet includes data from more than 50 million patients from across the US. PCORnet provides both data resources and study support services for researchers. PCORnet will freely provide a consultation about using PCORnet for a study, a study feasibility review, and potentially a data network request, which involves determining what available data in PCORnet would relate to a proposed study and providing site-level aggregate data. If applicable, PCORnet can support data queries, although these will have associated costs. Unlike many centralized databases, PCORnet uses a distributed research network. Participating institutions retain control of their local patient data behind institutional firewalls. Researchers submit standardized queries across the network, and participating sites return approved results.

https://pcornet.org/data/common-data-model/ https://pcornet.org/module-3-playbook/

5.5.1.1.1 What population do the data represent?

Data from PCORnet comes from electronic health records of more than 50 million patients in the healthcare system. These data primarily come from partnerships with PCORnet Clinical Research Networks (CRNs), which are groups of healthcare institutes across the US. A list of the CRNs and their associated healthcare institutions can be found here.

5.5.1.1.2 What types of data are included?

PCORnet includes EHR data, including:

  • Demographics
  • Healthcare encounters
  • Diagnoses
  • Procedures
  • Medications
  • Laboratory data
  • Vital signs
  • Immunizations
  • Cause of death

https://pcornet.org/wp-content/uploads/2025/05/PCORnet_Common_Data_Model_v70_2025_05_01.pdf

5.5.1.1.3 How to access the data

PCORnet provides several free resources to streamline research before getting to paid data queries. First, you can do a consultation with the PCORnet Front Door to access PCORnet. Next, you can request a study feasibility review, to understand what PCORnet data are available for a project and to develop a budget for a project using PCORnet data. After this, you can submit a data network request, in which you will work with PCORnet partners to determine where relevant data exist across the PCORnet network, and to provide high-level data trends for study planning. Finally, if you end up using PCORnet data in your study, you will work with PCORnet CRNs, which will require funding. More information about these resources can be found here.

https://pcornet.org/module-3-playbook/

5.5.1.2 ENACT

ENACT is an EHR platform that facilitates research with EHR for researchers at participating Clinical and Translational Science Award (CTSA) institutions. A list of participating institutions can be found here. ENACT includes over 150 million patient records, including records from 90% of the CTSA consortium. ENACT is primarily intended for cohort discovery and study feasibility assessment. Researchers can estimate how many patients meeting specific eligibility criteria exist across participating institutions before launching a clinical study.

https://enact-network.org/about/faqs/ https://enact-network.org

5.5.1.2.1 What population do the data represent?

ENACT includes over 150 million patient records, including records from 90% of the CTSA consortium. A list of participating institutions can be found here.

5.5.1.2.2 What types of data are included?

ENACT data includes the following:

  • Demographics
  • Healthcare encounters
  • Diagnoses
  • Procedures
  • Medications
  • Laboratory data

https://enact-network.org/media/rb3fsitp/act_ontology_data_dictionary.pdf

5.5.1.2.3 How to access the data

To request access to ENACT, you must be a member of a participating CTSA institution. If you are, you can reach out to the contact at your local institution, which you can find here.

5.5.1.3 Epic Cosmos

Epic Cosmos is an EHR platform that includes data from health systems using Epic software, and is available to researchers affiliated with and approved by a Cosmos participating organization. A list of organizations can be found here. If you are not part of a participating organization, you can search for a collaborator at a participating organization or submit ideas for a study here.

5.5.1.3.1 What population do the data represent?

Epic Cosmos includes EHR data from patients that are part of health systems that use Epic software. Epic reports that the demographic composition of Cosmos is similar to that of the U.S. population, although it only includes patients receiving care through participating Epic health systems (https://cosmos.epic.com/about/).

5.5.1.3.2 What types of data are included?

Epic Cosmos includes the following data:

  • Demographics
  • Diagnoses
  • Medications
  • Vital signs
  • Patient-generated health data
  • Birth records
  • Social determinants of health

https://cosmos.epic.com/about/

5.5.1.3.3 How to access the data

If you are affiliated with a Cosmos participating organization, you can request access here. You will need to submit a form with your contact information and a description of your inquiry.

5.5.2 Commercial EHR platforms

5.5.2.1 TriNetX

TriNetX is a commercial EHR platform that provides data for pharma companies, healthcare providers, and academic researchers. It connects researchers to harmonized patient data from participating healthcare organizations that have mapped their data to a common data model. TriNetX sources data directly from healthcare organizations, from more than 14,000 clinical sites. TriNetX is a federated EHR network, which means that patient data is queried within healthcare systems and answers are returned without the patient-level data leaving the healthcare system. A major use of TriNetX is to assess the feasibility of proposed clinical studies.

https://trinetx.com/data/ https://trinetx.com/solutions/trinetx-live/

5.5.2.1.1 What population do the data represent?

TriNetX includes records from more than 65 million patients, across 20 countries. These data come from a network of healthcare organizations across the world.

https://trinetx.com/data/

5.5.2.1.2 What types of data are included?

TriNetX data includes the following:

  • Demographics
  • Diagnoses
  • Procedures
  • Medications
  • Laboratory data
  • Vital signs
  • Outcomes
5.5.2.1.3 How to access the data

Data can be accessed by contacting TriNetX here.

5.5.2.2 Truveta

Truveta is another commercial EHR platform, with more than 130 million patient records. The data are standardized across healthcare systems and mapped to a common data model. Truveta was created by and is governed by a collective of US health systems, including the systems listed here. Data are delivered daily from these health systems to the EHR platform.

5.5.2.2.1 What population do the data represent?

Truveta includes records from more than 130 million patients, which are sourced from more than 900 hospitals and 20,000 clinics. Some affiliated health systems are listed here.

https://www.truveta.com/data/

5.5.2.2.2 What types of data are included?

Truveta data includes the following:

  • Demographics
  • Diagnoses
  • Procedures
  • Medications
  • Laboratory data
  • Immunizations
  • Clinical notes
  • Imaging data

https://www.truveta.com/data/

5.5.2.2.3 How to access the data

Data can be accessed by contacting Truveta here.

5.5.3 Summary

Provide summary table comparing these resources!

5.6 Linked Clinical and Genomic Data Resources

A major limitation of most registries and EHR platforms is a lack of detailed tumor molecular data. Clinico-genomic data platforms link clinical data with tumor sequencing data. Both clinical data and sequencing data are collected as part of the healthcare process. By linking clinical information with sequencing results, these resources allow researchers to study how genomic alterations relate to treatment response and clinical outcomes across cancer types.

https://pmc.ncbi.nlm.nih.gov/articles/PMC11869879/

5.6.1 Public platforms

5.6.1.1 Project GENIE

Project GENIE is a publicly accessible cancer-focused clinico-genomic data resource that combines genomic sequencing data and clinical outcomes, funded by the American Association for Cancer Research (AACR). A major goal of Project GENIE is to collect a large repository of sequencing and clinical cancer data for precision oncology research. Precision medicine customizes a patient’s treatment based on their individual genomic, behavioral, and environmental characteristics, and requires a large amount of data. Project GENIE includes clinical and cancer genome sequencing data from most cancer patients treated at twenty international cancer centers. Through partnerships with Sage Bionetworks and cBioPortal, the project aggregates and harmonizes the data and provides it freely for research purposes. Project GENIE also has a research collaboration with the Biopharma Collaborative (BPC), a coalition of ten biopharmaceutical companies, in order to gather additional clinical data to the existing Project GENIE data.

https://www.aacr.org/professionals/research/aacr-project-genie/ https://www.aacr.org/professionals/research/aacr-project-genie/bpc/

5.6.1.1.1 What population do the data represent?

The data in Project GENIE are collected from cancer patients from twenty international cancer centers. A list of participating centers can be found here.

5.6.1.1.2 What types of data are included?

Project GENIE includes the following data:

  • Demographics
  • Cancer characteristics (stage, histology, etc.)
  • Treatment
  • Genomic data
  • Outcomes

https://www.aacr.org/professionals/research/aacr-project-genie/aacr-project-genie-data-model/ https://www.aacr.org/wp-content/uploads/2026/02/19.0_public_data_guide.pdf

5.6.1.1.3 How to access the data

There are multiple ways to access the Project GENIE data. Users can either access the data directly through cBioPortal, or download the data from Synapse.

cBioPortal is a platform for basic analysis and visualization of cancer genomic data. It provides access to many publicly available cancer datasets (a list of available datasets can be found here). To access Project GENIE data on cBioPortal, you must make a free account and then request access here.

Synapse is a platform for data sharing. To access Project GENIE data with Synapse, follow these steps. This involves creating a Synapse account, agreeing to terms of use, submitting a request for access for each desired dataset, and downloading the data using R, Python, or a command line client.

5.6.1.2 MSK-CHORD

Memorial Sloan Kettering Cancer Center (MSK) curated the MSK-CHORD dataset, which combines deidentified longitudinal clinical data with genomic data from genomic profiling tests of tumors. This dataset includes cases of many common and rare cancer types from over 24,000 patients. The dataset includes natural language processing annotations of unstructured data, along with structured clinical data and genomic data. MSK-CHORD is available through cBioPortal.

https://www.mskcc.org/research-advantage/support/digital-health-projects/msk-chord-clinico-genomic-data-msk-cancer https://www.nature.com/articles/s41586-024-08167-5

5.6.1.2.1 What population do the data represent?

MSK-CHORD includes data from patients who received the MSK-IMPACT tumor-sequencing test at MSK.

https://www.mskcc.org/msk-impact

5.6.1.2.2 What types of data are included?

MSK-CHORD includes the following data. Some of the clinical annotations are derived from natural language processing.

  • Demographics
  • Cancer characteristics
  • Treatment
  • Genomic data

https://www.cbioportal.org/study/summary?id=msk_chord_2024

5.6.1.2.3 How to access the data

MSK-CHORD data can be accessed through cBioPortal.

5.6.2 Commercial platforms

5.6.2.1 Flatiron

Flatiron health is a commercial platform that contains oncology data, as well as additional resources. Flatiron includes more than five million patient records, which are linked with imaging, genomic, pathology, and claims data. Flatiron health has partnered with Foundation Medicine (FMI), a precision medicine company, to create their Clinico-Genomic Database (CGDB), which combines patient-level clinical data with genomic data from genomic profiling tests.

https://flatiron.com/real-world-evidence/solutions-for-clinical-development-teams https://jamanetwork.com/journals/jama/fullarticle/2730114

5.6.2.1.1 What population do the data represent?

Flatiron collects oncology EHR data from patients receiving care at more than 220 cancer practices. Approximately 70% of patients come from community cancer centers and 30% of patients come from academic medical centers. The CGDB includes over 110,000 patients.

https://flatiron.com/database-characterization

5.6.2.1.2 What types of data are included?

Flatiron data include the following:

  • Demographics
  • Cancer characteristics
  • Treatment
  • Clinical notes and unstructured data (processed with natural language processing, machine learning, and large language models)

The CGDB additionally links these records to genomic profiling data.

https://flatiron.com/database-characterization https://resources.flatiron.com/clinico-genomic-data-is-a-game-changer-for-precision-oncology/

5.6.2.1.3 How to access the data

Flatiron data can be accessed by contacting them here.

5.6.2.2 Tempus

Tempus is a healthcare technology company that maintains a large clinico-genomic database and provides genomic sequencing services. Tempus collects EHR data from healthcare systems, normalizes and harmonizes the data, links it with molecular genomic data, and uses AI agents and human experts to process unstructured clinical data.

https://www.tempus.com/research/ https://www.tempus.com/tech-blog/The-Tempus-data-pipeline-architecting-the-future-of-precision-medicine/#overview

5.6.2.2.2 What types of data are included?

Tempus includes the following data:

  • Clinical data
  • Clinical notes (unstructured)
  • Genomic and molecular data
    • DNA sequencing
    • RNA sequencing
  • Imaging data
  • Claims data
  • Outcomes

https://www.tempus.com/solutions/real-world-data/

5.6.2.2.3 How to access the data

Tempus data are available to participating institutions. Tempus can be contacted here.

5.6.3 Summary

Provide summary table comparing these resources

5.7 Summary

Provide overall summary of different data sources!

Note that more data resources can be found in this table.