Chapter 3 Specific Clinical Data Types

3.1 Learning Objectives

Learning Objectives: 1. Physiological (to-do), 2. Monitoring data (to-do), 3. Radiology, 4. Pathology, 5. Understand the potential uses for and risks of synthetic data

3.2 Physiological

3.3 Monitoring data

This section was written by: Jennifer Kelleher, Ph.D.1; Abigail S. Robbertz, Ph.D.1; and Meghan E. McGrady, Ph.D.1,2

NOTE: Jennifer Kelleher, Ph.D.1 and Abigail S. Robbertz, Ph.D.1 contributed equally.

1 Center for Adherence and Self-Management, Division of Behavioral Medicine and Clinical Psychology, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH, USA

2 Department of Pediatrics, University of Cincinnati College of Medicine, Cincinnati, OH, USA

The work discussed in this section was also supported by the National Cancer Institute at the National Institutes of Health (R21CA263704, K07CA200668) to MEM. JK and ASR are supported by the Eunice Kennedy Shriver National Institute of Child Health & Human Development at the National Institutes of Health (T32HD068223). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.

Electronic monitoring devices are digital tools that can be used to track health behaviors over time such as:

  • Sleep
  • Physical activity
  • Medication adherence
  • Calorie intake

Electronic monitoring devices can also be used to assess physical health indicators including: * Blood glucose levels * Blood pressure * Heart rate and heart rate variability * Oxygen saturation

Electronic monitoring devices enable researchers to track day-to-day health behaviors in the patient’s “real-world” setting. This allows researchers to explore patterns or changes in a patient’s health behavior and provides a richer understanding of daily behavior over time.

3.3.1 Benefits of Monitoring Data

  1. Electronic monitoring devices often include data transmission abilities that enable healthcare providers or researchers to access these data in near real-time potentially informing intervention and/or medical decision-making.

  2. Electronic monitoring devices also have the potential to produce more accurate estimates of health behaviors than alternative strategies (e.g., self-report) as they are not subject to recall bias and can detect efforts to inflate adherence due to social desirability.

3.3.2 Considerations

This section is not exhaustive. Research teams are strongly encouraged to consult with experts with experience and training in collecting and analyzing data from specific devices.

To ensure the outcome variables are aligned with the research question of interest and ethical and age/developmental considerations (Psihogios et al. (2024); Modi et al. (2012)) have been appropriately accounted for, readers are encouraged to consult with researchers in their field who have integrated these measurement strategies into their work.

3.3.2.1 Medication Adherence

There are three major components of medical adherence (the tracking of taking medication):

  • Initiation: Starting a prescribed regimen
  • Implementation: The amount of which a patient’s medication-taking behavior corresponds with the treatment regimen or protocol
  • Discontinuation: Stopping a perscribed regimen

For more information see:

3.4 Radiology

3.5 Pathology

3.6 Synthetic Data

Synthetic data are artificially generated datasets designed to reflect the structure and patterns of real-world populations, without containing information about actual patients (Susser et al. 2024). There are many ways to generate synthetic data. Some approaches are simple, like creating partially synthetic data by replacing potentially identifying pieces of information with synthetic values. Others are more complex, like using AI systems to fully generate synthetic patient notes. The goal of synthetic data is to mimic the characteristics and patterns in real patient data, while avoiding privacy concerns, cost, or biases of using real datasets.

3.6.1 Uses of synthetic data

There are several possible uses for synthetic data:

  • Testing or validation: synthetic data may be easier to create or access than real clinical data, while being similar enough to test or validate new tools, methods, or data handling pipelines. It can also be shared between institutions, and used to compare the performance of competing tools. However please see the consideration section about some limitations of this approach.
  • Augmenting or balancing datasets: real patient data can have subgroups with low representation, which can lead to algorithmic bias in some applications. Synthetic data can be used to balance data so that there are more observations that represent these groups.
  • Supplementing controls groups in clinical trials: Conducting clinical trials requires a lot of time and money. Using synthetic data has been suggested to supplement control arms in clinical trials, in order to reduce costs or realign funds to add more patients to the experimental arm (Elvatun et al. 2025).
  • Augmenting data for rare disease areas: Rare disease research often suffers from small sample sizes of patient data and can be especially susceptible to privacy concerns, which can lead to underpowered studies. Synthetic data have the potential to be used for detection and development of treatment for rare diseases (Mendes, Barbar, and Refaie 2025).

Synthetic data has several potential uses, including testing new methods, balancing datasets, supplementing control arms, and augmenting rare disease data.

3.6.2 Resources for generating synthetic data

There are a wide range of ways to generate synthetic data for different uses. Some tools to access or generate synthetic clinical data include:

  • Synthea: a synthetic patient population simulator, that generates synthetic medical histories. This tool can be used without restriction for research, industry, or government (Walonoski et al. 2018).
  • SyntheticMass: a synthetic dataset of residents of Massachusetts (generated using Synthea), which statistically mirrors the real population in terms of demographics, disease burden, and healthcare interaction. This data is free of PII and PHI (Walonoski et al. 2018).
  • simstudy: an R package used to simulate datasets. The user specifies a set of relationships between covariates and type of study, and the tool generates synthetic data (Goldfeld and Wujciak-Jens 2020).
  • synthpop: an R package that takes patient data with sensitive values, and replaces those values with synthetic values, while minimally distorting the statistical summaries of the dataset (Nowok, Raab, and Dibben 2016).

3.6.3 Benefits of synthetic data

Using synthetic data has a variety of potential benefits.

  • Reducing privacy risks: synthetic data are intended to have little or no information about individuals, and therefore reduce privacy risks associated with using clinical data. This also means they can more easily be shared across institutions and without needing to be granted access.
  • Lower cost to generate: synthetic data typically cost less to generate than the cost to collect real patient data, and once an AI system has been trained to generate synthetic data then it can keep producing more. This is particularly useful for training AI systems, which need large sample sizes.

3.6.4 Concerns about synthetic data

Although there is a lot of potential for synthetic data to revolutionize fields that use clinical data, there are also quite a few practical and ethical concerns.

  • Accuracy and reliability: it is challenging to assess how well synthetic data actually mimic real data. Therefore, it is often unknown exactly how reliable statistical tests are based on synthetic data, as opposed to real patient data.
  • Privacy risks: if synthetic data are built from real patient data, there are still threats of identifying patients in some cases, or small subgroups of patients in other cases.
  • Bias: using AI systems in many applications has been shown to lead to algorithmic bias. It is possible for synthetic data to reinforce biases from original patient data or to create new biases.
  • Regulations are still in development: because synthetic data are often not considered to be personally identifiable information (PII) or protected health information (PHI), they are not regulated by the same guidelines as other types of clinical data. Different types of synthetic data have different ethical and privacy risks, and regulations are still in development to govern the use and sharing of these data (Nisevic, Milojevic, and Spajic 2025).
  • Model collapse: when generative models are repeatedly trained on synthetic data rather than real data, the generated data may progressively diverge from true real-world patterns, reducing their utility.

While synthetic data have promise for several areas of clinical research, any synthetic data use needs to be carefully considered and validated.

3.7 Summary

References

Elvatun, Severin, Daan Knoors, Simon Brant, Christian Jonasson, and Jan F Nygård. 2025. “Synthetic Data as External Control Arms in Scarce Single-Arm Clinical Trials.” PLOS Digital Health 4 (1): e0000581.
Goldfeld, Keith, and Jacob Wujciak-Jens. 2020. “Simstudy: Illuminating Research Methods Through Data Generation.” Journal of Open Source Software 5 (54): 2763. https://doi.org/10.21105/joss.02763.
Mendes, Jorge M, Aziz Barbar, and Marwa Refaie. 2025. “Synthetic Data Generation: A Privacy-Preserving Approach to Accelerate Rare Disease Research.” Frontiers in Digital Health 7: 1563991.
Modi, Avani C., Ahna L. Pai, Kevin A. Hommel, Korey K. Hood, Sandra Cortina, Marisa E. Hilliard, Shanna M. Guilfoyle, Wendy N. Gray, and Dennis Drotar. 2012. “Pediatric Self-Management: A Framework for Research, Practice, and Policy.” Pediatrics 129 (2): e473–85. https://doi.org/10.1542/peds.2011-1635.
Nisevic, Maja, Dusko Milojevic, and Daniela Spajic. 2025. “Synthetic Data in Medicine: Legal and Ethical Considerations for Patient Profiling.” Computational and Structural Biotechnology Journal 28: 190–98.
Nowok, Beata, Gillian M Raab, and Chris Dibben. 2016. “Synthpop: Bespoke Creation of Synthetic Data in r.” Journal of Statistical Software 74: 1–26.
Psihogios, Alexandra M., Sara King-Dowling, Jonathan A. Mitchell, Meghan E. McGrady, and Ariel A. Williamson. 2024. “Ethical Considerations in Using Sensors to Remotely Assess Pediatric Health Behaviors.” American Psychologist 79 (1): 39–51. https://doi.org/10.1037/amp0001196.
Susser, Daniel, Daniel S Schiff, Sara Gerke, Laura Y Cabrera, I Glenn Cohen, Megan Doerr, Jordan Harrod, et al. 2024. “Synthetic Health Data: Real Ethical Promise and Peril.” Hastings Center Report 54 (5): 8–13.
Walonoski, Jason, Mark Kramer, Joseph Nichols, Andre Quina, Chris Moesel, Dylan Hall, Carlton Duffett, Kudakwashe Dube, Thomas Gallagher, and Scott McLachlan. 2018. “Synthea: An Approach, Method, and Software Mechanism for Generating Synthetic Patients and the Synthetic Electronic Health Care Record.” Journal of the American Medical Informatics Association 25 (3): 230–38.