Evidence brief

“More patients” is not a panacea

Additional patients can improve a study's statistical precision. They do not necessarily improve its validity. Once a study is adequately powered, more patients narrow the confidence interval without addressing whether the data observe the exposures, outcomes, and confounders the question requires.

This evidence brief summarizes peer-reviewed findings on data completeness and cohort size, and explains why more patients cannot correct the systematic error that missing data sources can produce.

What the brief covers

Real-world evidence studies draw on data sources that record a patient's exposures, outcomes, and clinical context. Completeness is the share of those expected sources present per patient-year, and it governs whether an estimate is biased rather than merely imprecise.

The brief summarizes evidence that patient count and completeness act on different sources of error, and that raising completeness through stricter source requirements can reduce the analyzable cohort.

  • Why sample size reduces random error but not systematic error
  • Completeness as the presence of expected data sources per patient-year
  • The cohort-size tradeoff observed in the TRUST study, and why it is configuration-specific
  • A four-criterion framework for evaluating whether a cohort is optimized on the right axis
  • The limits of the findings and when the tradeoff should not be generalized

Why this matters

A study can report a large cohort and a narrow confidence interval and still rest on data that omit the sources where its events are documented. For pharmacoepidemiology, that distinction matters because a precise estimate of a biased quantity reads as a confident result.

The practical question is not whether large cohorts are useful. They are. The question is whether patient count is being read as evidence of reliability that only completeness can provide.

Inside the brief

The evidence brief summarizes a large peer-reviewed assessment of real-world data reliability across 120,616 patients with asthma, 58 US hospitals, and more than 1,180 outpatient clinics.

Raising completeness required multi-source linkage, and the higher-completeness cohort was smaller:

46.1%
Completeness, traditional
119,526 patients under a claims-only approach
96.6%
Completeness, advanced
98,372 patients under a multi-source linked approach
18%
Cohort attrition
Patients lost to linkage while completeness roughly doubled

When the events under study depend on the dropped sources, the larger cohort is the more biased one, and its narrow confidence interval reflects only the random error that more patients reduce.

Who should read it

This brief is intended for pharmacoepidemiologists, health economics and outcomes research teams, real-world evidence leaders, epidemiology methods teams, and others evaluating whether a dataset is fit for a specific research use.

It is especially relevant when patient count is being used as the primary measure of a dataset's value.