“More patients” is not a panacea
Additional patients can improve a study's statistical precision. They do not necessarily improve its validity. Once a study is adequately powered, more patients narrow the confidence interval without addressing whether the data observe the exposures, outcomes, and confounders the question requires.
This evidence brief summarizes peer-reviewed findings on data completeness and cohort size, and explains why more patients cannot correct the systematic error that missing data sources can produce.
What the brief covers
Real-world evidence studies draw on data sources that record a patient's exposures, outcomes, and clinical context. Completeness is the share of those expected sources present per patient-year, and it governs whether an estimate is biased rather than merely imprecise.
The brief summarizes evidence that patient count and completeness act on different sources of error, and that raising completeness through stricter source requirements can reduce the analyzable cohort.
- Why sample size reduces random error but not systematic error
- Completeness as the presence of expected data sources per patient-year
- The cohort-size tradeoff observed in the TRUST study, and why it is configuration-specific
- A four-criterion framework for evaluating whether a cohort is optimized on the right axis
- The limits of the findings and when the tradeoff should not be generalized
Why this matters
A study can report a large cohort and a narrow confidence interval and still rest on data that omit the sources where its events are documented. For pharmacoepidemiology, that distinction matters because a precise estimate of a biased quantity reads as a confident result.
The practical question is not whether large cohorts are useful. They are. The question is whether patient count is being read as evidence of reliability that only completeness can provide.
Who should read it
This brief is intended for pharmacoepidemiologists, health economics and outcomes research teams, real-world evidence leaders, epidemiology methods teams, and others evaluating whether a dataset is fit for a specific research use.
It is especially relevant when patient count is being used as the primary measure of a dataset's value.