In Search of a p < 0.05: The Problem of Multiple Comparisons

Authors

DOI:

https://doi.org/10.51987/Rev.Hosp.Ital.B.Aires.v46i3.1301

Keywords:

statistics as topic, surgical stress, statistical significance, analysis of variance, clinical trials as topic, artificial intelligence

Abstract

This article examines the role of chance and statistics in medical research, focusing on statistical significance and the p-value. Although randomization in clinical trials helps control randomness, the risk of false positive findings persists when multiple comparisons are performed without appropriate statistical adjustment.
This issue is illustrated through an analogy involving a winning casino bet, in which a successful outcome is reported without disclosing previous failed attempts, a situation comparable to uncorrected post hoc analyses. Strategies to mitigate this bias are discussed, including the a priori definition of hypotheses and endpoints, the application of statistical corrections for multiple testing, and the validation of findings in independent cohorts. Finally, the article emphasizes that artificial intelligence is also susceptible to this problem when data are not properly managed.

Downloads

Download data is not yet available.

References

Pocock SJ, Stone GW. The primary outcome fails - what next? N Engl J Med. 2016;375(9):861-870. https://doi.org/10.1056/NEJMra1510064. DOI: https://doi.org/10.1056/NEJMra1510064

Sugitani T, Bretz F, Maurer W. A simple and flexible graphical approach for adaptive group-sequential clinical trials. J Biopharm Stat. 2016;26(2):202-216. https://doi.org/10.1080/10543406.2014.972509. DOI: https://doi.org/10.1080/10543406.2014.972509

Freemantle N, Calvert MJ. Interpreting composite outcomes in trials. BMJ. 2010;341:c3529. https://doi.org/10.1136/bmj.c3529. DOI: https://doi.org/10.1136/bmj.c3529

Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B Stat Methodol. 1995;57(1):289-300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x DOI: https://doi.org/10.1111/j.2517-6161.1995.tb02031.x

Bauer P. Multiple testing in clinical trials. Stat Med. 1991;10(6):871-889; discussion 889-890. https://doi.org/10.1002/sim.4780100609. DOI: https://doi.org/10.1002/sim.4780100609

Perneger TV. What's wrong with Bonferroni adjustments. BMJ. 1998;316(7139):1236-1238. https://doi.org/10.1136/bmj.316.7139.1236. DOI: https://doi.org/10.1136/bmj.316.7139.1236

Pocock SJ, Assmann SE, Enos LE, et al. Subgroup analysis, covariate adjustment and baseline comparisons in clinical trial reporting: current practiceand problems. Statist. Med.2002;21:2917-2930. https://doi.org/10.1002/sim.1296 DOI: https://doi.org/10.1002/sim.1296

Bender R, Lange S. Adjusting for multiple testing--when and how? J Clin Epidemiol. 2001;54(4):343-349. https://doi.org/10.1016/s0895-4356(00)00314-0. DOI: https://doi.org/10.1016/S0895-4356(00)00314-0

Ioannidis JPA. Machine learning-based medical prediction: challenges and opportunities. Lancet Digit Health. 2022;4(8):e540–e548.

Mills JL. Data torturing. N Engl J Med. 1993;329(16):1196-1199. https://doi.org/10.1056/NEJM199310143291613. DOI: https://doi.org/10.1056/NEJM199310143291613

Wasserstein RL, Schirm AL, Lazar NA. Moving to a world beyond “p < 0.05”. Am Stat. 2019;73(Suppl 1):1-19. https://doi.org/10.1080/00031305.2019.1583913. DOI: https://doi.org/10.1080/00031305.2019.1583913

Feynman RP. The pleasure of finding things out. Cambridge (MA): Perseus Books;

Downloads

Published

2026-09-09

Issue

Section

Notes on statistics and research

How to Cite

1.
Knorre ME, Bergier M, Cagide A. In Search of a p < 0.05: The Problem of Multiple Comparisons. Rev Hosp Ital B.Aires [Internet]. 2026 Sep. 9 [cited 2026 Sep. 10];46(3):e0001301. Available from: https://ojs.hospitalitaliano.org.ar/index.php/revistahi/article/view/1301

Most read articles by the same author(s)