In Search of a p < 0.05: The Problem of Multiple Comparisons
DOI:
https://doi.org/10.51987/Rev.Hosp.Ital.B.Aires.v46i3.1301Keywords:
statistics as topic, surgical stress, statistical significance, analysis of variance, clinical trials as topic, artificial intelligenceAbstract
This article examines the role of chance and statistics in medical research, focusing on statistical significance and the p-value. Although randomization in clinical trials helps control randomness, the risk of false positive findings persists when multiple comparisons are performed without appropriate statistical adjustment.
This issue is illustrated through an analogy involving a winning casino bet, in which a successful outcome is reported without disclosing previous failed attempts, a situation comparable to uncorrected post hoc analyses. Strategies to mitigate this bias are discussed, including the a priori definition of hypotheses and endpoints, the application of statistical corrections for multiple testing, and the validation of findings in independent cohorts. Finally, the article emphasizes that artificial intelligence is also susceptible to this problem when data are not properly managed.
Downloads
References
Pocock SJ, Stone GW. The primary outcome fails - what next? N Engl J Med. 2016;375(9):861-870. https://doi.org/10.1056/NEJMra1510064. DOI: https://doi.org/10.1056/NEJMra1510064
Sugitani T, Bretz F, Maurer W. A simple and flexible graphical approach for adaptive group-sequential clinical trials. J Biopharm Stat. 2016;26(2):202-216. https://doi.org/10.1080/10543406.2014.972509. DOI: https://doi.org/10.1080/10543406.2014.972509
Freemantle N, Calvert MJ. Interpreting composite outcomes in trials. BMJ. 2010;341:c3529. https://doi.org/10.1136/bmj.c3529. DOI: https://doi.org/10.1136/bmj.c3529
Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B Stat Methodol. 1995;57(1):289-300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x DOI: https://doi.org/10.1111/j.2517-6161.1995.tb02031.x
Bauer P. Multiple testing in clinical trials. Stat Med. 1991;10(6):871-889; discussion 889-890. https://doi.org/10.1002/sim.4780100609. DOI: https://doi.org/10.1002/sim.4780100609
Perneger TV. What's wrong with Bonferroni adjustments. BMJ. 1998;316(7139):1236-1238. https://doi.org/10.1136/bmj.316.7139.1236. DOI: https://doi.org/10.1136/bmj.316.7139.1236
Pocock SJ, Assmann SE, Enos LE, et al. Subgroup analysis, covariate adjustment and baseline comparisons in clinical trial reporting: current practiceand problems. Statist. Med.2002;21:2917-2930. https://doi.org/10.1002/sim.1296 DOI: https://doi.org/10.1002/sim.1296
Bender R, Lange S. Adjusting for multiple testing--when and how? J Clin Epidemiol. 2001;54(4):343-349. https://doi.org/10.1016/s0895-4356(00)00314-0. DOI: https://doi.org/10.1016/S0895-4356(00)00314-0
Ioannidis JPA. Machine learning-based medical prediction: challenges and opportunities. Lancet Digit Health. 2022;4(8):e540–e548.
Mills JL. Data torturing. N Engl J Med. 1993;329(16):1196-1199. https://doi.org/10.1056/NEJM199310143291613. DOI: https://doi.org/10.1056/NEJM199310143291613
Wasserstein RL, Schirm AL, Lazar NA. Moving to a world beyond “p < 0.05”. Am Stat. 2019;73(Suppl 1):1-19. https://doi.org/10.1080/00031305.2019.1583913. DOI: https://doi.org/10.1080/00031305.2019.1583913
Feynman RP. The pleasure of finding things out. Cambridge (MA): Perseus Books;
Downloads
Published
Issue
Section
License
Copyright (c) 2026 María E. Knorre, Mariano Bergier, Arturo Cagide

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.










