← Charlie Yan

Working paper · 2026

How Much Tail Prediction Could We Have Detected? A Power Accounting for a 1,450-Test Search

A null from a large signal search is uninterpretable unless the searcher reports what the search could have detected.

Headline result At n_eff ≈ 116 and Bonferroni over 1,450 tests, 80% power requires |IC| ≥ 0.44. The best candidates sat at 0.21–0.23.

Reporting “we found nothing” after 1,450 tests says nothing on its own. What it could have found is the number that makes the null readable.

This supplies that accounting, and proposes a tail-IC/mean-IC ≥ 2 pre-registration gate to remove premium-confound false positives. Event classification, on roughly 15 episodes, requires AUC ≥ 0.828 to clear the same bar.

Reproducibility

Fully reproducible, no data needed. python code/p3_power_table.py regenerates the detectability table analytically — Fisher-z for IC, Hanley–McNeil for AUC.

multiple testing · statistical power · pre-registration