Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

How Scientific Multiple Comparisons Work | When Many Tests Create False Discoveries

Three students in school uniforms work through open books at a classroom table, with textbooks and stationery nearby and study notes on the whiteboard behind them.

Science Education Systems · Article 75. Maya, Jia Jun, Hana and Ethan remain fictional Punggol learners. This article follows the multiple-comparisons layer: how Science protects itself when one dataset is asked too many questions at once.

The 50-second parent route

Test one true null hypothesis at a 5% significance level and a false positive can still occur.

Test many hypotheses and the chance that at least one looks significant by chance increases.

The route is:

scientific question → hypothesis family → number of tests → error criterion → correction → discoveries → effect sizes → uncertainty → independent validation → replication

The key question is:

How many opportunities did we give chance to produce an apparently exciting result?

This article extends How Scientific Hypothesis Testing Works, How Scientific Statistical Significance Works and How Scientific Bias Works.


1. One test and many tests are different scientific situations

If one outcome is prespecified, the error rule applies to that test.

If one hundred outcomes are searched and the smallest p-value is reported, the search process is much larger than the visible result.


2. Multiplicity is the hidden denominator

One significant result out of one test means something different from one significant result out of one thousand tests.

The number and structure of opportunities matter.


3. Maya’s multiplicity error is forgetting the other tests

She reports the only p < 0.05 result from twenty outcomes.

Her repair:

report how many outcomes were tested and apply the planned error-control strategy.


4. Jia Jun’s multiplicity error is correcting everything as one giant family

He applies an extremely severe correction to every exploratory and confirmatory question in an entire research programme.

His repair:

define scientifically meaningful hypothesis families rather than treating unrelated questions as automatically identical.


5. Hana’s multiplicity error is assuming correction makes every discovery true

A corrected significant result can still be biased, confounded or poorly measured.

Her repair:

multiplicity control protects one statistical error pathway, not the whole scientific system.


6. Ethan’s multiplicity error is inventing subgroups after seeing the answer

He splits by age, location, baseline score, time, instrument and dozens of other variables until one subgroup looks dramatic.

His repair:

treat post-hoc subgroup discoveries as exploratory until independently validated.


7. Family-wise error rate asks whether any false rejection occurs in a family

It is a strict error criterion.

The goal is to keep the probability of one or more false positives across the family below a chosen level under the method assumptions.


8. Bonferroni correction is simple and conservative

If m hypotheses form one family, a common Bonferroni rule tests each at alpha/m.

It controls family-wise error without requiring independence, but can reduce power when many tests are included.


9. Holm’s procedure improves on simple Bonferroni in many settings

Sort the p-values.

test them sequentially against adjusted thresholds.

Holm controls family-wise error while often rejecting more false nulls than plain Bonferroni.


10. False discovery rate asks a different question

Instead of trying to avoid any false positive, FDR procedures control the expected proportion of false discoveries among the rejected hypotheses under their assumptions.

This can be useful in high-dimensional discovery work.


11. Benjamini-Hochberg is a common FDR procedure

It orders p-values and compares them with rank-dependent thresholds.

Its purpose differs from family-wise error control:

discover more signals while limiting the expected false-discovery proportion.


12. The right error criterion depends on consequences

Safety-critical confirmatory decision?

Family-wise control may be appropriate.

Large exploratory screen where follow-up validation is expected?

FDR may be more useful.

Scientific purpose determines the trade-off.


13. Primary Science can learn multiplicity through repeated guessing

If a student guesses one hidden card, success is surprising.

If the student gets one hundred guesses, one lucky hit is less impressive.

Opportunity count changes interpretation.


14. Primary 3 can learn the search-space idea

Look for one planned pattern in a table.

Then search freely for anything unusual.

The second task gives chance many more ways to surprise us.


15. Primary 4 can distinguish prediction from discovery

Prediction before data:

stronger confirmatory test.

Pattern noticed afterward:

useful new hypothesis, but it needs another test.


16. Primary 5 can understand why replication matters after exploration

Use one dataset to discover a pattern.

Use a new dataset to test it.

The second dataset protects against lucky pattern searching.


17. Primary 6 can learn selective-reporting danger

If ten experiments were performed but only the successful one is shown, the evidence story is incomplete.

Transparent reporting preserves the denominator.


18. Secondary Science can formalise multiplicity

many outcomes.

many subgroups.

many time points.

many models.

many genes.

many features.

Students can see how repeated testing changes false-positive risk.


19. Multiple outcomes create multiplicity

A study tests:

score.

attendance.

motivation.

sleep.

stress.

Each outcome adds another opportunity for an apparently significant result.


20. Multiple treatment groups create multiplicity

Treatment A versus control.

B versus control.

C versus control.

A versus B.

A versus C.

The comparison family grows quickly.


21. Multiple time points create multiplicity

Week 1.

Week 2.

Week 3.

Week 4.

If the researcher reports only the one week where significance appears, false-positive risk is hidden.


22. Subgroup analyses create multiplicity

Boys.

girls.

high baseline.

low baseline.

each school.

each age.

Large subgroup searches can manufacture apparently special effects.


23. Model specification searching creates multiplicity too

Different covariates.

different transformations.

different outlier rules.

different time windows.

Every analytical choice can become another implicit test.


24. Researcher degrees of freedom are a multiplicity system

If many defensible analyses exist and only the most favourable is reported, the nominal p-value no longer represents the full search process.

Transparency is part of error control.


25. P-hacking is multiplicity without honest accounting

Keep trying analyses until p < 0.05.

Then present the final analysis as if it were the only one.

This inflates false-positive risk.


26. Preregistration reduces hidden multiplicity

Primary outcomes.

main hypothesis.

exclusion rules.

analysis plan.

Recording these before seeing the data distinguishes planned inference from exploration.


27. Exploratory analysis remains valuable

Science needs discovery.

Unexpected patterns can reveal new mechanisms.

The repair is not to ban exploration.

It is to label exploration honestly and validate discoveries independently.


28. Discovery and confirmation should form a loop

Explore → discover → formulate hypothesis → collect new data → confirm or reject → refine.

This converts accidental patterns into testable science.


29. Holdout datasets protect against repeated searching

Use one dataset to develop hypotheses and models.

Keep another untouched for final validation.

The holdout becomes valuable because it has not been searched repeatedly.


30. Reusing the holdout destroys its protection gradually

Test.

adjust model.

test again on the same holdout.

Eventually the holdout becomes part of the training process.

A fresh evaluation set may be needed.


31. Cross-validation manages model selection differently

Data is partitioned repeatedly so models can be assessed across folds.

But extensive model tuning can still overfit the cross-validation process.

Independent final validation remains valuable.


32. Genome-wide studies make multiplicity obvious

Millions of genetic variants may be tested.

Ordinary 0.05 thresholds would generate enormous numbers of chance findings.

Much stricter or FDR-based approaches are needed depending on the goal.


33. Imaging studies also face massive search spaces

Thousands of pixels or regions can be tested.

Spatial dependence and multiple testing must be handled together.


34. High-throughput biology depends on false-discovery control

Genes.

proteins.

metabolites.

microbes.

Large screens are powerful precisely because they test many candidates, which makes multiplicity central.


35. Multiple comparisons and statistical power trade off

Stricter correction reduces false positives but makes true effects harder to detect.

Study design may need larger samples or stronger measurements to preserve power.


36. Multiplicity and effect size should be separated

A corrected p-value says something about error control.

It does not tell us whether the effect is large enough to matter.

Magnitude remains essential.


37. Multiplicity and confidence intervals can be combined

Simultaneous confidence intervals widen ordinary intervals so the collection achieves a chosen coverage property across the family.

The cost of many questions is wider uncertainty.


38. Multiplicity and Bayesian analysis are handled differently

Bayesian hierarchical models can partially pool many related effects and account for the broader structure of the problem.

This does not remove the need for model checking or guard against selective reporting automatically.


39. Hierarchical testing uses scientific structure

Test an overall family first.

Then descend into subgroups or components if evidence supports it.

Tree structure can reduce unnecessary testing while matching the scientific question.


40. Gatekeeping protects primary claims

A trial may define one primary outcome that must succeed before secondary claims are formally tested.

This preserves error control around the most important decision.


41. Multiplicity correction should be planned, not chosen for convenience afterward

Different procedures protect different error criteria.

Choose according to the scientific decision, dependence structure and confirmatory status.


42. Correlated tests complicate simple counting

Ten nearly identical outcomes do not create the same effective search space as ten independent outcomes.

Some procedures account for dependence better than others.


43. Bonferroni remains valid under broad dependence conditions

Its strength is simplicity and strong family-wise control.

Its weakness is conservatism when the family is large or tests are highly correlated.


44. False discovery rate is useful when discoveries will be followed up

Screen broadly.

accept that some flagged candidates may be false.

then validate the candidates independently.

The discovery pipeline can tolerate a controlled proportion of false leads.


45. Independent replication is the final multiplicity firewall

A lucky pattern discovered after thousands of tests is unlikely to reproduce consistently in fresh data if it has no underlying effect.

Replication turns discovery into evidence.


46. Publication bias interacts with multiplicity

Many research teams test many hypotheses.

Positive findings are more likely to be written up and published.

The literature becomes a filtered subset of a much larger testing universe.


47. Meta-analysis inherits the hidden search process

If included studies report only favourable outcomes, pooling published estimates cannot fully recover the missing tests.

Research transparency matters upstream.


48. AI multiplies hypotheses faster than humans ever could

A model can propose thousands of correlations, transformations and subgroups.

That creative power makes rigorous validation more important, not less.


49. Automated data mining needs explicit discovery control

Training set.

validation set.

test set.

prespecified metrics.

correction procedures.

Repeated evaluation without boundaries creates hidden multiplicity.


50. AI benchmarks have multiplicity problems

Many models.

many prompts.

many datasets.

many seeds.

many metrics.

If only the best-looking run is reported, benchmark uncertainty is understated.


51. Leaderboard overfitting is repeated testing against the same benchmark

Teams adapt systems to public benchmark feedback repeatedly.

The benchmark becomes part of the development loop.

Fresh hidden evaluations are needed to estimate real generalisation.


52. AI can help teach multiplicity

Useful prompts:

“Simulate 100 null tests at alpha = 0.05 and count false positives.”

“Compare Bonferroni and false-discovery-rate goals.”

“Create a subgroup-search example where one significant result appears by chance.”

“Design a discovery-and-confirmation workflow.”


53. Parents can teach multiplicity with everyday pattern hunting

A child tries ten study changes in one week and notices one good test score.

Which change caused it?

Too many simultaneous experiments make attribution weak.

Change fewer things and repeat.


54. Small-group tuition can run a false-discovery simulation

Give students many random datasets with no real differences.

Let them test repeatedly.

Some “significant” patterns will appear.

The class sees why multiplicity control exists.


55. A compact multiple-comparisons checklist

  1. How many hypotheses were tested?
  2. Which tests form one scientific family?
  3. Were outcomes and subgroups prespecified?
  4. What error criterion matters: family-wise error or false discovery rate?
  5. Which correction procedure was chosen?
  6. Are tests independent or correlated?
  7. How did correction affect statistical power?
  8. Were exploratory tests labelled honestly?
  9. Were non-significant results reported?
  10. Was a holdout or fresh validation set preserved?
  11. What effect sizes and intervals accompany the discoveries?
  12. Have discoveries replicated independently?

56. Frequently asked questions

What is the multiple-comparisons problem?

It is the increase in false-positive opportunity that occurs when many hypotheses, outcomes, subgroups or analyses are tested and only the most favourable results are emphasised.

What is Bonferroni correction?

It is a simple family-wise error method that divides the desired significance level by the number of tests in the defined family.

What is false discovery rate?

It is an error criterion focused on controlling the expected proportion of false positives among declared discoveries under a specified procedure and assumptions.

Does correction solve p-hacking?

Not automatically. Hidden model choices, selective reporting and unreported analyses can still make the real search process larger than the corrected family.

Is exploration bad Science?

No. Exploration is essential for discovery. The important distinction is to validate exploratory findings on fresh evidence before treating them as confirmed.

How does multiplicity connect to Secondary Science?

It deepens understanding of probability, significance, repeated testing, false positives and why scientific claims need transparent methods and replication.


57. Continue the Science Education Systems series


Conclusion: More questions create more opportunities for chance to answer one convincingly

Maya finds the exciting result.

Jia Jun counts how many tests were tried.

Hana chooses the error criterion.

Ethan saves fresh data for validation.

Science needs all four.

Count the search.

define the family.

control the error.

preserve magnitude and uncertainty.

validate on new evidence.

Then let discovery become knowledge only after it survives outside the search that found it.

Continue from here: Start Here · Tuition · Education · Pathways · Parenting 101 · All Site Routes

eduKate Punggol

Contact

83 Punggol Central, Singapore 828761

edu|Kate Bukit Timah

8 Fourth Avenue, Singapore 268674

By Appointment +65 8823 1234
admin@edukatesg.com

Email Us

When a child finally understands, school becomes less frightening and the future opens wider. Email us for the latest schedules and fees.

← 返回

感谢您的回复。 ✨

了解 eduKate Punggol 的更多信息

立即订阅以继续阅读并访问完整档案。

继续阅读