Science Education Systems · Article 74. Maya, Jia Jun, Hana and Ethan remain fictional Punggol learners. This article follows the statistical-significance layer: how Science uses a decision threshold without allowing the threshold to replace scientific judgement.
The 50-second parent route
Statistical significance is a technical decision rule, not a certificate of truth.
The route is:
null model → test statistic → sampling distribution → p-value → significance threshold → statistical decision → effect size → confidence interval → validity → replication → scientific meaning
The key question is:
How surprising is this result under the tested null model, and what does that surprise actually justify?
This article extends How Scientific Hypothesis Testing Works, How Scientific Effect Size Works and How Scientific Confidence Intervals Work.
1. Statistical significance is defined relative to a threshold
A p-value is compared with a prespecified significance level such as 0.05.
If the p-value falls below that threshold, the result is called statistically significant under that testing framework.
2. The threshold is a convention, not a natural boundary
p = 0.049 and p = 0.051 are almost identical pieces of numerical evidence.
Scientific meaning should not flip from “real” to “not real” simply because one lies on the opposite side of an arbitrary convention.
3. Maya’s significance error is binary thinking
Significant = true.
not significant = false.
Her repair:
read the p-value together with effect size, interval width and study design.
4. Jia Jun’s significance error is probability reversal
He says p = 0.03 means there is a 3% probability the null hypothesis is true.
His repair:
the p-value conditions on the null; it does not directly calculate the probability of the null.
5. Hana’s significance error is practical-importance confusion
A tiny effect has p < 0.001.
She assumes it must matter greatly.
Her repair:
ask how large the effect is in meaningful units.
6. Ethan’s significance error is threshold shopping
He changes the analysis, subgroup or outcome until something crosses 0.05.
His repair:
prespecify confirmatory tests and report exploratory analyses transparently.
7. A p-value is conditional
It asks:
assuming the null model and statistical assumptions are correct, how unusual would data at least this incompatible with the null be under the chosen test?
The condition matters.
8. A p-value does not measure effect magnitude
Large sample plus tiny effect can yield a small p-value.
Small sample plus meaningful effect can yield a large p-value.
Magnitude and detectability are different.
9. A p-value does not measure replication probability
p = 0.01 does not mean the study has a 99% chance of replicating.
Replication depends on true effect, design, measurement, population, bias and future sampling variation.
10. A p-value does not measure the probability the result was “due to chance”
Chance is built into the null sampling model.
The p-value quantifies how extreme the observed result is under that model.
It does not partition reality into “chance” and “not chance.”
11. A p-value does not prove causation
A highly significant association can still be confounded.
Causation depends on study design and alternative explanations.
12. Internal validity comes before significance
If the treatment and control groups were systematically different, statistical significance can simply detect that bias precisely.
See How Scientific Internal Validity Works.
13. Effect size comes after significance
If evidence is inconsistent with the null, ask:
How large is the difference?
Is it scientifically meaningful?
See How Scientific Effect Size Works.
14. Confidence intervals show compatible effect ranges
A point estimate plus interval often communicates more than a binary significant/not-significant label.
It shows both magnitude and uncertainty.
15. Significance depends on sample size
As sample size increases, standard errors often shrink.
Smaller effects become detectable.
This is scientifically useful but makes practical interpretation essential.
16. Significance depends on variability
Noisy outcomes produce larger uncertainty.
A fixed effect becomes harder to detect.
Better measurement can therefore change significance without changing the underlying phenomenon.
17. Significance depends on the test model
Different assumptions.
different error structures.
different dependence models.
different test statistics.
A p-value is not independent of analytical choices.
18. Significance depends on one-sided versus two-sided testing
A one-sided test can produce a smaller p-value in the prespecified direction.
Changing direction after seeing results is invalid threshold manipulation.
19. Primary Science can learn the deeper idea without p-values
A small difference from one trial may be natural variation.
A large, repeated difference across fair tests is more convincing.
The intuition is signal versus expected noise.
20. Primary 3 can compare repeated versus isolated results
One unusual trial should not dominate five consistent ones automatically.
Students learn to judge patterns rather than single surprises.
21. Primary 4 can connect repetition to confidence
If the same direction appears repeatedly under controlled conditions, the pattern becomes harder to explain by ordinary variation alone.
22. Primary 5 can distinguish “different” from “meaningfully different”
Two temperatures differ by 0.1°C.
If the thermometer resolves only 1°C reliably, the apparent difference may not support a meaningful conclusion.
23. Primary 6 can learn calibrated conclusion language
Evidence supports.
evidence suggests.
evidence is insufficient.
The language should match the strength and precision of the observations.
24. Secondary Science can formalise significance
null hypothesis.
test statistic.
p-value.
significance level.
Type I error.
power.
Students can see where the binary rule fits inside the wider evidence system.
25. The significance level controls a long-run error rate
Under repeated tests where the null model is true, a 5% significance rule is designed to reject about 5% of the time under the method assumptions.
It does not mean 5% of significant findings are necessarily false.
26. False discovery proportion depends on more than alpha
How many tested hypotheses are truly null?
How much power do studies have?
How many analyses are performed?
The proportion of false positives among significant results is not equal automatically to the significance threshold.
27. Base rates matter
If researchers test many implausible hypotheses, a larger fraction of significant findings may be false positives even with conventional error control.
Prior plausibility and research design matter.
28. Significance can be manipulated through repeated testing
Try twenty outcomes.
twenty subgroups.
several transformations.
different stopping points.
The chance of finding something below 0.05 rises.
29. Multiple comparisons require correction or explicit exploratory framing
The next article, How Scientific Multiple Comparisons Work, follows this problem.
30. Selective reporting hides the denominator
One significant result is shown.
Ninety-nine non-significant tests disappear.
The visible p-value no longer describes the full search process honestly.
31. Preregistration helps preserve the denominator
Confirmatory hypotheses can be recorded before data are seen.
Exploratory analyses can still be reported, but readers know which were discovered after searching.
32. Replication is a better filter than threshold worship
A result that crosses p < 0.05 once may disappear later.
Independent replication shows whether the pattern survives new data and new implementation.
33. Meta-analysis can move beyond single-study significance
Several imprecise studies may each be non-significant but jointly estimate a coherent effect.
Conversely, several significant small studies may reveal publication bias or heterogeneity.
34. Dichotomising p-values loses information
p = 0.049 and p = 0.051 are almost the same.
Reporting exact p-values, effect sizes and intervals preserves more information.
35. Extremely small p-values still require validity checks
p < 10⁻¹⁰ does not immunise a study against confounding, measurement bias or data leakage.
Systematic error can produce extremely strong-looking statistical evidence.
36. Extremely large p-values do not prove equality
p = 0.9 can occur because the true effect is tiny.
Or because the data are extremely noisy.
Use confidence intervals and equivalence bounds to distinguish these situations.
37. Equivalence testing can provide evidence of practical similarity
Define a range of effects small enough to be unimportant.
Then test whether the data rule out effects outside that range sufficiently.
This is more informative than simply failing to reject zero.
38. Non-inferiority asks a one-sided practical question
Is the new method no worse than the standard by more than a prespecified acceptable margin?
The margin must be scientifically justified.
39. Significance is useful when embedded in a complete workflow
Prespecified question.
credible design.
appropriate test.
effect size.
interval.
replication.
The threshold then becomes one evidence component rather than the conclusion itself.
40. Statistical significance can be useful for screening
In large experiments, it can help identify results less compatible with null models.
But screening results should be followed by validation, multiplicity control and practical interpretation.
41. AI creates enormous significance-search capacity
Automated systems can test thousands of relationships rapidly.
Without correction and provenance, false discoveries can multiply.
42. AI can p-hack at machine speed if instructed badly
Try transformations.
subgroups.
outcomes.
stopping rules.
Report only the smallest p-value.
This produces persuasive-looking but unreliable evidence.
43. AI can also improve significance literacy
Useful prompts:
“Explain this p-value without saying the null has a probability.”
“Give me two studies with the same effect size but different p-values.”
“Show a significant but practically tiny effect.”
“Show a non-significant result whose interval still includes a large meaningful effect.”
44. AI benchmarks often overfocus on tiny rank differences
Model A = 92.1%.
Model B = 91.9%.
Without uncertainty and repeated evaluation, that ranking may be unstable.
Statistical significance is only one part of comparison quality.
45. Parents can teach significance intuition with repeated coin flips
A fair coin can produce four heads in a row.
Rare-looking sequences happen sometimes.
The right question is whether the observed pattern is sufficiently inconsistent with the fair-coin model across the whole experiment.
46. Small-group tuition can compare evidence packages
Study A:
tiny effect, huge sample, p < 0.001.
Study B:
larger effect, small sample, p = 0.08.
Students decide which questions remain unanswered.
47. A compact statistical-significance checklist
- What null model is being tested?
- What test statistic is used?
- What assumptions does the test require?
- What is the p-value?
- What significance threshold was prespecified?
- How many tests were performed?
- Was the analysis exploratory or confirmatory?
- What effect size was observed?
- What does the confidence interval include?
- Is the study internally valid?
- Is the effect practically meaningful?
- Has the finding replicated?
48. Frequently asked questions
What does statistically significant mean?
It means the p-value fell below a prespecified threshold under the chosen statistical test and null model.
Does statistical significance mean the hypothesis is true?
No. It indicates incompatibility with the tested null model under the assumptions; truth requires broader scientific evidence.
Does p < 0.05 mean there is less than a 5% chance the result is false?
No. The p-value is not the posterior probability that the finding is false.
Can a tiny effect be statistically significant?
Yes. Large samples can detect very small effects.
Can an important effect be non-significant?
Yes. Small or noisy studies can remain too imprecise to detect meaningful effects reliably.
How does significance thinking help Secondary Science?
It teaches students to distinguish detectability from magnitude and to interpret statistical evidence alongside design and uncertainty.
49. Continue the Science Education Systems series
- How Scientific Hypothesis Testing Works
- How Scientific Multiple Comparisons Work
- How Scientific Bayesian Updating Works
Conclusion: Statistical significance is a threshold inside Science, not the throne above it
Maya sees the label.
Jia Jun reads the conditional probability correctly.
Hana checks magnitude and interval.
Ethan asks how many analyses were tried.
Science needs all four.
Use the threshold.
do not worship it.
preserve effect size.
preserve uncertainty.
preserve study validity.
Then let replication decide whether the result deserves to become knowledge.
