Science Education Systems · Article 73. Maya, Jia Jun, Hana and Ethan remain fictional Punggol learners. This article follows the hypothesis-testing layer: how Science converts an idea into a prediction that evidence can challenge.
The 50-second parent route
A hypothesis is not merely a guess.
It is a proposed explanation or relationship that should generate testable consequences.
The route is:
question → hypothesis → null model → alternative model → prediction → study design → data → test statistic → p-value or other evidence measure → effect size → uncertainty → scientific judgement
The key question is:
What evidence should we expect if this hypothesis is wrong, and what evidence would make us revise it?
This article extends How Science Inquiry Works, How Scientific Falsification Works, How Scientific Effect Size Works and How Scientific Confidence Intervals Work.
1. A scientific hypothesis makes a claim about the world
More light increases growth rate within a specified range.
A material conducts heat faster than another under matched conditions.
A treatment changes an outcome relative to control.
The claim must be clear enough to confront evidence.
2. Testability requires observable consequences
If the hypothesis is true, what should happen?
If it is false, what pattern should appear instead?
Without discriminating predictions, the hypothesis cannot be tested strongly.
3. Maya’s hypothesis error is vague prediction
She says:
“Plant A will do better.”
Her repair:
define the outcome, time period, comparison and direction precisely.
4. Jia Jun’s hypothesis error is data-first storytelling
He studies the graph, notices a pattern, then writes a hypothesis as if it had been predicted in advance.
His repair:
distinguish exploratory discovery from confirmatory testing.
5. Hana’s hypothesis error is treating failure to reject as proof
The test is not statistically significant.
She says the null hypothesis has been proven true.
Her repair:
ask what effect sizes remain compatible with the data and how precise the study was.
6. Ethan’s hypothesis error is treating rejection as final truth
A small p-value appears.
He declares the alternative hypothesis proven forever.
His repair:
combine statistical evidence with effect magnitude, design validity, replication and mechanism.
7. Null hypotheses provide reference models
A common null model says there is no difference, no association or no effect of a specified kind.
The observed data is compared against what this reference model predicts.
8. The null is not always literally “nothing happens”
It may specify a particular parameter value, equality, boundary or model.
Good testing starts by stating exactly what the null model claims.
9. Alternative hypotheses describe competing possibilities
Difference exists.
association is positive.
treatment exceeds a meaningful threshold.
The alternative should match the scientific question.
10. One-sided and two-sided tests answer different questions
A two-sided test considers departures in both directions.
A one-sided test concentrates error probability in one prespecified direction.
The direction should be justified before inspecting the outcome.
11. The study design comes before the test
Randomisation.
control.
sampling.
measurement.
blinding.
No statistical test can repair a fundamentally invalid comparison automatically.
12. A test statistic compresses evidence into a comparison quantity
Difference in means.
standardised difference.
correlation.
likelihood ratio.
Different tests use different statistics suited to the model.
13. The sampling distribution defines what counts as surprising
Under the null model, repeated comparable studies would produce a range of test statistics by chance.
The observed statistic is located within that distribution.
14. A p-value measures compatibility with the null model
It is the probability, assuming the null model and test assumptions are correct, of obtaining a test statistic at least as extreme as the observed one according to the chosen test.
It is not the probability that the null hypothesis is true.
15. Statistical thresholds are decision conventions
A threshold such as 0.05 is commonly used in many fields.
It is not a universal law of nature.
The scientific meaning of evidence should not jump magically at one decimal boundary.
16. Type I error is false rejection under the null
If the null model is true, a testing procedure can still reject it by chance.
The significance level controls that long-run false-positive rate under the test assumptions.
17. Type II error is missed detection of a specified alternative
A real effect exists, but the study fails to reject the null.
This risk depends on effect size, sample size, variability and test design.
18. Statistical power is one minus Type II error for a specified scenario
Power planning asks whether the study can detect effects worth caring about.
See How Scientific Statistical Power Works.
19. Primary Science already uses hypothesis logic
Prediction:
“If this material is a better conductor, the temperature should change faster under the same conditions.”
Then test.
The statistical machinery can come later.
20. Primary 3 can separate question and prediction
Question:
Which material keeps water warm longest?
Prediction:
The material with lower heat transfer will show the smallest temperature drop over the chosen interval.
21. Primary 4 can identify evidence that would contradict a prediction
If the predicted group repeatedly performs worse, the learner should revise the hypothesis rather than rewrite the observation.
Disconfirmation is part of learning.
22. Primary 5 can compare competing hypotheses
Why did one plant grow faster?
More light?
more water?
different starting health?
Design the next test to discriminate among explanations.
23. Primary 6 can connect hypotheses to variables and controls
A useful hypothesis identifies what should change, what should be measured and what conditions must remain comparable.
24. Secondary Science can formalise hypothesis testing
null and alternative hypotheses.
test statistics.
significance levels.
Type I and Type II errors.
confidence intervals.
Students begin linking experimental design to statistical decision rules.
25. Hypothesis tests should match the data structure
Independent groups.
paired measurements.
counts.
continuous outcomes.
clustered observations.
Different structures require different models.
26. Independence is an assumption in many common tests
Ten readings from one sensor are not always equivalent to ten independent sensors.
Thirty students in one classroom are not fully independent observations of school-level effects.
Dependence changes uncertainty.
27. Distributional assumptions matter
Some tests rely on approximate normality, equal variances or large-sample behaviour.
Strong analysis checks whether those assumptions are reasonable enough for the purpose.
28. Nonparametric tests provide alternatives in some settings
They may rely on fewer or different distributional assumptions.
But they are not assumption-free.
Every method has a model.
29. Exact tests can be useful in small samples
When asymptotic approximations are weak, exact or permutation-based methods can use the study’s combinatorial or randomisation structure more directly.
30. Randomisation tests connect design and inference
Reassign labels according to the original randomisation rule.
Ask how often a statistic as extreme as the observed one appears.
The experimental design becomes the probability engine.
31. Confidence intervals often communicate more than binary testing
Reject or fail to reject is one decision.
An interval shows which effect magnitudes remain compatible with the data.
See How Scientific Confidence Intervals Work.
32. Effect size should accompany hypothesis testing
A tiny effect can become statistically detectable with enough data.
The magnitude determines whether the result matters scientifically.
33. Practical thresholds can be more meaningful than zero
Instead of testing whether an effect differs from exactly zero, ask whether it exceeds a meaningful engineering, clinical or educational threshold.
The scientific decision becomes better aligned with real consequence.
34. Equivalence testing reverses the usual emphasis
Define a range of effects small enough to be practically unimportant.
Then test whether the data support effects lying inside that range.
This can provide evidence of practical similarity.
35. Non-inferiority asks whether something is not unacceptably worse
A new method may be cheaper or safer.
The question becomes whether its performance falls below a prespecified acceptable margin.
The margin should be scientifically justified.
36. Hypothesis testing is vulnerable to optional stopping
Collect data.
check significance.
continue only if necessary.
stop immediately when the threshold is crossed.
Naive repeated peeking changes false-positive behaviour.
37. Sequential methods can handle repeated looks correctly
Planned sequential designs adjust decision boundaries so evidence can be monitored while controlling error rates appropriately.
38. Multiple hypotheses create another problem
Test enough null hypotheses and some will appear significant by chance.
The next article on multiple comparisons follows this layer.
39. Selective reporting distorts the testing record
Ten outcomes tested.
one significant.
only that one reported.
The reader sees an inflated impression of evidence.
40. Preregistration can separate planned from exploratory tests
Recording the main hypothesis and analysis before seeing outcomes reduces ambiguity about which tests were confirmatory.
Exploratory work remains valuable when labelled honestly.
41. Replication is a second test of the hypothesis
An effect that appears once may reflect chance or local bias.
Independent replication tests whether the predicted pattern returns.
42. Falsification is deeper than one significance test
A scientific theory can generate many predictions.
One failed test may expose a boundary, auxiliary assumption or measurement problem.
Model revision requires scientific diagnosis, not mechanical rejection.
43. Hypothesis tests do not rank all explanations automatically
Rejecting one null does not prove one specific alternative if several alternatives predict the same result.
Discriminating tests should separate competing explanations.
44. Likelihood compares how well models explain the observed data
Instead of asking only whether one null is surprising, likelihood methods can compare relative support among parameter values or models.
This creates a bridge toward Bayesian updating.
45. Bayesian tests answer different inferential questions
Prior information and the likelihood are combined to update probability over hypotheses or parameters.
The series will treat that framework separately rather than collapsing it into frequentist testing.
46. AI can generate hypotheses quickly
But idea generation is not evidence.
A fluent model can propose dozens of plausible relationships.
Each requires scientific testing.
47. AI creates a multiple-hypothesis explosion
If a model proposes hundreds of correlations and humans report only the interesting ones, false discoveries become likely.
Automation increases the need for error control.
48. AI can help design discriminating tests
Useful prompts:
“Give me three hypotheses that explain this pattern.”
“For each, predict one outcome the others do not.”
“Design a fair test that separates them.”
“List assumptions behind the statistical test.”
49. AI can misinterpret p-values confidently
Common errors include saying a p-value is the probability the null is true or the probability the result happened by chance.
Readers should restore the conditional logic explicitly.
50. Parents can teach hypothesis testing through prediction before observation
Before opening the oven:
Which loaf should have risen more?
Why?
What result would make us rethink that explanation?
Prediction before observation reduces hindsight storytelling.
51. Small-group tuition can run hypothesis tournaments
Give one surprising result.
Each student proposes a different explanation.
Then design one experiment whose outcomes distinguish among them.
The strongest hypothesis is the one that survives hard tests, not the one told most confidently.
52. A compact hypothesis-testing checklist
- What precise scientific question is being asked?
- What is the hypothesis?
- What null or reference model is used?
- What alternative is scientifically relevant?
- What prediction distinguishes the models?
- Was the test planned before inspecting outcomes?
- Does the design support the comparison?
- Are assumptions reasonable?
- What effect size was observed?
- How uncertain is the estimate?
- What error rates or multiplicity issues apply?
- What result would genuinely change the model?
53. Frequently asked questions
What is hypothesis testing?
It is a framework for comparing observed data with predictions from a reference hypothesis or model using a prespecified statistical procedure.
What is a null hypothesis?
It is the reference claim against which the observed data is evaluated, often specifying no effect or a particular parameter value.
Does failing to reject the null prove it?
No. The study may simply lack enough information to distinguish the null from meaningful alternatives.
Does rejecting the null prove the alternative?
No. It shows the data are difficult to reconcile with the tested null under the model assumptions; scientific interpretation still depends on design, effect size, alternatives and replication.
How does hypothesis logic help PSLE Science?
It strengthens prediction, fair-test design, evidence evaluation and willingness to revise explanations when results contradict expectations.
How does it change in Secondary Science?
Students can connect hypotheses to variables, statistical models, uncertainty, significance, power and competing explanations more formally.
54. Continue the Science Education Systems series
- How Scientific Statistical Significance Works
- How Scientific Multiple Comparisons Work
- How Scientific Bayesian Updating Works
Conclusion: A hypothesis earns strength by surviving predictions it could have failed
Maya proposes the explanation.
Jia Jun specifies the prediction.
Hana checks the uncertainty.
Ethan designs the test that could prove them wrong.
Science needs all four.
State the claim.
define the reference.
predict before observing.
test fairly.
measure magnitude and uncertainty.
Then revise the model according to the evidence rather than the preference.
