Science Education Systems · Article 69. Maya, Jia Jun, Hana and Ethan remain fictional Punggol learners. This article follows the internal-validity layer: how Science decides whether the effect observed inside a study is genuinely caused by the factor being tested rather than by bias, confounding or procedural failure.
The 50-second parent route
Before asking whether a result travels, ask whether the study earned the result in the first place.
The route is:
question → causal contrast → assignment → control → blinding → measurement → attrition check → analysis → alternative explanation audit → internal-validity judgement
The key question is:
Did the study isolate the causal effect it claims to have measured?
This article extends How Scientific Randomisation Works, How Scientific Blinding Works, How Scientific Confounding Works and How Scientific External Validity Works.
1. Internal validity is about causal credibility inside the study
If Group A improves more than Group B, can the difference be attributed to the tested intervention?
Or could something else explain it?
Internal validity is the discipline of answering that question.
2. A strong result needs a strong counterfactual
What would have happened to the treated group if treatment had not occurred?
We cannot observe that directly.
A well-designed control group approximates it.
3. Maya’s internal-validity error is before-and-after certainty
Scores improve after an intervention.
She assumes the intervention caused the improvement.
Her repair:
ask what would have happened over the same period without the intervention.
4. Jia Jun’s internal-validity error is control-group worship
A control group exists, so he assumes the study must be valid.
His repair:
check whether the groups were comparable, treated similarly apart from the intervention, and measured consistently.
5. Hana’s internal-validity error is perfectionism
One design limitation appears.
She declares the study useless.
Her repair:
judge how much the limitation plausibly changes the causal estimate.
6. Ethan’s internal-validity error is statistical rescue
He believes a sufficiently complex model can repair any weak design.
His repair:
recognise that analysis cannot fully reconstruct experimental information that was never collected.
7. Randomisation protects baseline comparability
Random assignment reduces systematic differences between treatment groups before the intervention begins.
It is one of the strongest internal-validity tools.
8. Controls protect the causal contrast
A valid reference condition helps distinguish intervention effects from time, maturation, handling, background trends and other changes.
9. Blinding protects behaviour and measurement
Participants, investigators, outcome assessors or analysts may behave differently if they know the assignment.
Blinding reduces expectation pathways where feasible.
10. Standardisation protects procedure
Same instructions.
same timing.
same equipment.
same outcome definitions.
Unequal procedures can create artificial treatment differences.
11. Measurement quality is internal validity
If the outcome is measured badly, the study cannot identify the causal effect accurately even if assignment is perfect.
12. Misclassification can hide or manufacture effects
Participants are placed into the wrong exposure category.
Outcomes are recorded inconsistently.
Measurement error changes the estimated difference.
13. Differential measurement is especially dangerous
If one group is measured more carefully than another, systematic bias can appear.
Measurement procedures should not depend on treatment status unless the scientific question requires it.
14. Attrition threatens internal validity
Participants drop out.
If dropout differs by treatment and outcome risk, the remaining groups may no longer be comparable.
15. Missing outcomes are not merely smaller sample size
Who is missing matters.
If the weakest performers disappear disproportionately from one group, the observed result can look artificially strong.
16. Protocol deviations can change the treatment contrast
Wrong dose.
missed sessions.
crossovers.
extra support.
What was assigned and what was received may differ.
17. Intention-to-treat protects the original randomised contrast
Participants are analysed according to the group assigned initially.
This preserves key benefits of randomisation when estimating the effect of assignment.
18. Per-protocol analyses answer different questions
They focus on participants who adhered sufficiently to the intervention.
This may estimate treatment under adherence but can reintroduce selection bias.
19. Contamination reduces contrast
The control group receives part of the treatment.
The treatment group adopts control behaviours.
The difference between groups shrinks or becomes harder to interpret.
20. Co-interventions can distort the result
Treatment Group A receives the intervention plus extra coaching.
Control Group B does not.
Which component caused the outcome?
Internal validity requires separation.
21. History threatens before-after designs
An external event happens during the study.
Policy changes.
weather shifts.
another programme begins.
The outcome changes for reasons unrelated to the tested intervention.
22. Maturation threatens longitudinal intervention studies
Children develop.
patients recover naturally.
organisms age.
Systems change even without treatment.
23. Regression to the mean threatens extreme-group studies
Select the very lowest scorers.
Some improve on the next test simply because the first score contained random downward fluctuation.
A control group helps distinguish this effect.
24. Instrumentation changes can create false effects
Before treatment, one instrument is used.
After treatment, a more sensitive instrument is used.
The apparent improvement may partly be measurement-system change.
25. Testing effects can change participants
Repeated tests create practice.
Repeated surveys change awareness.
The act of measurement can influence later outcomes.
26. Selection bias can weaken comparability before analysis
If treatment access depends on baseline prognosis, groups begin differently.
Randomisation or careful design is needed to prevent confounding by selection.
27. Confounding threatens internal validity in observational studies
A third variable influences exposure and outcome.
The measured association differs from the true causal effect.
See How Scientific Confounding Works.
28. Time order protects causal logic
The cause should precede the effect.
If outcome measurement occurs before exposure, the causal story requires rethinking.
29. Reverse causation can mimic treatment effects
Outcome state influences exposure rather than exposure causing outcome.
Longitudinal timing and design help distinguish the directions.
30. Primary Science can learn internal validity through fair tests
Change one factor.
keep comparison conditions similar.
measure the same outcome.
repeat.
Fair-test logic is the foundation.
31. Primary 3 can identify unfair comparisons
One plant gets more light and more water.
The child should reject the simple causal conclusion.
32. Primary 4 can inspect baseline equality
Were the objects or organisms similar enough before the test began?
Starting differences matter.
33. Primary 5 can inspect measurement consistency
Was the same ruler used?
same measuring point?
same timing?
Method consistency protects the comparison.
34. Primary 6 can critique alternative explanations
What else changed?
Could natural variation explain the difference?
Would another control help?
Students begin thinking like causal auditors.
35. Secondary Science can formalise validity threats
selection.
history.
maturation.
instrumentation.
testing.
attrition.
confounding.
Students can map where a causal study might fail.
36. Statistical significance does not repair weak internal validity
A tiny p-value from a biased study can still support the wrong conclusion.
Statistical certainty about a flawed comparison is not causal certainty.
37. Large samples do not repair systematic design bias automatically
More data narrows random uncertainty.
It can also make a biased estimate look precisely wrong.
38. Effect size should be interpreted only after validity is considered
A large observed difference is impressive only if the study design makes that difference causally meaningful.
The next article, How Scientific Effect Size Works, follows this layer.
39. Statistical power is also downstream of validity
A highly powered study can detect tiny systematic errors.
Power answers whether a real effect can be detected, not whether the study measured the right causal contrast.
40. Confidence intervals quantify sampling uncertainty, not every bias
A narrow interval can surround a biased estimate.
Unmeasured systematic error often lies outside the statistical interval.
41. Internal validity and reproducibility are different
A flawed analysis can be perfectly reproducible.
Reproducibility makes the path visible.
Internal validity judges whether the path supports the causal claim.
42. Internal validity and replication are connected
If independent studies with strong internal validity produce compatible effects, confidence increases.
Replication protects against one local implementation failure.
43. Internal validity and triangulation are connected
Different study designs carry different threats.
If a causal conclusion survives methods with distinct weaknesses, the robust core becomes stronger.
44. Internal and external validity can trade off
Tight laboratory control can increase internal validity while reducing realism.
Pragmatic field settings increase realism while introducing more uncontrolled variation.
Science often needs both kinds of study.
45. The best research programme separates the questions
First:
Can the effect be identified under strong control?
Then:
Does the effect survive realistic conditions?
Internal and external validity become sequential tests.
46. Internal validity matters in education research
A new teaching method appears to improve scores.
But did the treatment group receive a stronger teacher?
more time?
more practice?
different marking?
The intervention must be isolated.
47. Internal validity matters in AI evaluation
Model A appears better than Model B.
But did Model A receive more tools, longer context, easier prompts or more human editing?
Comparison conditions define the causal claim.
48. A/B tests can have internal-validity failures
Users see both variants.
assignment changes mid-test.
metrics are redefined.
bots contaminate traffic.
Randomisation alone does not protect every layer.
49. AI can help audit internal validity
Useful prompts:
“List alternative explanations for this treatment effect.”
“Identify threats from selection, history, maturation, attrition and measurement.”
“Which variable should have been controlled?”
“What control group best represents the counterfactual?”
50. AI can also create false confidence
A model can generate a polished causal explanation without noticing that the study was observational, unblinded or badly confounded.
Method comes before narrative.
51. Parents can teach internal validity through household experiments
Try a new study routine.
Do not simultaneously change sleep, tuition time, subject difficulty and reward system if you want to know what caused improvement.
One clear contrast teaches the principle.
52. Small-group tuition can run validity audits
Give three experimental designs.
One has poor randomisation.
one has biased measurement.
one has differential dropout.
Students identify the threat and propose a repair.
53. A compact internal-validity checklist
- What causal effect is being estimated?
- What is the comparison condition?
- Were groups comparable at baseline?
- Was assignment protected from bias?
- Were procedures standardised?
- Was measurement comparable across groups?
- Was blinding used where appropriate?
- Did attrition differ?
- Could contamination or co-interventions occur?
- Could history or maturation explain the change?
- Could confounding or reverse causation remain?
- Does the analysis preserve the design?
- What alternative explanation is still plausible?
54. Frequently asked questions
What is internal validity?
Internal validity is the degree to which a study supports the claimed causal relationship within the population and conditions actually studied.
What threatens internal validity?
Selection bias, confounding, history, maturation, poor controls, differential measurement, attrition, contamination and analytical errors can all weaken it.
Does randomisation guarantee internal validity?
No. It strongly protects baseline comparability, but measurement, attrition, contamination and analysis can still fail.
Does a narrow confidence interval prove high internal validity?
No. Confidence intervals mainly quantify sampling uncertainty under a statistical model; systematic design bias can remain.
How does internal validity help PSLE Science?
It deepens fair-test reasoning by asking whether the observed difference really comes from the changed variable rather than another difference in the setup.
How does internal validity change in Secondary Science?
Students can evaluate randomisation, controls, blinding, attrition, confounding, measurement quality and alternative causal explanations more formally.
55. Continue the Science Education Systems series
- How Scientific Effect Size Works
- How Scientific Statistical Power Works
- How Scientific Confidence Intervals Work
Conclusion: Internal validity is the study’s first promise
Maya sees the improvement.
Jia Jun checks the comparison.
Hana audits the threats.
Ethan asks what alternative cause remains.
Science needs all four.
Build the causal contrast.
protect assignment.
control the procedure.
measure consistently.
track attrition.
challenge every alternative explanation.
Only then ask how large the effect is or how far it travels.

