Science Education Systems · Article 78. Maya, Jia Jun, Hana and Ethan remain fictional Punggol learners. This article follows the meta-analysis layer: how Science combines effect estimates from multiple studies without pretending every study is identical.
The 50-second parent route
A meta-analysis is not a vote count.
It does not ask how many studies were “positive.”
It asks how large the effects were, how uncertain each estimate was, how much the studies differ, and what pooled quantity is scientifically meaningful.
The route is:
systematic evidence base → comparable effect measures → study weights → pooled estimate → confidence interval → heterogeneity → publication-bias checks → sensitivity analysis → interpretation → update
The key question is:
What does the combined evidence estimate, and what important differences would a single pooled number hide?
This article extends How Scientific Systematic Review Works, How Scientific Effect Size Works and How Scientific Confidence Intervals Work.
1. Meta-analysis begins after study selection
The evidence base should already have been searched systematically and screened using explicit criteria.
Statistical pooling cannot repair an incomplete search.
2. Studies must estimate sufficiently compatible quantities
Same broad outcome.
same direction of comparison.
similar scientific question.
If studies estimate fundamentally different things, pooling can create a meaningless average.
3. Maya’s meta-analysis error is vote counting
Five significant studies and three non-significant studies.
She declares “five wins to three.”
Her repair:
use effect estimates and uncertainty rather than significance labels.
4. Jia Jun’s meta-analysis error is equal weighting
He gives a ten-person study and a ten-thousand-person study identical influence.
His repair:
weight studies according to their statistical information under the chosen model.
5. Hana’s meta-analysis error is pooled-number worship
She sees one combined effect and stops reading.
Her repair:
inspect heterogeneity, study design, population differences and prediction intervals.
6. Ethan’s meta-analysis error is pooling everything
Different outcomes, doses, species and designs are forced into one number.
His repair:
ask whether the combined estimand has a coherent scientific meaning.
7. Effect measures must be aligned
Mean difference.
standardised mean difference.
risk ratio.
odds ratio.
hazard ratio.
correlation.
Different measures require appropriate transformations before combination.
8. Direction must be consistent
If positive values mean benefit in one study and harm in another, signs must be harmonised before pooling.
A simple coding mistake can reverse the synthesis.
9. Study weights reflect precision
More precise estimates generally receive more weight than less precise estimates.
Under common inverse-variance approaches, smaller variance means larger weight.
10. Weight is not the same as scientific quality
A huge biased study can be very precise.
Statistical weight should not erase risk-of-bias judgement.
11. Fixed-effect models make a strong common-effect assumption
One underlying true effect is assumed for all included studies, and observed differences arise from sampling variation.
This can be reasonable in narrow settings but inappropriate when real effects differ across populations or implementations.
12. Random-effects models allow true effects to vary
Each study can estimate a different underlying effect drawn from a distribution of effects.
The pooled estimate then describes an average across that distribution.
13. Random effects do not solve heterogeneity
They model it statistically.
Scientists should still ask why effects differ.
14. Heterogeneity is scientific information
Different doses.
different populations.
different instruments.
different follow-up times.
Variation can reveal effect modifiers and boundary conditions.
15. Cochran-style heterogeneity tests ask whether variation exceeds sampling expectations
But such tests depend on the number and precision of studies.
A non-significant heterogeneity test does not prove studies are identical.
16. I-squared summarises relative heterogeneity
It describes the proportion of observed variation associated with between-study heterogeneity rather than within-study sampling error under the model.
It should not be interpreted as a universal measure of scientific importance.
17. Tau-squared estimates between-study variance
This quantity sits on the scale of the meta-analytic model and helps describe how much true study effects vary.
18. Prediction intervals answer an external-validity question
Instead of asking only where the average effect lies, a prediction interval estimates a range in which a future study’s true effect may plausibly fall under the model.
This can be more useful when heterogeneity is substantial.
19. Forest plots make the evidence visible
Each study shows:
effect estimate.
confidence interval.
weight.
The pooled estimate appears separately.
A good forest plot lets the reader see agreement and disagreement simultaneously.
20. Meta-analysis should not erase study identity
Every estimate should remain traceable to its source, design and population.
Provenance prevents the pooled number from becoming detached from the evidence that created it.
21. Primary Science can learn meta-analytic thinking informally
Three class groups measure the same effect.
Do not simply average every group if one used a much less precise method.
Ask which measurements carry more information.
22. Primary 3 can compare repeated group results
Group A sees a strong difference.
Group B sees a moderate difference.
Group C sees almost none.
The class asks whether the pattern is broadly consistent.
23. Primary 4 can compare ranges as well as averages
One group’s estimate is precise.
another varies widely.
The same centre does not mean the same certainty.
24. Primary 5 can recognise outlying studies
One class result differs sharply from the others.
Check method, sample and measurement before deleting it.
25. Primary 6 can ask whether studies are similar enough to combine
Same material?
same temperature?
same outcome?
same time interval?
Scientific comparability comes before arithmetic pooling.
26. Secondary Science can formalise meta-analysis
effect estimates.
standard errors.
weights.
fixed-effect and random-effects models.
heterogeneity.
forest plots.
Students can see evidence synthesis as quantitative modelling.
27. Publication bias threatens meta-analysis
If studies with null or unfavourable findings are less likely to appear, the pooled result can be biased upward.
28. Small-study effects can create suspicious patterns
Smaller studies may show larger effects because of publication bias, design differences or genuine population differences.
The pattern deserves investigation rather than automatic correction.
29. Funnel plots are one diagnostic
Effect size is plotted against a measure of study precision.
Asymmetry can suggest publication bias or small-study effects, but many other mechanisms can create asymmetry.
30. Statistical tests for funnel asymmetry are not proof
Low study counts reduce reliability.
heterogeneity can distort patterns.
Diagnostics should be interpreted with context.
31. Sensitivity analysis tests robustness
Remove high-risk studies.
change the statistical model.
change effect measure.
exclude extreme assumptions.
If the conclusion changes dramatically, fragility should be reported.
32. Leave-one-out analysis checks study influence
Repeat the meta-analysis while omitting each study in turn.
If one study controls the entire conclusion, the evidence base is less stable than the pooled sample size suggests.
33. Subgroup meta-analysis can explore effect modification
Age group.
dose.
setting.
study design.
But many subgroup searches create multiplicity and should be interpreted cautiously.
34. Meta-regression models study-level moderators
Effect size is related to study characteristics.
This can generate hypotheses about heterogeneity.
Because the number of studies is often small and study-level variables can be confounded, conclusions require caution.
35. Ecological bias can affect meta-regression
A study-level average may not represent individual-level relationships.
Associations among studies should not be transferred automatically to individuals.
36. Individual participant data meta-analysis can go deeper
Instead of combining only published summary estimates, researchers may obtain participant-level datasets from multiple studies.
This can improve harmonisation and subgroup analysis, but requires major collaboration and careful governance.
37. Data harmonisation is a major challenge
Different variable names.
different units.
different outcome definitions.
different missing-value codes.
Interoperability is part of meta-analysis.
38. Meta-analysis inherits every upstream measurement problem
Biased instruments.
poor sampling.
confounding.
selective reporting.
Pooling cannot turn weak inputs into strong evidence automatically.
39. Garbage in, precision out is possible
A meta-analysis can produce a narrow confidence interval around a biased pooled estimate.
Precision is not validity.
40. Risk-of-bias assessment should affect interpretation
A pooled effect dominated by high-risk studies should be described differently from one supported by several independently strong designs.
41. Meta-analysis and systematic review are different jobs
The systematic review defines and retrieves the evidence base.
The meta-analysis combines comparable estimates statistically.
Search before pooling.
42. Meta-analysis and synthesis are connected
Quantitative pooling is one form of scientific synthesis.
Narrative and mechanistic synthesis may still be required to explain heterogeneity and external validity.
43. Meta-analysis and triangulation are different
Meta-analysis often combines similar effect estimates.
Triangulation seeks convergence across methods with different bias structures.
Both can strengthen evidence in different ways.
44. Meta-analysis and external validity are connected
Variation across sites and populations helps map where an effect travels.
A pooled mean without heterogeneity can hide that question.
45. Meta-analysis and effect size are inseparable
The fundamental object being pooled is usually an effect estimate, not a significance label.
Magnitude remains central.
46. Meta-analysis and confidence intervals are inseparable
Each study contributes an estimate and uncertainty.
The pooled estimate also requires uncertainty.
47. Meta-analysis can be prospective
Studies can agree on outcomes and analysis before results are known, then combine evidence later.
This reduces some selective-reporting problems.
48. Network meta-analysis can compare several interventions
Direct and indirect evidence can be connected across a network of comparisons under additional assumptions.
Consistency between evidence pathways must be checked.
49. Indirect comparisons require transitivity
If A was compared with B in one population and B with C in another very different population, the A-versus-C inference may be distorted by effect modifiers.
Network structure does not remove external-validity questions.
50. AI can accelerate extraction and harmonisation
It can help identify effect measures, sample sizes, outcomes and study characteristics.
But automated extraction errors can propagate into pooled estimates.
51. AI can fabricate precision
If one extracted number is wrong, statistical software may still produce an elegant pooled result.
Every critical value should remain traceable to the source.
52. AI can help with meta-analytic diagnostics
Useful prompts:
“Explain why these effect measures are not directly comparable.”
“List plausible sources of heterogeneity.”
“Show how one influential study changes the pooled estimate.”
“Compare fixed-effect and random-effects interpretations.”
53. AI summaries should not replace forest plots
A sentence saying “overall evidence favours X” can hide study spread, uncertainty and outliers.
The quantitative evidence structure should remain visible.
54. Parents can teach meta-analysis intuition through repeated school evidence
One test score improves.
Another does not.
A third improves strongly.
Instead of choosing the favourite result, compare the entire pattern and how reliable each measurement was.
55. Small-group tuition can build a mini forest plot
Give three experiments with:
effect estimate.
uncertainty range.
sample size.
Students place them on one number line and discuss which result is most precise and whether the effects appear consistent.
56. A compact meta-analysis checklist
- Was the evidence base identified systematically?
- Are the studies estimating comparable quantities?
- Are effect directions aligned?
- What weighting model is used?
- Is a common-effect or varying-effect model scientifically appropriate?
- How much heterogeneity exists?
- What explains that heterogeneity?
- What does the prediction interval imply?
- Are any studies disproportionately influential?
- Could publication bias or small-study effects matter?
- How does risk of bias affect interpretation?
- Do sensitivity analyses preserve the conclusion?
- What populations or settings remain outside the evidence?
57. Frequently asked questions
What is a meta-analysis?
A meta-analysis is a statistical synthesis that combines sufficiently comparable effect estimates from multiple studies using an explicit weighting and modelling framework.
Is meta-analysis the same as systematic review?
No. Systematic review determines how evidence is found and assessed. Meta-analysis is the quantitative pooling step used when studies are suitable for combination.
What is heterogeneity?
Heterogeneity is variation in true or observed effects across studies beyond what would be expected from sampling uncertainty alone.
Why are larger studies often weighted more?
Because their effect estimates are usually more precise, although statistical weight should not be confused with methodological quality.
Can meta-analysis be wrong?
Yes. Incomplete searches, biased studies, incompatible outcomes, publication bias, incorrect extraction and inappropriate statistical models can all produce misleading pooled results.
How does meta-analytic thinking help Secondary Science?
It teaches learners to combine effect magnitude, uncertainty and between-study variation rather than counting significant studies.
58. Continue the Science Education Systems series
- How Scientific Systematic Review Works
- How Scientific Preregistration Works
- How Scientific Registered Reports Work
Conclusion: A pooled number should summarise evidence, not erase it
Maya counts studies.
Jia Jun weights information.
Hana inspects heterogeneity.
Ethan asks which study designs and populations created the spread.
Science needs all four.
Find the whole evidence base.
align the effects.
weight transparently.
show uncertainty.
preserve heterogeneity.
Then let the pooled estimate be a map of the evidence—not a curtain drawn across its differences.
