Science Education Systems · Article 97. Maya, Jia Jun, Hana and Ethan are fictional learners used to make scientific reasoning visible. This article owns one distinct scientific job: model selection—choosing among competing scientific models without confusing better fit with better explanation. It does not replace cross-validation, regularisation or validation. Its job is to decide which candidate model earns the strongest support for the purpose at hand.
The 50-second parent route
A more complicated model almost always has more ways to fit the data. That does not make it a better scientific model.
The route is:
question → candidate models → assumptions → fit → complexity → residuals → information criterion or validation score → uncertainty → scientific plausibility → selection → external check → revision
The fastest diagnostic is to ask: Did the model win because it captured real structure, or because it had more freedom to chase this dataset?
This article extends How Scientific Validation Works, How Scientific Robustness Works, How Scientific Identifiability Works and How Scientific Cross-Validation Works.
1. Model selection starts with a scientific question
Do we want explanation, prediction, compression, control or decision support? The best model for one purpose may not be best for another. A highly interpretable mechanistic model may be preferred for explanation; a flexible predictive model may win for forecasting.
2. Candidate models should represent meaningful alternatives
Model selection is not strongest when hundreds of arbitrary formulas are thrown at the data. It is strongest when candidates encode scientifically plausible hypotheses or defensible levels of complexity.
3. Maya’s first error is “highest R² wins”
She chooses the model with the largest training R². Her repair is to recognise that additional parameters often raise in-sample fit even when they capture noise.
4. Jia Jun’s first error is “simplest always wins”
He prefers a straight line because it is easy. His repair is to see that parsimony does not mean refusing complexity when the phenomenon genuinely requires it.
5. Hana’s first error is “AIC is truth”
She reports the lowest AIC model as proven. Her repair is to treat information criteria as comparative evidence under assumptions, not as absolute truth machines.
6. Ethan’s first error is choosing after seeing every result
He tests dozens of models and reports only the winner. His repair is to account for selection uncertainty and protect evaluation data from repeated tuning.
7. Fit measures how closely a model reproduces observed data
Residual sum of squares, likelihood, deviance and prediction error are examples. Good fit is necessary for many purposes but rarely sufficient.
8. Complexity measures model flexibility
Parameter count is one crude proxy. Effective degrees of freedom, tree depth, basis richness, hidden units and flexible interactions can also contribute complexity.
9. Flexible models can fit random noise
If a model has enough freedom, it can learn idiosyncrasies that will not repeat in new data. This is overfitting.
10. Underfitting is the opposite failure
A model is too simple to capture important structure. Residuals retain systematic pattern and predictions are biased.
11. Model selection balances underfitting and overfitting
The scientific challenge is not maximal fit or minimal complexity. It is enough complexity to capture reproducible structure without spending degrees of freedom on noise.
12. The bias–variance trade-off is one model-selection lens
Simpler models can have more bias but lower variance. More flexible models can reduce bias but become unstable across samples.
13. Training error usually favours complexity
A more flexible model can adapt to the training set. Therefore training error alone cannot fairly compare models with different flexibility.
14. Penalised criteria correct for optimism
AIC, BIC and related criteria reward fit while penalising complexity. They formalise different ideas about what constitutes a good model.
15. AIC balances likelihood and parameter count
For a common form, AIC = −2 log-likelihood + 2k, where k is the number of estimated parameters. Lower values are preferred among candidates fit to the same data and likelihood framework.
16. AIC is motivated by predictive information loss
Its logic is closer to estimating which model will minimise expected information loss under repeated sampling, not to identifying a literally true model.
17. BIC uses a stronger sample-size-dependent penalty
BIC = −2 log-likelihood + k log n in common settings. As n grows, the complexity penalty generally becomes stronger than AIC’s 2k penalty.
18. AIC and BIC can select different models
That disagreement is not necessarily an error. They optimise different asymptotic goals under different theoretical motivations.
19. The “lowest AIC” rule needs context
AIC differences matter more than absolute AIC values. The number itself has no simple standalone scientific meaning.
20. AIC differences can be turned into relative support
Akaike-style weights summarise relative evidence within the candidate set under the criterion’s assumptions.
21. Relative support is candidate-set dependent
If every candidate model is poor, one still receives the best AIC. Winning a weak race does not guarantee adequacy.
22. Residual diagnostics remain necessary
A model with the best criterion can still leave nonlinearity, heteroscedasticity, autocorrelation or systematic subgroup error.
23. Model selection should not rescue a misspecified candidate set
If the true mechanism requires a threshold and every candidate is linear, selecting the best line does not solve the scientific problem.
24. Nested models differ by additional parameters
One model is a restricted special case of another. Likelihood-ratio tests can compare some nested models under suitable regularity conditions.
25. Non-nested models require other comparison strategies
A polynomial, mechanistic differential equation and random forest may not be related by simple parameter restriction. Cross-validation or information criteria can provide comparison depending on the objective.
26. Scientific plausibility should constrain model selection
A model violating conservation laws or impossible parameter signs should not be celebrated merely because the fit score is excellent.
27. Mechanistic consistency can break statistical ties
Two models predict equally well. One aligns with known physics and generates correct intermediate predictions. The other is an opaque curve fit. For explanatory Science, the mechanistic model may deserve preference.
28. Prediction can break explanatory ties too
If both models are scientifically plausible, performance on genuinely new data can distinguish their generalisation ability.
29. Cross-validation is a model-selection tool, not the definition of model selection
Cross-validation estimates out-of-sample performance. Model selection is the broader decision about which candidate to prefer, possibly using cross-validation, information criteria, theory, residuals and constraints.
30. Cross-validation score has uncertainty
Different fold assignments can change the ranking of models. If two candidates differ by less than the variability of the score, declaring a clear winner may be premature.
31. Repeated cross-validation can reveal ranking instability
Repeat the validation split and inspect the distribution of performance differences. Stable winners are more convincing than one lucky split.
32. The one-standard-error rule favours simpler models
When several models perform within one estimated standard error of the best, choose the simplest candidate among them. This is a practical parsimony heuristic.
33. Simplicity should be defined scientifically
Fewer parameters may be simpler numerically but harder to interpret mechanistically. Complexity has several dimensions.
34. Effective complexity can exceed parameter count intuition
A smoothing model with many basis coefficients may have fewer effective degrees of freedom because regularisation constrains them strongly.
35. Regularisation turns model selection into a continuum
Instead of choosing between a few discrete models, vary a penalty strength that controls complexity continuously. Article 99 owns that process.
36. Hyperparameter tuning is model selection
Choosing tree depth, regularisation strength, number of neighbours or kernel width selects among models even if the algorithm name stays the same.
37. Feature selection is model selection
Every chosen subset of predictors defines a different candidate model.
38. Feature selection outside validation causes optimism
If all data are used to choose features before cross-validation, the held-out folds have already influenced the model design.
39. Preprocessing can be part of model selection
Scaling, imputation, transformation and dimensionality reduction can alter performance and should be fit inside training folds when tuned from data.
40. A pipeline is the true candidate model
The scientific model-selection object often includes preprocessing + feature construction + estimator + hyperparameters + threshold.
41. Leakage turns model selection into self-congratulation
If evaluation data influence tuning, reported performance describes the dataset-search process, not honest generalisation.
42. Test data should remain untouched until final evaluation
If the test set is checked repeatedly while adjusting models, it becomes another training signal.
43. Nested cross-validation separates tuning from evaluation
An inner loop selects models; an outer loop estimates performance of that selection procedure. This prevents the same validation results from serving both jobs.
44. Model-selection bias grows with search breadth
Try enough models and one will win partly by chance. The larger the search, the more important independent evaluation becomes.
45. Winner’s curse applies to models
The selected model often has an overly optimistic estimated score because it won partly through favourable noise.
46. Selection uncertainty should be reported
Which models were close? How often did each win across resamples? Would a slightly different dataset select another candidate?
47. Model averaging can respect selection uncertainty
Instead of choosing one winner, combine predictions or inferences across several plausible models with justified weights.
48. Model averaging is not always appropriate
Combining incompatible mechanistic interpretations may obscure scientific meaning even if predictive performance improves.
49. Bayesian model comparison offers another route
Bayes factors or posterior model probabilities compare models under priors and likelihoods. The results can be sensitive to prior specification.
50. Posterior predictive checks remain necessary
A model can receive high relative posterior probability among poor candidates and still fail to reproduce important data features.
51. WAIC and leave-one-out approaches estimate predictive performance
They account for effective complexity using the posterior distribution and can support Bayesian model comparison.
52. No criterion eliminates judgement
AIC, BIC, cross-validation, Bayes factors and predictive scores quantify different aspects. Scientific purpose chooses which aspect matters.
53. Primary Science can learn model selection through competing explanations
Why did one ice cube melt faster? Sunlight? smaller size? warmer plate? Students list candidate explanations and design observations that discriminate them.
54. Primary 3 can compare simple models
One model says plant growth depends only on water. Another adds light. Which better explains several observations without adding unnecessary assumptions?
55. Primary 4 can use counterexamples
A rule explains three cases but fails a fourth. A revised rule explains all four. The learner sees why model selection responds to evidence.
56. Primary 5 can learn parsimony
Two models explain the same pattern. One needs one mechanism; another invents five unobserved causes. Prefer the simpler model unless evidence supports the extra complexity.
57. Primary 6 can learn holdout thinking
Build a rule from four observations, then test on a fifth observation not used to create it.
58. Secondary Science can formalise fit versus complexity
Students can compare residuals, parameter count, AIC-like penalties and held-out prediction error.
59. Secondary Science can compare nested models
Does adding a quadratic term improve fit enough to justify another parameter?
60. Secondary Science can compare mechanistic and empirical models
One model may fit slightly worse but have stronger physical interpretation and better extrapolation.
61. Worked case: linear versus quadratic growth
Five time points show accelerating growth. A line has systematic curved residuals. A quadratic fits better but adds one parameter.
62. Training fit favours the quadratic
It has more flexibility. The real question is whether curvature repeats in new observations.
63. AIC can reward the quadratic only if fit gain is sufficient
The criterion asks whether the log-likelihood improvement compensates for added complexity.
64. BIC may prefer the line in smaller or larger penalty contexts
Because its penalty depends on sample size, it can favour a simpler candidate more strongly.
65. Cross-validation may reveal which generalises better
If the quadratic consistently predicts held-out points more accurately, its added complexity earns support.
66. Worked case: ecological response curve
A species response to temperature could be linear, quadratic optimum-shaped or threshold-like.
67. Mechanism constrains the candidates
If physiology predicts an optimum, the quadratic-like model has scientific motivation rather than being mere curve flexibility.
68. Boundary conditions still matter
A quadratic can turn upward at extreme temperatures outside the observed range, creating biologically absurd extrapolation.
69. Worked case: learner diagnostic model
Model A says every error comes from content knowledge. Model B separates content, reading and explanation structure.
70. More detailed models need discriminating evidence
If the tasks cannot distinguish those failure modes, Model B’s extra detail is not identifiable.
71. Model selection therefore depends on identifiability
You cannot meaningfully select among models whose distinguishing parameters cannot be learned from the data.
72. Worked case: AI classifier
Candidate models range from logistic regression to deep neural network. The deep model achieves lower training loss.
73. Holdout performance may reverse the ranking
The simpler model can generalise better when data are limited and the deep model overfits.
74. Interpretability can be a model-selection criterion
In high-stakes scientific settings, the ability to inspect coefficients, mechanisms or failure conditions may matter beyond raw accuracy.
75. Calibration can be a selection criterion
Two classifiers have similar accuracy, but one produces probabilities that match observed frequencies better.
76. Robustness can be a selection criterion
A model may win average accuracy but fail catastrophically under modest distribution shift. Another may be slightly less accurate yet much more stable.
77. Fairness can be a selection criterion
Model performance across groups may differ. Selection should reflect the actual scientific and ethical decision context.
78. Computational cost can be a selection criterion
A model requiring one hour per prediction may be scientifically unusable if the operational decision takes seconds.
79. Measurement burden can be a selection criterion
A model needing fifty laboratory measurements may perform only slightly better than one needing five.
80. Simpler input requirements can improve transportability
A model based on commonly measured variables can be applied across more sites than one requiring specialised instruments.
81. Missing-data resilience can be a selection criterion
Some models fail when one predictor is unavailable. Others degrade gracefully.
82. Selection criteria should be declared before looking at final results
If researchers decide afterward which metric matters, the winner can be chosen opportunistically.
83. Preregistration can protect model-selection credibility
Specify candidate families, primary metric, validation scheme and decision rule before final data where appropriate.
84. Exploratory model search is still valuable
Exploration can discover useful patterns. The key is to label exploration honestly and validate selected models on new evidence.
85. Confirmation needs fresh data
Once the data helped inspire the model, those same data are no longer an independent confirmation.
86. Post-selection inference is difficult
Confidence intervals calculated as if the selected model had been specified in advance can be too optimistic.
87. Selection changes uncertainty
The model structure itself is uncertain, not just the parameters inside it.
88. Model uncertainty should propagate into decisions
If two plausible models make different policy recommendations, selecting one silently hides an important uncertainty source.
89. Ensemble methods are one response to model uncertainty
Random forests, boosting and stacked models combine many learners. Their success shows that one-model selection is not always optimal for prediction.
90. Ensembles trade simplicity for predictive stability
They may generalise well but become harder to explain mechanistically.
91. Stacking learns how to combine candidate predictions
Cross-validated predictions can be used to estimate ensemble weights, protecting against fitting the combiner on the same predictions used for training.
92. Blending without honest validation can overfit too
Combining models is another model-selection problem requiring held-out evidence.
93. The null model matters
A complex model should often be compared against simple baselines: mean prediction, last-value forecast, linear model or known physical rule.
94. Baselines prevent impressive-looking useless models
A sophisticated method that barely beats a naive benchmark may not justify complexity.
95. The baseline should match the task
Time-series forecasting needs time-aware naive models; classification may use majority class or calibrated simple models.
96. Model selection can be unstable in small samples
One observation can change the winner. Bootstrap or repeated validation can reveal this fragility.
97. Stability selection examines repeated selection behaviour
If a feature or model appears only in a small fraction of resamples, its apparent importance may be sample-specific.
98. Scientific conclusions should survive reasonable model alternatives
If the sign of an effect changes across plausible models, the effect is not robust enough for confident interpretation.
99. Specification curves can expose researcher degrees of freedom
Run many defensible model specifications and display how conclusions vary rather than presenting one convenient choice.
100. Multiverse analysis is related
It evaluates conclusions across many plausible analytical decisions, revealing whether results depend on one narrow specification.
101. More models can increase transparency or increase fishing
The difference is whether the candidate set is scientifically justified and the analysis reports the full selection process honestly.
102. Model-selection pipelines need version control
Record candidate definitions, data versions, preprocessing, hyperparameters and metrics so results can be reconstructed.
103. Model registries prevent forgotten losers
Only preserving the final winner erases the evidence about how much searching occurred.
104. AI can automate model search
AutoML systems can test algorithms, features and hyperparameters rapidly.
105. Automated search increases selection-bias risk
The easier it becomes to test thousands of pipelines, the more important nested or final independent evaluation becomes.
106. AI can help learners compare models
Useful prompts include: “Give me two models where the more complex one wins training fit but loses validation,” “Show AIC and BIC disagreement,” “Create residuals that reveal underfitting,” and “Challenge my chosen model with a simpler baseline.”
107. AI can produce fake sophistication
A model with fashionable terminology and many parameters may appear advanced. Scientific value comes from evidence, not architectural prestige.
108. Parents can use model-selection thinking in diagnosis
One simple story—“not enough effort”—may explain some performance. A richer model may include reading, retrieval, mechanism and timing. Choose the more detailed diagnosis only when evidence discriminates those causes.
109. Small-group tuition can compare explanation models
Three students explain the same wrong answer differently. The tutor tests which explanation predicts the learner’s behaviour on a new question.
110. The examination link
Students often choose between mental models implicitly. A question changes one variable; the correct model should predict the new outcome without needing a memorised case.
111. Transfer is model validation
A model built from familiar examples should work on unfamiliar but structurally similar cases if it captured the underlying Science.
112. Independent-attempt task 1: fit versus complexity
Draw data with slight curvature. Compare a line, quadratic and fifth-degree polynomial. Predict which has lowest training error and which may generalise best.
113. Independent-attempt task 2: AIC versus BIC
Give two candidate likelihoods and parameter counts. Calculate both criteria and explain why they can disagree.
114. Independent-attempt task 3: residual challenge
Create a model with good average error but systematic residual curvature. Explain why selection by one scalar score can miss misspecification.
115. Independent-attempt task 4: winner instability
Repeat a train-validation split several times and record which model wins. Explain what unstable ranking says about evidence strength.
116. Independent-attempt task 5: baseline test
Compare a sophisticated forecast with a seasonal-naive baseline. Decide whether the complexity earns its operational cost.
117. Diagnostic error: highest fit wins
Repair with complexity penalty, held-out prediction and residual checks.
118. Diagnostic error: lowest AIC means true
Repair by remembering that AIC ranks candidates under assumptions; it does not certify absolute truth.
119. Diagnostic error: simplest always wins
Repair by allowing complexity when reproducible structure requires it.
120. Diagnostic error: test set used during tuning
Repair by protecting final evaluation data or using nested cross-validation.
121. Diagnostic error: feature selection before validation
Repair by embedding feature selection inside each training fold.
122. Diagnostic error: one metric defines quality
Repair by matching metrics to scientific purpose: calibration, discrimination, error, robustness, interpretability or cost.
123. Diagnostic error: candidate set too narrow
Repair by asking whether a missing mechanism could explain residual structure.
124. Diagnostic error: candidate set too broad
Repair by constraining models with scientific plausibility and preregistered search logic.
125. Diagnostic error: model uncertainty ignored
Repair by reporting close competitors, selection stability or model-averaged conclusions where appropriate.
126. The independence test
Give a learner three competing models with training fit, validation scores, parameter counts and residual plots. Can they choose a model and justify the choice without simply picking the biggest number? That is transferable model-selection reasoning.
127. The evidence boundary
Model selection is conditional on the candidate set, metric, data, validation design and assumptions. A selected model is the best-supported option among what was seriously considered—not automatically the final truth of the system.
128. A compact model-selection checklist
- What scientific purpose should the model serve?
- What candidate models are scientifically plausible?
- What assumptions distinguish them?
- How is goodness of fit measured?
- How is complexity represented?
- What does residual structure show?
- Is AIC, BIC or another criterion appropriate?
- Is out-of-sample performance estimated honestly?
- Was tuning separated from evaluation?
- Were preprocessing and feature selection inside the validation loop?
- How stable is the ranking across resamples?
- Do models differ meaningfully or only trivially?
- What simpler baseline must be beaten?
- Does mechanistic plausibility alter the choice?
- How much model-selection uncertainty remains?
- Does the selected model survive new data?
129. Frequently asked questions
What is scientific model selection?
It is the process of comparing competing models and choosing among them using evidence about fit, complexity, predictive performance, residual behaviour, assumptions and scientific purpose.
What is AIC?
AIC is an information criterion that rewards likelihood fit while penalising the number of estimated parameters, aiming to balance fit and expected predictive information loss.
What is BIC?
BIC is a likelihood-based criterion with a complexity penalty that grows with sample size and is often more conservative toward added parameters than AIC.
Why not choose the model with the highest R²?
Because more flexible models often increase training fit by capturing noise, so in-sample fit alone rewards overfitting.
How does cross-validation help?
It estimates how candidate models perform on observations excluded from training, helping judge generalisation.
How does model selection help PSLE Science?
The formal criteria are advanced, but the habit is simple: compare explanations by how well they account for evidence without adding unnecessary assumptions.
How does it deepen in Secondary Science?
Students can compare residuals, complexity, validation performance, information criteria, uncertainty and mechanistic plausibility more formally.
130. Continue the Science Education Systems series
- How Scientific Cross-Validation Works
- How Scientific Regularisation Works
- How Scientific Dimensionality Reduction Works
Conclusion: The winning model must earn its complexity
Maya sees the best fit.
Jia Jun asks what extra freedom bought that fit.
Hana checks the residuals and new-data performance.
Ethan asks whether another dataset would choose the same winner.
Science needs all four.
Compare models.
penalise unnecessary freedom.
protect evaluation data.
inspect residuals.
respect mechanism.
report selection uncertainty.
Then prefer the model that explains or predicts enough to deserve every degree of complexity it uses.

