Science Education Systems · Article 100. Maya, Jia Jun, Hana and Ethan are fictional learners used to make scientific reasoning visible. This article owns one distinct scientific job: dimensionality reduction—representing high-dimensional data with fewer variables or latent dimensions while preserving scientifically useful structure. It does not replace regularisation, feature selection or visualisation. Its job is to reduce representation complexity and make the cost of that compression explicit.
The 50-second parent route
Modern Science can measure hundreds, thousands or millions of variables at once. More columns do not automatically mean more understanding.
The route is:
high-dimensional data → scientific purpose → scaling → dependence structure → linear or nonlinear reduction → retained dimensions → explained structure → reconstruction or neighbourhood check → cross-validation → interpretation → uncertainty → bounded use
The fastest diagnostic is to ask: What information was discarded, and could that discarded structure matter to the scientific question?
This article extends How Scientific Regularisation Works, How Scientific Cross-Validation Works, How Scientific Scale Works and How Science Knowledge Networks Work.
1. Dimension means a variable or coordinate needed to describe a data point
A plant measured by height and mass has two dimensions. Add leaf area, water content, colour channels, gene expression and environmental measurements, and the dimensionality rises quickly.
2. High dimensionality can hide simple structure
Many variables may move together because they reflect a smaller number of underlying processes.
3. Maya’s first error is “more variables means more information”
She measures fifty nearly identical features. Her repair is to distinguish redundant measurements from genuinely new information.
4. Jia Jun’s first error is projecting to two dimensions and trusting the picture literally
His repair is to recognise that every 2D embedding distorts some relationships.
5. Hana’s first error is using PCA without scaling
One variable is measured in thousands and dominates the variance. Her repair is to decide deliberately whether raw scale or standardised scale matches the scientific meaning.
6. Ethan’s first error is choosing the number of components by appearance
He keeps enough components to make a pleasing plot. His repair is to use explained variance, reconstruction, downstream validation and scientific interpretability.
7. The curse of dimensionality changes geometry
As dimensions increase, data become sparse relative to the volume of the space.
8. Distances can become less informative
In high dimensions, nearest and farthest neighbours can become surprisingly similar in distance under some distributions.
9. More dimensions demand more data
To sample high-dimensional space densely, the required number of observations grows rapidly.
10. High dimensionality increases overfitting risk
With many features and few observations, models can find accidental correlations.
11. Dimensionality reduction can act as complexity control
Reducing variables can make downstream models more stable.
12. Dimensionality reduction is not automatically regularisation
It changes representation. Regularisation constrains model fitting. They often work together but are distinct jobs.
13. Feature selection keeps original variables
Choose a subset of measured features and discard the rest.
14. Feature extraction creates new variables
PCA, autoencoders and manifold methods combine original variables into new coordinates.
15. Feature selection is easier to interpret
The retained variables keep their original physical meaning.
16. Feature extraction can compress redundancy more efficiently
Several correlated variables can become one latent axis.
17. Principal Component Analysis is the classic linear method
PCA finds orthogonal directions of maximum variance in centred data.
18. The first principal component captures the largest possible variance among unit-length linear directions
It is a weighted combination of the original variables.
19. The second principal component captures the largest remaining variance orthogonal to the first
Subsequent components continue this pattern.
20. PCA rotates the coordinate system
It does not merely delete columns. It finds a new basis aligned with variance structure.
21. PCA components are linear combinations
PC1 = a1X1 + a2X2 + … + apXp.
22. Loadings describe variable contribution to a component
Large positive or negative loadings indicate strong association with the component direction.
23. Component signs are arbitrary
Multiplying a principal component and all its loadings by −1 describes the same axis.
24. Scores locate observations in component space
Each sample receives a coordinate along PC1, PC2 and subsequent components.
25. Explained variance quantifies retained variation
Each principal component accounts for a proportion of total variance.
26. Cumulative explained variance helps choose dimension
Keep enough components to retain a target fraction of variance when that criterion matches the purpose.
27. Explained variance is not explained scientific importance
A low-variance feature can be scientifically critical, especially for rare events or small treatment effects.
28. High variance can be nuisance
Lighting conditions, batch effects or body size may dominate variance while the scientific signal is subtle.
29. PCA is unsupervised
It ignores the outcome label when finding components.
30. Therefore PCA can discard predictive directions
A low-variance direction strongly related to the outcome may be removed.
31. Supervised dimensionality reduction uses outcome information
Linear Discriminant Analysis, partial least squares and supervised embeddings can seek dimensions relevant to labels or responses.
32. Supervised reduction must remain inside cross-validation
Using all labels to build the representation before evaluation causes leakage.
33. PCA should usually be fitted inside each training fold too
Even without labels, held-out feature covariance should not shape the training representation during honest evaluation.
34. Scaling changes PCA profoundly
PCA on the covariance matrix favours variables with larger numerical variance.
35. Standardised PCA uses the correlation structure
Each variable is centred and divided by its standard deviation, giving equal marginal variance before decomposition.
36. Standardisation is not always scientifically correct
If absolute variance is meaningful, forcing every variable to equal variance can overemphasise noisy measurements.
37. Units matter
Mixing millimetres, kilograms and percentages without scaling can make the representation depend on arbitrary unit choice.
38. Log transforms can change covariance structure
Multiplicative biological or chemical variables may become more interpretable after log transformation.
39. Transformations should occur before fitting PCA inside the pipeline
Any data-derived transform parameters should be estimated using training data only.
40. Singular Value Decomposition underlies many PCA computations
SVD factors a data matrix into orthogonal directions and singular values.
41. SVD works directly on the data matrix
It can be numerically convenient when variables outnumber observations or the covariance matrix is large.
42. Truncated SVD keeps only leading components
It is widely used for sparse matrices such as text term-document representations.
43. Reconstruction error measures information loss
Project data into fewer dimensions, reconstruct back to the original space and measure the discrepancy.
44. Low reconstruction error means the reduced coordinates preserve much original variation
But not necessarily the scientifically important variation.
45. Scree plots show eigenvalue decline
An elbow can suggest where additional components contribute diminishing variance.
46. Scree elbows are often subjective
Several observers may choose different cut points.
47. Parallel analysis provides a stronger reference
Compare observed eigenvalues with eigenvalues from random data under a chosen null model.
48. Cross-validation can choose component count
Select the dimensionality that gives best downstream held-out performance.
49. Component count is a hyperparameter
It should be tuned without touching final test data.
50. Too few components underfit the representation
Important structure is discarded.
51. Too many components retain noise
Compression benefits disappear and downstream overfitting can return.
52. PCA assumes linear structure
If data lie on a curved manifold, a small linear subspace may represent them poorly.
53. Kernel PCA introduces nonlinear feature mappings
It performs PCA in an implicit transformed space defined by a kernel.
54. Kernel choice changes the geometry
Different kernels represent different similarity assumptions.
55. Nonlinear dimensionality reduction often targets neighbourhood preservation
Methods such as t-SNE and UMAP emphasise local relationships rather than global variance.
56. t-SNE builds probability-like neighbourhood similarities
Nearby points in high-dimensional space are encouraged to remain nearby in the low-dimensional embedding.
57. t-SNE is powerful for visualisation
Clusters and local neighbourhoods can become visible in two dimensions.
58. t-SNE global distances are easy to overinterpret
The distance between separated clusters in the plot may not faithfully represent high-dimensional distance.
59. t-SNE cluster size can be misleading
Visual area in the embedding does not necessarily represent population variance or density.
60. Perplexity changes t-SNE neighbourhood scale
Different perplexity settings can reveal different local structures.
61. Random initialisation can change t-SNE layout
Stable scientific conclusions should not depend on one attractive run.
62. UMAP also constructs a neighbourhood graph
It seeks a low-dimensional representation preserving local fuzzy topological relationships under its assumptions.
63. UMAP often preserves more broad structure than t-SNE in practice
But global geometry should still be interpreted cautiously.
64. UMAP has important hyperparameters
Number of neighbours controls local-versus-broader structure; minimum distance affects embedding compactness.
65. Different UMAP settings can create different visual stories
Scientific interpretation should test parameter stability.
66. Nonlinear embeddings can invent apparent clusters
Visual separation may result partly from the embedding objective rather than discrete biological or scientific groups.
67. Cluster appearance is not cluster proof
Use independent labels, density analysis, stability tests or external measurements before declaring new categories.
68. Neighbourhood preservation can be quantified
Trustworthiness and continuity-like metrics assess how well local relationships survive projection.
69. Global structure can be quantified separately
Distance correlations, reconstruction proxies or downstream tasks can evaluate broader preservation.
70. One embedding cannot preserve every relationship
Compression forces trade-offs among local distances, global distances, density and interpretability.
71. Manifold learning assumes lower-dimensional structure
High-dimensional observations may lie near a lower-dimensional curved surface.
72. Isomap preserves approximate geodesic distances
It builds a neighbourhood graph and estimates distances along the manifold.
73. Isomap can fail when neighbourhood graphs disconnect
Too small a neighbourhood misses paths; too large a neighbourhood shortcuts curvature.
74. Locally Linear Embedding preserves local reconstruction relationships
Each point is represented by its neighbours, then those weights are reproduced in lower dimensions.
75. Spectral embedding uses graph eigenvectors
It represents data using the structure of a similarity graph.
76. Graph construction becomes a scientific assumption
Which points count as neighbours and how similarities are weighted determine the embedding.
77. Autoencoders learn nonlinear compression
A neural network maps inputs to a lower-dimensional latent code and reconstructs them.
78. The bottleneck dimension controls compression
Too narrow loses signal; too wide can reproduce inputs without learning useful structure.
79. Autoencoder regularisation matters
Denoising, sparsity, weight decay and variational objectives can shape latent representations.
80. Variational autoencoders model a latent distribution
They add probabilistic structure to the latent space under assumptions about priors and decoders.
81. Deep latent variables are harder to interpret
A latent dimension may not correspond neatly to one physical quantity.
82. Interpretability is a scientific design goal
If the purpose is mechanism discovery, a slightly less compressed but more interpretable representation may be preferable.
83. Independent Component Analysis seeks statistically independent sources
ICA is useful when observed mixtures may arise from underlying independent signals.
84. PCA and ICA optimise different criteria
PCA decorrelates and orders variance; ICA seeks independence and often non-Gaussian structure.
85. Non-negative matrix factorisation preserves additive parts
When data are non-negative, NMF represents observations as non-negative combinations of non-negative components.
86. NMF can yield part-based representations
In images, spectra or counts, additive components can be easier to interpret than signed PCA loadings.
87. Matrix factorisation is a broad dimensionality-reduction family
Low-rank models express a large matrix through smaller factor matrices.
88. Low rank is an assumption about redundancy
It says much observed variation can be generated by a smaller number of latent dimensions.
89. Rank selection is a model-selection problem
Too low loses structure; too high retains noise.
90. Missing data complicates factorisation
Naive PCA requires complete matrices, but specialised methods can estimate latent structure with missing entries under assumptions.
91. Imputation before PCA can create false structure
If imputation smooths values toward group means, components may reflect the imputation model.
92. Batch effects can dominate components
In biological data, laboratory batch can explain more variance than the scientific condition.
93. A PCA plot can reveal batch effects
That is valuable diagnosis, but the batch component should not be mistaken for biology.
94. Removing batch effects changes the data
Correction methods should preserve real scientific differences and remain inside validation pipelines when outcome information is used.
95. Worked case: plant measurements
Height, stem diameter, leaf count, leaf area and biomass are strongly correlated.
96. PCA may produce a “size” component
All growth variables load positively on PC1, summarising overall plant size.
97. A second component can represent allocation pattern
Leaf area versus stem thickness may separate plants with different growth strategies.
98. The component is a scientific hypothesis, not a fact
Calling PC1 “vigour” requires external evidence that the latent axis really corresponds to vigour.
99. Worked case: environmental sensors
Ten air-quality sensors measure overlapping pollutants and meteorological variables.
100. Dimensionality reduction can reveal common pollution patterns
One component may represent combustion-related variation, another weather-driven dispersion.
101. But source interpretation requires chemistry and context
Loadings alone do not prove which source emitted the pollutants.
102. Worked case: gene expression
Thousands of genes are measured across a few hundred samples.
103. PCA can show dominant sample variation
Disease status, tissue type or batch may separate along leading components.
104. High-dimensional biology needs strict validation
Feature selection and dimensionality reduction must occur inside training folds to avoid optimistic prediction.
105. Worked case: learner skill profile
A Science diagnostic contains twenty correlated subskill measures.
106. A latent “scientific reasoning” component may emerge
Several evidence-evaluation and explanation tasks move together.
107. Construct validity still matters
The component should not be labelled “reasoning” merely because the name sounds useful.
108. Primary Science can learn dimensionality reduction through grouping observations
Many details can be summarised into fewer patterns without forgetting what was left out.
109. Primary 3 can sort many objects by two useful features
Colour, size, material or magnetism. Ask which features actually distinguish the groups.
110. Primary 4 can identify redundant measurements
Two measures always move together. Do we need both for this question?
111. Primary 5 can create a composite score cautiously
Combine several related measurements, then ask what information the summary hides.
112. Primary 6 can compare two-dimensional plots
Project three variables into two axes and discuss which relationship is lost.
113. Secondary Science can formalise PCA
Students can centre data, understand covariance, loadings, scores and explained variance.
114. Secondary Science can compare linear and nonlinear embeddings
PCA preserves global linear variance; t-SNE and UMAP emphasise local neighbourhoods differently.
115. Dimensionality reduction and feature selection are different
Selection keeps original variables. Extraction creates new coordinates.
116. Dimensionality reduction and regularisation are different
Reduction compresses representation; regularisation constrains model fitting.
117. Dimensionality reduction and clustering are different
Reduction changes coordinates. Clustering assigns groups. A 2D embedding that looks clustered has not automatically performed valid clustering.
118. Dimensionality reduction and visualisation are different
Visualisation is one use. Reduction can also improve storage, denoising, modelling and measurement design.
119. Dimensionality reduction and model selection are connected
The number of components and reduction method are hyperparameters requiring validation.
120. Dimensionality reduction and identifiability are connected
Compression can remove variables needed to distinguish mechanisms or parameters.
121. Dimensionality reduction and uncertainty are connected
Estimated component directions vary across samples. A component is not known exactly.
122. Bootstrap PCA can assess loading stability
Resample observations, refit PCA and inspect how loadings and subspaces change.
123. Near-equal eigenvalues make component directions unstable
The subspace may be stable while individual component axes rotate substantially.
124. Sign alignment is needed across bootstrap components
Because component signs are arbitrary, direct averaging without alignment can cancel equivalent solutions.
125. Component order can swap
When eigenvalues are close, PC2 and PC3 can exchange order across samples.
126. Scientific interpretation should focus on stable subspaces when needed
Do not over-label one unstable axis.
127. Outliers can dominate PCA
Because variance is squared-distance sensitive, extreme observations can rotate components.
128. Robust PCA variants reduce outlier influence
They separate low-rank structure from sparse corruption under specific assumptions.
129. An outlier can be scientifically important
Do not remove it automatically. Determine whether it is measurement error, rare regime or new phenomenon.
130. Sparse PCA encourages simpler loadings
Many loadings become zero, making components easier to interpret.
131. Sparse PCA trades reconstruction for interpretability
The component may explain less variance but use fewer variables.
132. Factor analysis differs from PCA
Factor analysis models latent causes plus unique measurement error, while PCA is primarily a variance-decomposition transformation.
133. PCA components should not automatically be called latent causes
They are directions of variance, not proof of hidden mechanisms.
134. Factor models require stronger assumptions
Loadings, residual independence and latent structure become explicit model components.
135. Canonical correlation finds paired low-dimensional relationships
It seeks linear combinations of two variable sets that are maximally correlated.
136. Partial least squares uses outcome covariance
It extracts components that explain predictor structure relevant to a response.
137. Supervised methods can outperform PCA for prediction
But their use of labels increases leakage risk and can reduce general scientific interpretability.
138. Random projection offers fast approximate distance preservation
Project high-dimensional data through a random matrix into fewer dimensions.
139. Johnson–Lindenstrauss intuition supports random projection
A sufficiently large reduced dimension can approximately preserve pairwise distances for a finite set of points.
140. Random projection sacrifices interpretability
The axes are random mixtures rather than scientifically meaningful directions.
141. Hashing tricks reduce feature dimensions in text and streaming systems
Features are mapped into a fixed number of bins using hash functions.
142. Hash collisions are controlled information loss
Different features can map to the same bin, trading interpretability for memory efficiency.
143. AI embeddings are dimensional representations
Text, images and other data are encoded into vectors whose dimensions are learned rather than hand-defined.
144. Embedding dimensions are usually not individually interpretable
Meaning emerges from vector relationships across many coordinates.
145. Reducing embeddings can aid visualisation or retrieval speed
But compression may alter nearest neighbours and semantic structure.
146. Nearest-neighbour retrieval depends on geometry
Dimensionality reduction that distorts distances can change which documents or images are retrieved.
147. AI benchmark visualisations can mislead
A colourful UMAP plot of model embeddings may suggest crisp categories that are not stable under different seeds or parameters.
148. AI can help learners compare methods
Useful prompts include: “Create data where PCA works well and where it fails,” “Show t-SNE plots with different perplexities,” “Give me a PCA component dominated by scale,” and “Ask what information a two-dimensional projection discarded.”
149. AI can generate attractive but meaningless embeddings
Plot aesthetics should never replace quantitative preservation checks and scientific validation.
150. Parents can understand dimensionality reduction through report cards
A single overall mark compresses reading, knowledge, reasoning, timing and answer precision. Useful summary can also hide the cause of weakness.
151. Diagnostics need both summary and detail
A compact learner profile helps navigation, but intervention often requires returning to the original subskill dimensions.
152. Small-group tuition can build concept maps before compression
Group related errors, then ask whether one underlying weakness explains several surface mistakes.
153. Educational compression should preserve actionable distinctions
If “Science ability” merges reading and scientific reasoning, the summary may be too coarse to guide teaching.
154. Independent-attempt task 1: scale trap
Create two correlated variables with the same pattern but one measured in units 1000 times larger. Compare PCA before and after standardisation.
155. Independent-attempt task 2: explained variance
Given eigenvalues 6, 3, 1, calculate the variance explained by each component and decide how many to retain for several purposes.
156. Independent-attempt task 3: low-variance signal
Create a classification problem where the outcome depends on a low-variance direction. Explain why unsupervised PCA may discard it.
157. Independent-attempt task 4: embedding instability
Run a thought experiment where t-SNE produces three visual clusters under one perplexity and two under another. What conclusions remain justified?
158. Independent-attempt task 5: leakage
Place PCA before versus inside cross-validation and explain which workflow gives an honest performance estimate.
159. Diagnostic error: PCA means “important variables”
Repair by remembering that PCA ranks directions by variance, not scientific importance.
160. Diagnostic error: first two PCs are the whole dataset
Repair by checking cumulative explained variance and omitted structure.
161. Diagnostic error: visible t-SNE cluster equals biological subtype
Repair with stability, external labels and independent evidence.
162. Diagnostic error: global t-SNE distance interpreted literally
Repair by restricting interpretation mainly to local neighbourhood structure unless validated otherwise.
163. Diagnostic error: PCA fit before cross-validation
Repair by fitting the transformation inside every training fold.
164. Diagnostic error: component labels treated as causes
Repair by validating latent interpretations against external measurements and mechanisms.
165. Diagnostic error: component count chosen by pretty plot
Repair with explained variance, reconstruction, stability and held-out performance.
166. Diagnostic error: nonlinear embedding used for exact quantitative distance
Repair by checking what geometry the method preserves and what it distorts.
167. Diagnostic error: dimension reduction applied to units with incompatible meaning
Repair with scaling, transformations and scientific review before decomposition.
168. The independence test
Give a learner a high-dimensional dataset and three candidate reductions. Can they choose a method based on purpose, scale, labels, geometry and validation rather than on which plot looks nicest? That is transferable dimensionality-reduction reasoning.
169. The evidence boundary
Every reduced representation is conditional on preprocessing, distance assumptions, retained dimensions, hyperparameters and algorithm. A low-dimensional pattern is an analytical view of the data—not the data’s one true shape.
170. A compact dimensionality-reduction checklist
- Why is dimensionality reduction needed?
- Is the goal visualisation, prediction, compression, denoising or interpretation?
- Are variables on comparable scales?
- Should the method keep original features or create latent ones?
- Is linear structure plausible?
- How much variance or reconstruction quality is retained?
- Could low-variance scientific signal be lost?
- Is the method supervised or unsupervised?
- Is reduction fitted inside cross-validation?
- How is the number of dimensions chosen?
- Are nonlinear embedding results stable to hyperparameters and seeds?
- What local and global geometry is preserved?
- Are apparent clusters independently validated?
- How stable are components across resamples?
- Can the latent dimensions be interpreted scientifically?
- What important information has been discarded?
171. Frequently asked questions
What is dimensionality reduction?
It is the process of representing data using fewer variables or latent dimensions while preserving structure relevant to a scientific or computational purpose.
What is PCA?
Principal Component Analysis is a linear method that creates orthogonal directions ordered by the amount of variance they explain.
What is explained variance?
It is the proportion of total data variance captured by a principal component or set of components.
How is PCA different from feature selection?
PCA creates new linear combinations of variables; feature selection keeps a subset of original variables.
How are t-SNE and UMAP different from PCA?
They are nonlinear neighbourhood-oriented embedding methods designed mainly to preserve local relationships rather than global linear variance.
Can a t-SNE or UMAP cluster prove a real scientific category?
No. The cluster should be tested for stability and validated with independent evidence.
How does dimensionality reduction help PSLE Science?
The formal methods are advanced, but the habit is accessible: summarise many observations while asking what important detail the summary hides.
How does it deepen in Secondary Science?
Students can connect covariance, PCA, explained variance, scaling, low-rank structure, nonlinear embeddings and cross-validation more formally.
172. Continue the Science Education Systems series
- How Scientific Model Selection Works
- How Scientific Cross-Validation Works
- How Scientific Regularisation Works
Conclusion: Compression is useful only when we remember what it threw away
Maya sees hundreds of variables.
Jia Jun finds the shared structure.
Hana checks what the projection distorted.
Ethan asks whether the reduced representation still answers the scientific question.
Science needs all four.
Scale deliberately.
choose the geometry.
compress only what is redundant.
validate the retained structure.
show the information loss.
Then let fewer dimensions create clarity without pretending the discarded dimensions never mattered.

