One question can be useful.
It can expose a sign error.
Reveal a misconception.
Show that a learner cannot start independently.
But one question is rarely enough to define the learner.
Mira misses one quadratic question.
Is factorisation weak?
Or was this one item unusually awkward?
Jonas overclaims one inference.
Is evidence calibration generally weak?
Or did one unfamiliar word distort the passage model?
Nadia mishandles one experiment.
Is variable reasoning unstable?
Or did the diagram contain an unfamiliar convention?
Training sampling is the deliberate selection of enough representative observations to infer a learner’s current capability without pretending that one item, one day or one score is the whole state.
Quick Read: Sample the Capability, Not Just the Worksheet
Hypothesis → Representative Items → Varied Conditions → Pattern Check → Decision
A useful sample should ask:
- Does the same error recur across more than one item?
- Does it appear in different representations?
- Does the learner succeed when the method is named but fail when it is not?
- Does performance change with time pressure?
- Does support level change the result?
- Does the pattern survive a fresh task later?
Sampling Is Different From More Testing
More questions do not automatically create better evidence.
Twenty nearly identical items may tell us less than six carefully chosen ones.
If every item is blocked by topic, the sample may never test recognition.
If every item uses the same representation, transfer remains unknown.
If every task is completed immediately after teaching, durability remains unknown.
Sampling is about coverage of the capability, not page count.
The Sampling Problem
Any observation contains two things:
- something about the learner;
- something about the particular task and conditions.
A learner can succeed because the item is familiar.
Fail because the wording is unusually confusing.
Guess correctly.
Make one arithmetic slip despite strong conceptual understanding.
That is why measurement design matters.
A 2025 systematic review in School Effectiveness and School Improvement examined the measurement quality of student outcomes and highlighted the importance of reliability and validity when educational conclusions are built from observed performance. The context is formal research, but the practical lesson is useful for tutoring: if the sample does not represent the capability well, the conclusion becomes fragile.
Sampling Across Items
Do not diagnose from one item when several item forms could expose the same capability.
Mira is suspected of weak simultaneous-equation method selection.
Sample:
- one case where elimination is obvious;
- one where substitution is efficient;
- one where both are reasonable;
- one word problem requiring equations to be formed first;
- one mixed item where the method is not named.
Now the sample tests method selection rather than one memorised routine.
Sampling Across Representations
A skill may look stable in one form and fragile in another.
For direct proportion, sample:
- equation;
- table;
- graph;
- word problem.
If Mira succeeds only with equations, the relationship may not yet be representation-independent.
The existing Representation Switching article owns the broader mechanism. Sampling simply asks whether the capability survives more than one representation before we declare it stable.
Sampling Across Support Levels
A learner may perform under prompts and fail independently.
Sample both.
Jonas answers an inference correctly after the tutor asks, “Which line gives you evidence?”
That is useful.
Now give a fresh passage without the cue.
If performance collapses, evidence selection remains externally triggered.
This connects to Training Cue Hierarchy and Training Independence.
Sampling Across Time
Immediate success can be real and still temporary.
Sample again after delay.
The learner should not need to reproduce the exact item.
Use a fresh item sampling the same underlying capability.
This distinguishes immediate lesson performance from more durable availability.
It also aligns with mastery-learning practice, where retesting and cumulative checks are used to determine whether learning remains sufficient to move forward.
Sampling Across Difficulty
A learner may be stable at one difficulty level and fragile at another.
Sample near the intended performance range.
If Nadia succeeds only on extremely clean experimental diagrams, the sample should include one realistic unfamiliar setup.
But do not jump immediately to maximum complexity.
The sample should reveal the current envelope, not manufacture failure.
Sampling Across Error Classes
Sometimes the question is not “Can the learner do this?”
It is:
Which kind of failure is recurring?
For Mathematics, sample opportunities for:
- recognition;
- method selection;
- execution;
- sign control;
- checking.
For English:
- literal retrieval;
- inference;
- evidence selection;
- claim strength;
- paragraph development.
For Science:
- concept retrieval;
- variable identification;
- data interpretation;
- mechanism explanation;
- evidence-bounded conclusion.
The sample becomes a map of where performance breaks.
One Question Can Still Be High Value
“One question is not a diagnosis” does not mean every diagnosis needs twenty questions.
A well-chosen item can sharply reduce uncertainty.
The existing How Tuition Works | The Diagnostic Probe owns that mechanism: one discriminating question can separate two possible causes.
Training Sampling owns the next question:
How much evidence is enough before we treat the pattern as the learner state rather than an isolated observation?
Probe to discriminate.
Sample to establish pattern.
Sampling and Training Baselines
Training Baselines needs a sampling strategy.
A baseline built from one easy item is fragile.
A baseline built from fifty redundant items is wasteful.
The right sample is the smallest set that represents the dimensions relevant to the training decision.
Sampling and Training Branching
Sampling can happen adaptively.
Use Training Branching.
Mira succeeds on two clean items.
Branch to a mixed item.
She fails.
Branch to a contrast pair.
The sample grows only where the evidence remains uncertain.
This is more efficient than giving every learner the same diagnostic battery regardless of what early responses reveal.
Sampling and Training Case Families
A Training Case Family is an ideal sampling reservoir because it can contain:
- clean examples;
- near-misses;
- representation changes;
- distractors;
- transfer cases;
- boundary cases.
The tutor does not need to use every case.
Select enough cases to answer the current diagnostic question.
Mathematics Sampling: Mira
The tutor suspects Mira has a quadratic-method selection problem.
Instead of another full paper, the tutor uses six short items.
- two factorisable quadratics;
- one awkward quadratic;
- one problem where completing the square is useful;
- one graph-to-equation connection;
- one mixed item where no method is named.
Mira executes methods accurately once selected.
Selection fails on three mixed cases.
The sample is enough to justify a narrow training decision.
More full-paper testing would add cost without much diagnostic value.
English Sampling: Jonas
Jonas appears weak at inference.
The tutor samples four passages:
- one narrative;
- one article;
- one speech;
- one passage containing an unfamiliar word.
Jonas succeeds on the first three and fails only when vocabulary blocks the key sentence.
The original hypothesis changes.
Inference is not the main weakness.
Vocabulary access is the dependency in that context.
Science Sampling: Nadia
Nadia is suspected of weak experimental reasoning.
The tutor samples:
- one familiar plant experiment;
- one unfamiliar heat setup;
- one flawed experiment;
- one data interpretation;
- one explanation task.
Nadia handles variable logic well but overclaims conclusions in three contexts.
The sampling pattern points to evidence calibration, not experimental design generally.
Adequate Sampling in Mastery Learning
A practical review of mastery learning recommends that formative assessments be appropriately blueprinted so they adequately sample the competencies being judged, and that retesting preserve comparable coverage of learning objectives.
This comes from higher-education literature and should not be copied mechanically into small-group tuition.
But the principle is useful:
If a decision depends on whether a capability is ready, sample the capability broadly enough that the decision is not resting on one narrow item.
Do Not Oversample
Sampling has a cost.
Time spent measuring is time not spent learning.
Stop sampling when the remaining uncertainty is unlikely to change the next training decision.
If six carefully chosen items already show that method selection is the bottleneck, another forty may not be useful.
The goal is sufficient evidence.
Not maximal measurement.
Do Not Undersample
The opposite error is equally common.
One wrong answer becomes:
You don’t understand fractions.
One successful answer becomes:
Done. Mastered.
Both conclusions may be premature.
Add one or two discriminating observations before changing the whole training programme.
Do Not Confuse Repetition With Independent Samples
The learner solves the same question three times.
That is useful repetition.
It is weak evidence about transfer because memory for the exact item may support later attempts.
For diagnosis, use fresh items sampling the same underlying capability.
Do Not Sample Only Easy Success
If all sample items are easier than the real performance demand, the learner state will look artificially strong.
Include at least some tasks near the intended level.
But do not turn sampling into a stress test unless robustness under pressure is the capability being measured.
The Parent Sampling Audit
- Are we making a large conclusion from one question?
- Have we seen the same pattern more than once?
- Were the tasks different enough to test the underlying skill?
- Did support level change the result?
- Did the pattern survive a fresh task later?
- Is more testing likely to change the next decision?
The Tutor Sampling Audit
- What hypothesis am I testing?
- Which small set of tasks best samples it?
- What item variation is necessary?
- What conditions should change or remain stable?
- How much evidence is enough for the next decision?
- Am I collecting redundant observations?
- What would falsify my current diagnosis?
The Deeper Idea: Diagnosis Is an Inference From Samples
No tutor sees the learner’s entire knowledge state directly.
We see attempts.
Answers.
Explanations.
Hesitations.
Errors.
Transfer.
From those samples we infer what is likely happening underneath.
Good training keeps that inference humble enough to update when new evidence arrives.
One question can open the diagnosis. A pattern across well-chosen questions is what makes the diagnosis worth acting on.
Research Foundations
Useful anchors include the 2025 systematic review of measurement quality in student outcomes, the practical mastery-learning review which discusses blueprinting and adequate sampling of competencies, and the wider educational-measurement literature on reliability, validity and task sampling. The practical claim here is simple: sampling should be representative enough to support the instructional decision, but compact enough that diagnosis does not displace learning.
Continue Through How Training Works
Read this alongside Training Baselines, Training Signals, Training Case Families, Training Branching, Training Dependencies and How Tuition Works | The Diagnostic Probe.
