Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

How High Performance Learning Works | Proxy Failure — When the Metric Improves but the Capability Does Not

Nadia’s score rose.

Her tutor was not satisfied.

That sounded unreasonable until he changed the questions.

For three weeks, Nadia had practised one narrow question family. She became fast. She recognised the surface cues. She learned which sentence frames earned marks. On the familiar worksheet, performance looked transformed.

Then the same underlying skill appeared in an unfamiliar format.

The score disappeared.

The metric had improved.

The capability had not travelled.

A measure is useful only while it still tells the truth about the thing we care about.

The 60-Second Route

Proxy failure occurs when a measure, indicator or stand-in for an underlying goal becomes a poor guide to that goal—especially after people begin optimising the proxy itself.

Learning is full of proxies because capability is partly invisible.

  • Marks proxy understanding and performance.
  • Questions completed proxy practice.
  • Minutes studied proxy effort.
  • Homework submission proxies consistency.
  • Vocabulary-list size proxies lexical knowledge.
  • Confidence ratings proxy self-knowledge.
  • Mock-exam scores proxy future examination performance.
  • Tutor prompts needed proxy independence.

These measures can be useful.

The danger begins when improving the measure becomes easier than improving the underlying capability.

The Proxy Is Not the Goal

A student does not ultimately need a high practice count.

They need learning.

A student does not ultimately need beautiful notes.

They need usable knowledge.

A student does not ultimately need high confidence.

They need calibrated judgement and capable action.

A student does not ultimately need one impressive mock score.

They need reliable examination performance.

Proxies matter because they help us observe progress toward those goals.

They become dangerous when we forget the direction of the relationship.

We chose the number because it tracked the capability. We should not redefine the capability as whatever makes the number rise.

Goodhart’s Law, Campbell’s Law and Proxy Failure

A family of ideas across economics, social science and measurement warns about this problem.

Goodhart’s law is commonly summarised as the idea that when a measure becomes a target, it tends to become a worse measure.

Campbell’s law makes a related point about quantitative indicators used for consequential social decision-making: greater pressure on an indicator creates pressure to corrupt the indicator and distort the process it was meant to monitor.

A 2024 target article in Behavioral and Brain Sciences, Dead rats, dopamine, performance metrics, and peacock tails: Proxy failure is an inherent risk in goal-oriented systems, argues for a broad unifying perspective: when selection or incentives are based on an imperfect proxy for the true goal, pressure arises that can make the proxy diverge from the goal.

This article applies that systems idea to individual learning and tuition.

Proxy Failure Is Not “Metrics Are Bad”

Without metrics, learning decisions can become vague and impressionistic.

Scores, timing, retrieval rates, error counts and completion data can reveal genuine progress.

The correct response to proxy failure is therefore not to abandon measurement.

It is to use measures with humility.

  • Know what each measure can and cannot represent.
  • Use more than one measure when the goal is broad.
  • Retest under changed conditions.
  • Watch for behaviour adapting to the metric.
  • Periodically audit whether the measure still correlates with the capability.

Proxy Failure Is Not Construct Validity

Construct validity asks whether a measure represents the construct we think it represents.

Proxy failure includes a dynamic element.

A measure can begin as a reasonable indicator and become worse after it is targeted.

For example, a mock exam score may initially provide useful information about performance. Once every practice session trains specifically on the same paper style and familiar item patterns, the score may rise faster than general capability.

The measure did not necessarily start invalid.

Optimisation changed the relationship.

Proxy Failure Is Not Local Optimisation

Local Optimisation asks whether one component improves while the whole performance worsens.

Proxy failure asks whether the thing being optimised is still a trustworthy stand-in for the thing we care about.

The proxy can improve even if no obvious component gets worse.

A student simply learns to perform better on the measurement instrument itself.

The gap is epistemic before it is operational:

The number says “better.” Is the capability actually better?

Proxy Failure Is a Second-Order Effect

The earlier Batch 16 article on Second-Order Effects explains why behaviour changes after an intervention changes the system.

Proxy failure is one specific adaptation.

Once the measure matters, learners and adults naturally organise behaviour around it.

That can improve the underlying capability when the proxy is well aligned.

It can also create a wedge when the easiest way to improve the number differs from the best way to improve the goal.

The Proxy Gap

Define:

Goal: the capability we actually care about.

Proxy: the observable measure we use to estimate it.

Proxy gap: the difference between improvement in the measure and improvement in the underlying capability.

The proxy gap can widen because:

  • practice becomes too similar to the test;
  • students learn superficial cues;
  • measurement conditions become easier;
  • unmeasured skills are neglected;
  • people game the metric;
  • the goal changes while the measure stays the same;
  • the measure becomes saturated at a ceiling.

The Transfer Test Closes the Proxy Gap

One of the strongest ways to test whether a metric represents capability is to change the surface conditions.

Same skill.

Different question.

Different representation.

Different context.

Different order.

Cold start.

If the gain survives, the proxy is more likely tracking a genuine capability.

If the gain disappears, the measure may have become too tightly coupled to the training format.

The Independence Test

Another powerful proxy audit is to remove support.

A student scores highly with tutor prompts.

What happens without them?

A writer produces strong paragraphs using a visible template.

What happens cold?

A learner completes homework when a parent starts the session.

What happens when the cue is absent?

Supported performance is not fake.

It measures supported capability.

The error occurs when supported performance is interpreted as independent capability.

The Delay Test

Immediate performance can proxy learning poorly when familiarity is high.

A student completes ten questions immediately after an explanation and scores 90%.

That shows successful immediate performance.

Does it show durable learning?

Retest after delay.

Delay removes some short-lived contextual support and reveals whether the knowledge remains accessible.

The Mixed-Practice Test

Blocked practice can produce impressive scores because the method has already been selected by the page.

A student solving twenty simultaneous-equation questions does not need to decide whether each is a simultaneous-equation problem.

The exercise may measure execution well and selection poorly.

Mixed practice restores the missing decision.

If performance falls sharply, the blocked score was a proxy for a narrower capability than adults may have assumed.

The Parallel-Form Test

Repeated use of one test form creates familiarity.

Use a fresh parallel form that measures the same skill through different items.

If the gain survives, confidence increases that the measured capability is general.

If performance drops, investigate whether the learner memorised item patterns rather than the underlying structure.

Marks as a Proxy

Marks matter.

They summarise performance under a defined assessment.

But a mark is not identical to learning.

It reflects:

  • knowledge;
  • question sampling;
  • difficulty;
  • marking rules;
  • time;
  • state on that day;
  • examcraft;
  • sometimes luck.

A single mark should therefore be treated as evidence, not a complete learner model.

When Marks Become the Target

Once marks dominate every decision, behaviour can narrow.

Students avoid untested reading.

Teachers emphasise high-frequency item types.

Practice focuses on the scoring surface.

This can be rational near a high-stakes examination.

The danger is forgetting the scope of the proxy.

Exam scores measure examination performance on sampled content and skills.

They should not silently become the definition of education.

Practice Count as a Proxy

“Do fifty questions.”

The instruction is easy to monitor.

Question count can proxy exposure and practice volume.

But once count becomes the target, learners can improve it by:

  • choosing easier questions;
  • skipping review;
  • guessing quickly;
  • repeating familiar formats;
  • avoiding slow diagnostic problems.

The metric rises.

Learning per question may fall.

Replace Count With a Portfolio of Evidence

Instead of one count, use several signals:

  • questions attempted;
  • accuracy;
  • error type;
  • difficulty;
  • cold retrieval;
  • transfer;
  • time;
  • independent correction.

No single number becomes the whole goal.

Study Time as a Proxy

Minutes studied proxy effort and opportunity for learning.

They do not directly measure learning.

When hours become the target, students may:

  • sit longer with lower attention;
  • avoid breaks that would improve quality;
  • choose activities that feel easy to sustain;
  • equate exhaustion with productivity.

A better measure combines time with output quality and later retrieval.

Forty focused minutes that produce durable retrieval can be more valuable than two hours of passive exposure.

Homework Completion as a Proxy

Completion matters because unfinished work can signal weak routines or lost practice.

But completed homework does not necessarily mean understood homework.

The proxy can be gamed unintentionally through copying, answer-checking too early or dependence on help.

A more useful completion metric distinguishes:

  • independent attempt;
  • help required;
  • correction quality;
  • cold reattempt;
  • remaining uncertainty.

The goal is not merely a finished page.

Neat Notes as a Proxy

Neatness can improve usability.

Organisation can reduce search cost.

But visual quality is easy to observe, so adults may overvalue it.

A student can create beautiful notes that produce weak retrieval.

A messy-looking retrieval sheet may reveal far more learning.

Evaluate notes by what they enable:

  • finding key information;
  • reconstructing relationships;
  • supporting retrieval;
  • revealing gaps;
  • connecting topics.

Vocabulary Size as a Proxy

A learner knows five hundred advanced words.

What does “knows” mean?

  • recognises them?
  • can define them?
  • can retrieve them?
  • can use them naturally?
  • knows register and collocation?
  • can distinguish near-synonyms?

List size is a proxy whose meaning depends on the test.

If the target is expressive writing, production and contextual precision deserve direct testing.

Reading Speed as a Proxy

Reading speed can proxy fluency.

But faster is not always better.

If speed becomes the target, students may skim relationships, miss contrast and reduce inference quality.

The appropriate metric is conditional.

Can the learner maintain comprehension and adjust speed to text difficulty and purpose?

Adaptive pacing is closer to the capability than raw words per minute.

Writing Length as a Proxy

Longer answers often contain more opportunity for development.

Then students learn that longer looks stronger.

Proxy failure appears when they add:

  • repetition;
  • irrelevant examples;
  • decorative description;
  • unnecessary qualifiers;
  • empty transitions.

The word count rises.

Argument quality may not.

Measure function.

Does each paragraph advance the job?

Sophisticated Vocabulary as a Proxy

Advanced vocabulary can proxy language maturity.

Once targeted directly, students may insert rare words where simpler precise words would be stronger.

The proxy failure is visible in register, collocation and semantic precision.

The real goal is controlled expressive range.

A sophisticated word that damages meaning is negative performance, not evidence of vocabulary capability.

Speed as a Proxy in Mathematics

Speed matters in timed examinations.

It can also proxy automaticity.

But if speed becomes the dominant target, students may skip modelling, compress high-risk steps and choose familiar methods before classification.

The better capability is efficient correct performance:

  • fast where routine;
  • slow where branching risk is high;
  • able to recover when the first route fails.

Raw speed is a proxy for only one part of that system.

Error Count as a Proxy

Fewer errors sounds unambiguously good.

Then imagine two learners.

One attempts hard transfer questions and makes five errors.

Another repeats easy familiar questions and makes one.

The second error count is lower.

It does not prove stronger learning.

Error metrics need challenge context.

Sometimes a temporary rise in errors is evidence that the learner has entered a more informative practice zone.

Confidence as a Proxy

Confidence can proxy readiness and metacognitive judgement.

But confidence itself can be trained independently of capability.

Generic encouragement may raise confidence.

If accuracy and calibration do not improve, the proxy gap widens.

The high-performance target is not “feel confident.”

It is “act with confidence proportional to evidence.”

Help Requests as a Proxy

Fewer help requests might appear to indicate independence.

But students can reduce help-seeking by giving up, hiding confusion or avoiding difficult work.

More help requests might indicate dependency—or better metacognitive awareness.

The measure needs context.

A better capability measure asks whether the learner can:

  • attempt independently;
  • locate the blockage;
  • ask for the smallest useful help;
  • resume independently afterward.

Mock Exam Scores as a Proxy

Mock exams are strong proxies when they are representative.

The proxy weakens when:

  • items are familiar;
  • conditions differ materially from the final examination;
  • marking is inconsistent;
  • students receive hints;
  • content coverage is narrower;
  • the learner has practised the exact paper repeatedly.

A good mock does not imitate theatre.

It samples the capability under sufficiently representative conditions.

One Good Score as a Proxy for Reliability

A peak score shows capability.

It does not by itself show reliability.

This is why Performance Reliability asks whether the performance can be repeated.

Use distributions, not highlights.

Peak, median, range and catastrophic lows tell different stories.

The Score Inflation Problem

A score can rise because capability rises.

It can also rise because training becomes better matched to the specific measure.

In education policy, research on high-stakes testing has long examined score inflation and “teaching to the test.” Brookings has summarised this as a classic Campbell’s-law problem: when test scores carry strong accountability stakes, organisations may respond in ways that inflate the targeted scores or narrow the curriculum without equivalent gains in broader learning.

The individual-learning analogue is straightforward.

If practice becomes increasingly identical to the measurement, the score may become a better measure of familiarity with the instrument than of general capability.

The Teaching-to-the-Test Boundary

Preparing for a known examination is not automatically proxy failure.

Students should know the format, mark allocation and performance conditions.

Examcraft is a legitimate capability when the goal includes succeeding in that examination.

The proxy problem begins when narrow familiarity substitutes for the underlying knowledge and transfer the examination intends to sample.

Therefore:

  • teach the format;
  • train the underlying capability;
  • use fresh items;
  • mix contexts;
  • separate exam-specific strategy from subject understanding.

Proxy Failure and Practice Specificity

Practice Specificity says practice should resemble the target performance enough to prepare it.

Proxy failure creates the counter-pressure.

If practice is too specific to one measurement form, performance can become instrument-bound.

The solution is a two-layer design:

  1. representative practice for the real performance;
  2. variation that proves the underlying capability is not trapped in one surface form.

Proxy Failure and Transfer Distance

The farther a gain travels, the harder it is to explain as pure test familiarity.

If a method improves:

  • the trained item;
  • a fresh parallel item;
  • a changed representation;
  • an unfamiliar context;
  • a later cold test;

confidence grows that the underlying capability improved.

Transfer Distance is therefore one proxy audit.

Proxy Failure and Calibration

A metric can distort self-belief.

If a student repeatedly scores highly on familiar blocked practice, confidence rises.

Then mixed cold performance disappoints.

The learner appears overconfident.

The deeper problem may be that the proxy used to build confidence was too narrow.

Calibration depends on measurement quality.

Proxy Failure and Model Parsimony

A single metric is attractive because it simplifies decision-making.

“Score up means learning up.”

“Hours up means effort up.”

“Questions up means practice up.”

But simplicity can become underfitting.

Model Parsimony reminds us to use the simplest model that still fits the important evidence—not the simplest number available.

Proxy Failure and Reference Classes

A proxy may look impressive until compared with similar cases.

A learner completes 100 questions this week.

Is that good?

Compared with weeks of 50 high-quality mixed questions, perhaps not.

A mock score rises ten marks.

Did similar rises on familiar papers predict future cold performance?

Reference Class Reasoning supplies context that keeps proxies from floating free of outcomes.

Proxy Failure and Error Detectability

A good measurement system makes proxy divergence detectable.

If blocked-practice accuracy rises while mixed accuracy does not, the discrepancy is a signal.

If study hours rise while delayed retrieval remains flat, the discrepancy is a signal.

If vocabulary-list scores rise while composition usage remains unchanged, the discrepancy is a signal.

Error Detectability applies at the measurement-system level too.

The Multi-Measure Defence

One defence against proxy failure is triangulation.

Use measures that fail differently.

For a Mathematics skill:

  • blocked accuracy;
  • mixed selection accuracy;
  • cold retrieval;
  • time;
  • transfer to a new context;
  • independent error detection.

For writing:

  • task fulfilment;
  • idea structure;
  • evidence;
  • language accuracy;
  • timed completion;
  • performance on an unseen prompt.

When several independent measures move together, confidence in real improvement rises.

But Multiple Metrics Can Be Gamed Too

Adding more metrics is not a universal solution.

If all measures become explicit targets, students can optimise the dashboard instead of the underlying capability.

Keep some measures diagnostic rather than incentivised.

Use surprise fresh tests.

Rotate item forms.

Preserve professional judgement.

Metrics support judgement.

They do not eliminate the need for it.

The Hidden-Metric Defence

Sometimes the strongest validation measure should not be the same thing the learner is optimising directly.

Train with known practice objectives.

Validate with fresh transfer.

Train essay structure.

Validate on an unseen prompt.

Train algebraic fluency.

Validate inside a mixed modelling problem.

Train vocabulary.

Validate through reading and writing usage.

The validation task checks whether the capability escapes the training proxy.

The Metric Rotation Defence

Metrics should change as the bottleneck changes.

Early in learning, accuracy may matter most.

Later, delayed retrieval.

Then selection.

Then transfer.

Then speed under representative conditions.

Keeping one metric forever encourages over-optimisation after its diagnostic value falls.

The Threshold Defence

Sometimes a metric should be treated as a threshold rather than maximised.

Once handwriting is readable, making it prettier may have little learning value.

Once routine accuracy exceeds a stable threshold, more blocked practice may have diminishing returns.

Once attendance is consistent, the focus should shift to what happens during study.

Thresholds reduce pressure to game already-sufficient proxies.

The Balanced-Scorecard Defence

For a broad goal, track several dimensions with explicit trade-offs.

An examination-performance scorecard might include:

  • knowledge accuracy;
  • completion;
  • method selection;
  • error recovery;
  • timing;
  • transfer;
  • variance across papers.

If one rises while three fall, the system should not celebrate blindly.

The Capability Ledger

Instead of tracking only outcomes, track capability states.

For each important skill:

  • understood?
  • retrievable cold?
  • executable accurately?
  • selectable among alternatives?
  • transferable?
  • stable after delay?
  • independent?
  • robust under time pressure?

A final score is then interpreted through the capability ledger rather than standing alone.

The Proxy Audit

  1. What underlying capability do we actually care about?
  2. Why was this metric chosen as a proxy?
  3. Under what conditions does the metric correlate with the capability?
  4. How could the metric improve without the capability improving?
  5. What behaviour is the metric likely to incentivise?
  6. Has that behaviour started appearing?
  7. What fresh transfer test can validate the gain?
  8. What independent measure fails differently?
  9. Has the proxy saturated?
  10. Has the learning goal changed?
  11. Are we rewarding the measure directly?
  12. Could the learner game it without intending to deceive?
  13. What unmeasured skill might be crowded out?
  14. What threshold would be sufficient rather than maximal?
  15. When should this metric be retired or rotated?

The Parent Version: Ask What the Number Is Standing In For

Parents see many numbers.

Marks.

Percentiles.

Homework counts.

Study hours.

Vocabulary counts.

Before reacting, ask:

What is this number supposed to tell us about the child?

Then ask what other evidence should agree if the interpretation is true.

If the mark rose because understanding improved, fresh unseen questions should usually show some improvement too.

If study hours rose because discipline improved, independent task completion should become more reliable.

Look for convergence.

The Tutor Version: Never Let the Training Metric Become the Curriculum

Tutors naturally create metrics.

Accuracy.

Time.

Error count.

Questions completed.

These help organise intervention.

But the metric should follow the diagnostic job.

When the job changes, the metric should change.

Otherwise the learner becomes excellent at yesterday’s measurement problem.

The Student Version: Do Not Study the Dashboard

Students can become obsessed with trackers.

Streaks.

Hours.

Question counts.

Scores.

Use the dashboard to steer.

Do not make the dashboard the destination.

Every week, include at least one check that asks the capability directly.

Can I do it cold?

Can I do it in a new form?

Can I explain why?

Can I recover if I make an error?

The Anti-Gaming Design

Good measurement systems are harder to game because improvement requires more of the true capability.

Examples:

  • fresh items rather than repeated items;
  • mixed questions rather than labelled blocks;
  • transfer tasks rather than exact replicas;
  • cold starts rather than immediate post-teaching tests;
  • independent attempts rather than prompted success;
  • multiple representations rather than one familiar format.

The objective is not to trick the learner.

It is to ensure the easiest route to a better metric passes through genuine capability.

The Anti-Corruption Principle

A proxy is strongest when optimising it improves the goal by design.

If the metric is “percentage of fresh mixed questions solved independently after one week,” it is harder to raise without improving retrieval, selection and transfer than a metric such as “questions completed today.”

No measure is perfect.

But some are more aligned with the capability and less gameable.

The Metric Hierarchy

Use different metrics for different levels.

Activity metrics: minutes studied, questions attempted, sessions completed.

Process metrics: independent starts, error-detection rate, quality of explanations, checking allocation.

Capability metrics: delayed retrieval, mixed selection, transfer, performance reliability.

Outcome metrics: school assessments and final examinations.

A strong learning system does not confuse levels.

Activity supports process.

Process builds capability.

Capability supports outcomes.

But the chain is not perfect, so validation remains necessary.

The Leading and Lagging Indicator Distinction

Some proxies are useful precisely because they appear before the final outcome.

Retrieval accuracy this week can be a leading indicator.

The examination mark months later is a lagging outcome.

Leading indicators are valuable for steering.

But because they are proxies, they need periodic validation against lagging outcomes and transfer.

When the Proxy Becomes the Product

One warning sign is language change.

“We need to improve her Mathematics” becomes “We need to get the worksheet accuracy above 90%.”

“We need stronger writing” becomes “We need 800 words.”

“We need consistent study” becomes “We need two hours every night.”

The operational measure has replaced the underlying goal in the sentence.

That is a useful moment to audit.

When a Proxy Is Good Enough

Perfect measurement is impossible.

We often need a cheap proxy.

A ten-question quiz may be enough to decide whether a foundation needs repair.

A study-time log may be enough to reveal chronic under-allocation.

A homework-completion rate may be enough to identify a consistency problem.

The question is not whether the proxy is perfect.

It is whether it is accurate enough for the decision being made.

The Decision-Specific Proxy

A metric can be adequate for one decision and inadequate for another.

Blocked accuracy may be enough to decide whether basic execution is emerging.

It is not enough to decide whether the student is examination-ready.

A mock exam may be useful for pacing.

It may be insufficient to diagnose one conceptual misconception.

Match the proxy to the decision.

The Proxy Failure Ladder

  1. Useful proxy: measure tracks the capability sufficiently for the decision.
  2. Targeted proxy: learners begin optimising the measure.
  3. Adaptive response: behaviour shifts toward what raises the metric.
  4. Divergence: some metric gains no longer require equal capability gains.
  5. Corruption: the measure substantially loses its meaning.
  6. Audit: fresh transfer or independent validation reveals the gap.
  7. Repair: rotate, broaden or redesign the measure.

Not every targeted proxy reaches the later stages.

The purpose of the framework is to notice divergence early.

The Proxy Failure Matrix

Evaluate a metric on two dimensions.

Alignment: how closely does the metric represent the capability?

Gameability: how easily can the metric rise through behaviour that does not improve the capability?

High alignment + low gameability: strong operational metric.

High alignment + high gameability: useful diagnostically but risky as a target.

Low alignment + low gameability: weak but perhaps harmless.

Low alignment + high gameability: dangerous target.

Case: The 100-Question Week

Evan set a goal of one hundred Mathematics questions.

Week one: he completed eighty difficult mixed questions and reviewed them carefully.

Week two: determined to hit the target, he chose shorter familiar questions and completed one hundred and twelve.

The dashboard improved.

His delayed mixed test did not.

The family changed the metric.

Instead of raw count, they tracked three things:

  • representative question volume;
  • error repair completion;
  • cold mixed accuracy.

The count remained useful.

It stopped being sovereign.

Case: The Perfect Homework Record

Mira submitted every homework assignment for six weeks.

The record looked excellent.

Then a cold quiz revealed that many solutions had been completed with heavy answer-checking.

Submission was a valid measure of consistency.

It was an invalid stand-in for independent mastery.

The solution was not to stop valuing submission.

The family added a small cold reattempt from the week’s homework.

Now the measurement system could distinguish completed from learned.

Case: The Vocabulary Champion

Nadia could define hundreds of advanced words.

Her writing still used a much smaller active vocabulary.

The recognition test had become a proxy for productive vocabulary.

The programme added production checks:

  • write a natural sentence;
  • choose the right collocation;
  • distinguish a near-synonym;
  • use the word in an appropriate register;
  • retrieve it without seeing the list.

The vocabulary score initially fell.

The metric became harder.

The capability measure became more honest.

Case: The Fast Reader

Jonas increased reading speed dramatically.

Then inference accuracy fell.

Words per minute had become a proxy for fluency while comprehension was treated as constant.

It was not constant.

The new metric became adaptive reading:

  • speed on straightforward text;
  • accuracy on inference;
  • ability to slow at structural difficulty;
  • ability to summarise relationships afterward.

Speed remained part of the goal.

It no longer defined the goal.

Case: The High-Confidence Student

A programme tried to improve a student’s confidence through frequent praise and easier early successes.

The confidence rating rose.

When mixed difficult tasks returned, performance did not change enough.

Confidence had been successfully targeted as a proxy for readiness.

The repair paired confidence with calibration.

After each task, the student predicted performance and compared the prediction with outcome.

Confidence became evidence-responsive.

Case: The Mock-Exam Improvement

Evan’s mock score improved from 62 to 78.

The first reaction was celebration.

The second reaction was verification.

The paper contained several item structures he had practised repeatedly.

On a fresh parallel form, he scored 71.

The improvement was still real.

It was smaller than the original proxy suggested.

The fresh test did not erase the gain.

It calibrated it.

Proxy Failure and Path Dependence

The preceding article on Path Dependence explains how repeated choices shape later option costs.

Metrics can create paths.

If a student spends years optimising grades through narrow test familiarity, they may accumulate fewer transfer experiences.

If a student tracks retrieval and transfer, they build a different learning path.

What gets measured repeatedly can become what gets practised repeatedly.

Proxy Failure and Stability–Plasticity

Metrics stabilise behaviour.

If the metric is well aligned, that stability helps.

If the metric has drifted away from the goal, the learner can become highly stable at producing the wrong signal.

Stability–Plasticity therefore applies to measurement systems too.

Retain metrics that remain truthful.

Update metrics when the goal, bottleneck or behaviour changes.

The Metric Sunset Rule

No training metric should live forever by default.

At defined intervals, ask:

  • Is this still the bottleneck?
  • Does the metric still correlate with fresh performance?
  • Has the learner learned to game it?
  • Has the metric reached a ceiling?
  • Would another measure now be more diagnostic?

If the answers change, rotate the metric.

The Proxy Failure Early-Warning Signs

  • The metric improves unusually quickly while broader performance does not.
  • Students can explain how to raise the number without discussing learning.
  • Performance collapses on fresh items.
  • Unmeasured activities disappear from the schedule.
  • The metric reaches a ceiling but capability still has room to grow.
  • Adults argue about the number more than the underlying work.
  • The learner becomes anxious when the metric cannot be recorded.
  • Qualitative evidence and the metric increasingly disagree.
  • Metric improvement requires increasingly artificial conditions.
  • The measure predicts less well than it used to.

The Proxy Repair Ladder

  1. Restate the true goal.
  2. Identify why the current proxy was chosen.
  3. Document the divergence.
  4. Find the behaviour producing metric gains without capability gains.
  5. Add one independent validation measure.
  6. Reduce stakes on the corrupted proxy.
  7. Rotate or redesign the metric.
  8. Use fresh transfer to reconnect measurement with capability.
  9. Monitor for new gaming.
  10. Retire the metric when it is no longer needed.

Build Metrics That Fail Differently

If two measures share the same failure mode, using both adds little protection.

Two nearly identical worksheets are not independent validation.

Two confidence scales are not independent validation.

Better combinations include:

  • timed performance + cold transfer;
  • recognition + production;
  • accuracy + explanation;
  • teacher-scored work + independent fresh task;
  • activity count + delayed retrieval.

Independent failure modes make divergence easier to detect.

Do Not Incentivise Every Diagnostic Measure

A diagnostic measure can lose information when rewards become attached.

If students know every confidence rating is judged, they may stop reporting uncertainty honestly.

If every error count affects rewards, students may avoid difficult work.

Some measures should remain low-stakes sensors.

Measurement and incentive are different design decisions.

Do Not Punish Honest Uncertainty

If the system rewards only confident answers, uncertainty becomes costly to report.

Then confidence data become corrupted.

Students should be able to say:

I think this is right, but the final inference is less secure than the evidence location.

That is high-resolution metacognition, not weakness.

Do Not Punish Productive Errors

If every error lowers the visible performance score, learners may avoid exploration.

Training should distinguish:

  • errors during acquisition;
  • errors during diagnostic challenge;
  • errors during final performance.

The same error count means different things in different contexts.

Do Not Let the Measurement Change the Task Too Much

Measurement itself can alter behaviour.

A learner who knows every pause is timed may rush.

A learner who knows every explanation is scored may produce performative reasoning.

A learner who knows every study minute is logged may keep the timer running while attention drifts.

Keep measurement lightweight enough that it does not become the main task.

Proxy Failure in AI-Assisted Learning

AI-assisted work introduces new proxies.

A polished answer can proxy understanding.

A completed draft can proxy writing capability.

A correct solution can proxy mathematical reasoning.

When tools can generate high-quality outputs, output quality alone becomes a weaker measure of the learner’s independent capability unless the task explicitly allows that support as part of the target skill.

The solution is to define the target.

If the goal is tool-assisted production, evaluate tool use, judgement and final quality.

If the goal is independent examination performance, validate without the tool.

The proxy must match the performance environment.

Proxy Failure and Cognitive Offloading

Cognitive Offloading distinguishes productive tool use from outsourcing the thinking that the learner actually needs.

Proxy failure explains why this matters for measurement.

If a tool produces the observable performance, the output may no longer proxy the learner’s internal capability in the same way.

Again, this is not automatically bad.

It means the interpretation of the metric must change.

Measurement Invariance Across Support Conditions

When comparing performance over time, keep support conditions visible.

A 75% score with heavy prompting is not directly comparable to a 72% independent score.

The second may represent greater capability despite the lower number.

Record the support state:

  • independent;
  • minimal cue;
  • guided;
  • worked example available;
  • open notes;
  • tool-assisted.

Otherwise the metric can mislead through changing conditions rather than gaming.

Proxy Failure and Measurement Noise

Not every disagreement between metric and capability is proxy failure.

One score may simply be noisy.

A difficult paper can lower marks.

A lucky item set can raise them.

Proxy failure is more convincing when a systematic pattern appears:

  • targeted metric rises;
  • independent validation does not;
  • behaviour has adapted toward the metric;
  • the divergence persists across repeated observations.

The Metric-to-Mechanism Check

Whenever a metric improves, ask what mechanism changed.

Score rose because:

  • knowledge improved?
  • retrieval improved?
  • question familiarity improved?
  • time allocation improved?
  • marking became easier?
  • support increased?
  • items became more predictable?

The same numerical gain can represent different learning.

Mechanism keeps the metric interpretable.

The Proxy-Robust Learning Goal

Some goals are harder to corrupt because they bundle several dimensions.

Instead of “score 90% on the worksheet,” use:

Solve fresh mixed questions independently, explain the method choice, detect major errors and retain the skill after delay.

This goal is harder to measure with one number.

It is closer to the capability.

The Dashboard With No Single King

A robust learning dashboard has no single metric that can overrule every other signal.

Marks matter.

Transfer matters.

Independence matters.

Reliability matters.

Time matters.

The weighting changes with the learning stage.

Early acquisition may tolerate low speed.

Near an examination, speed becomes more important.

After the examination, breadth and transfer can rise again.

Dynamic weighting reduces permanent proxy capture.

Research: Proxy Failure as a General Systems Risk

The 2024 Behavioral and Brain Sciences target article Dead rats, dopamine, performance metrics, and peacock tails reviews Goodhart’s law, Campbell’s law and related phenomena across domains. Its central systems claim is that when an imperfect proxy becomes the basis of selection or incentive, pressure tends to arise that makes the proxy a worse approximation of the true goal.

The authors explicitly use standardised testing in education as one example: if test scores become the target, teaching can become increasingly optimised for the test, weakening the relationship between test performance and broader educational quality.

Research: Goodhart’s Law in Education

An ERIC-indexed paper, Goodhart’s Law and Performance Indicators in Higher Education, discusses how performance indicators can distort behaviour and argues that process indicators grounded in underlying processes may sometimes be less vulnerable than easily accessible outcome indicators.

At school-system level, Brookings has summarised evidence on test-based accountability through the lens of Campbell’s law, including inappropriate test preparation, strategic testing behaviour and reallocation of instructional time toward measured subjects or content.

These institutional examples are larger and more consequential than an individual student’s study tracker, but the mechanism is recognisable: a visible measure can reorganise behaviour around itself.

Research: Testing Can Still Be Valuable

Proxy-failure warnings should not be misread as evidence that testing is useless. Testing provides information and can strongly support learning when used appropriately. Even research on high-stakes accountability does not imply that every test score is corrupted or meaningless; some studies find strong correlations between high- and low-stakes measures in certain settings.

The correct lesson is conditional.

Measures become risky when stakes and optimisation pressure rise faster than alignment with the underlying goal.

The Proxy Failure Test

  1. What is the true capability or goal?
  2. What proxy are we using?
  3. Why should the proxy correlate with the goal?
  4. Can the proxy rise without equivalent capability growth?
  5. Has the learner learned how to do that?
  6. Are incentives attached to the metric?
  7. Has behaviour narrowed around the metric?
  8. Does fresh transfer confirm the gain?
  9. Does performance survive delay?
  10. Does it survive removal of prompts or support?
  11. Does it survive mixed selection?
  12. Do independent measures agree?
  13. Has an unmeasured capability been crowded out?
  14. Is the metric at a ceiling?
  15. Has the learning stage changed?
  16. Would a threshold be better than maximisation?
  17. Should the metric rotate?
  18. Can some validation remain low stakes?
  19. Does the measurement system preserve honesty about uncertainty and errors?
  20. Are we still improving the learner—or mainly improving the dashboard?

Batch Sixteen: Adapt, Anticipate, Preserve Options and Measure the Right Thing

  1. Stability–Plasticity: learn new things without unnecessarily overwriting what already works.
  2. Second-Order Effects: ask what an improvement changes next.
  3. Path Dependence: recognise how early choices alter later option costs without treating history as destiny.
  4. Proxy Failure: keep metrics connected to the capabilities they are supposed to represent.

Together they describe a learner and learning system that can change without becoming unstable, anticipate downstream consequences, preserve future options and resist the seductive mistake of confusing a rising number with a growing mind.

Research Notes and Evidence Boundary

Proxy failure is an established systems concept with close links to Goodhart’s law, Campbell’s law, goal displacement and measurement reactivity. The 2024 Behavioral and Brain Sciences article linked above provides a cross-disciplinary synthesis. Education research and policy analysis have documented examples where high stakes around test scores can alter instructional behaviour, score validity or participation patterns.

The student-level applications in this article—question counts, study hours, vocabulary counts, confidence ratings, mock scores, prompt dependence and similar metrics—are eduKatePunggol synthesis. They should not be interpreted as claims that every metric inevitably becomes corrupted. A proxy remains valuable when it is sufficiently aligned with the decision, when gaming is limited, and when fresh independent validation continues to agree.

Series Note

“High performance learning” is used descriptively throughout this eduKatePunggol series. The series does not claim affiliation with or reproduce any third-party branded educational framework using similar terminology.

Continue from here: Start Here · Tuition · Education · Pathways · Parenting 101 · All Site Routes

eduKate Punggol

Contact

83 Punggol Central, Singapore 828761

edu|Kate Bukit Timah

8 Fourth Avenue, Singapore 268674

By Appointment +65 8823 1234
admin@edukatesg.com

Email Us

When a child finally understands, school becomes less frightening and the future opens wider. Email us for the latest schedules and fees.

← 返回

感谢您的回复。 ✨

了解 eduKate Punggol 的更多信息

立即订阅以继续阅读并访问完整档案。

继续阅读