Science Education Systems · Transportability. Maya, Jia Jun, Hana and Ethan remain fictional Punggol learners. This article owns one narrow scientific job: how an effect estimated in one study population can be moved, cautiously and explicitly, to a different target population. The broader question of whether findings travel across people, settings, times and scales remains with How Scientific External Validity Works.
Wait, What? A study can be right and still give the wrong answer for the people you care about
Imagine a well-run study. The treatment was randomised. The measurements were careful. The analysis was correct. The estimated effect for the people who took part is trustworthy.
Now change the question.
Instead of asking, “What happened in this study?”, ask:
“What would the average effect be in a different population?”
That is not automatically the same number.
The people in the target population may differ in age, baseline risk, prior knowledge, disease severity, school context, geography, language, implementation conditions or other characteristics that modify the effect. A result can therefore be internally valid yet still require additional work before it is useful somewhere else.
Scientific transportability is the formal attempt to do that additional work.
The 50-second route
Transportability starts by naming two populations clearly:
source study population → target population
Then ask:
target definition → effect modifiers → source/target differences → overlap → measurement comparability → identification assumptions → transport method → diagnostics → sensitivity analysis → target effect estimate → uncertainty → validation → scope statement
The shortest useful rule is:
Do not move a result merely because two populations look similar. Move it only when the causal differences that matter are measured, sufficiently represented and handled transparently.
1. External validity, generalisability and transportability are related—but not identical
These terms overlap in ordinary scientific conversation, and the literature does not use them with perfect uniformity. A useful working distinction is this:
- External validity is the broad question of whether an inference applies beyond the original study.
- Generalisability often refers to extending a study result to a broader population from which the study sample could be regarded as a subset.
- Transportability usually refers to extending a causal effect estimate to a target population that is at least partly distinct from the study population.
This distinction is used in modern methodological reviews and guidance, although authors sometimes draw the boundary differently. The important habit is therefore not to fight over terminology. It is to state exactly what population produced the evidence and exactly what population you want the conclusion to describe.
For the broader conceptual owner, return to External Validity. This page goes narrower: it asks how to estimate a target-population effect rather than merely how to discuss whether a finding “travels”.
2. The target population comes first, not last
A transportability analysis cannot begin with “everyone else”.
The target must be defined before the method is chosen.
In medicine, the target might be adults with a particular condition treated in ordinary clinical practice. In education, it might be all schools of a defined type that could realistically implement a programme. In engineering, the target might be a fleet operating under a specified environment and duty cycle. In public policy, it might be residents affected by a particular programme in a defined jurisdiction.
A vague target creates a vague estimand.
That matters because the average effect depends on who is being averaged over.
The United States Institute of Education Sciences makes this point especially clearly for education research: define the target population of students and schools first, identify moderators that may change impact, and let that definition guide sampling, analysis and reporting.
3. Transportability is about an estimand, not merely a resemblance judgement
Suppose a randomised trial estimates an average treatment effect in its participants.
Call that the study average treatment effect.
But a decision-maker may care about the average effect in a different population. That target quantity is often called a population average treatment effect, or PATE.
The two quantities need not be identical.
If treatment effects vary across people and the composition of the source and target populations differs on those effect-modifying characteristics, the source average and target average can diverge even when the trial itself is flawless.
This is why transportability is a causal estimation problem rather than a decorative discussion section at the end of a paper.
4. The first question is not “How different are the populations?”
It is tempting to compare every available variable and count differences.
That can be misleading.
Two populations may differ dramatically on variables that do not modify the effect. Or they may appear similar overall while differing on one variable that changes the treatment effect strongly.
The more useful question is:
Which source–target differences could change the effect we are trying to transport?
Those variables are often called effect modifiers or moderators.
Age may matter for one treatment and not another. Baseline mathematics fluency may change the effect of a problem-solving intervention. Ambient humidity may modify the field performance of a material. Existing infrastructure may change the effect of a public-policy intervention.
Transportability therefore depends on causal structure, not demographic matching for its own sake.
5. Maya’s mistake: “The sample is large, so it must represent the target”
Maya sees a study with 50,000 participants.
She assumes the result must transport well.
But sample size mainly helps with precision inside the population actually studied. A very large non-representative study can estimate the wrong target quantity very precisely.
If the study contains almost no people like an important part of the target population, increasing the number of already-overrepresented participants does not solve the missing-support problem.
Repair: ask whether the source contains adequate information about the combinations of effect modifiers that occur in the target.
6. Jia Jun’s mistake: “Randomisation solves everything”
Randomisation is powerful because it protects treatment comparisons inside the trial.
But random assignment does not automatically make the trial participants representative of a wider target population.
A trial may randomly assign treatment among volunteers from a narrow age range, one healthcare system, a small number of schools or highly selected sites.
The causal comparison inside the trial can be strong while transport to another population remains uncertain.
Repair: separate internal validity from source-to-target selection.
See also How Scientific Randomisation Works and How Scientific Internal Validity Works.
7. Hana’s mistake: “If the groups differ, the study cannot be used”
Hana goes too far in the other direction.
She notices that the source and target populations differ and assumes transport is impossible.
But some differences may be irrelevant to the treatment effect. Others may be measurable and adjustable. Still others may be important but can be examined with sensitivity analyses.
Repair: distinguish three situations:
- differences unlikely to modify the effect;
- differences that may modify the effect and are adequately measured with source–target overlap;
- differences that may modify the effect but lack adequate measurement or overlap.
The third category is the dangerous one. It creates a genuine identification problem rather than a simple modelling inconvenience.
8. Ethan’s mistake: “A statistical model can fill any gap”
Ethan sees that the target includes groups missing from the study.
He proposes a sufficiently complicated machine-learning model.
But no algorithm can turn complete absence of relevant evidence into direct knowledge.
A model can extrapolate. Extrapolation may sometimes be scientifically reasonable. But it relies on stronger assumptions than interpolation within observed support.
Repair: treat missing support as missing evidence, not as an invitation to hide uncertainty inside a flexible model.
9. Overlap is the map of where transport has evidence beneath it
Suppose the source study contains participants aged 30–60, while the target population spans ages 20–90.
Transporting to ages 30–60 may be reasonably supported if other conditions hold. Transporting to people aged 85 is different: the study contains no directly comparable participants.
This idea is often discussed through positivity or overlap.
Informally:
For every scientifically relevant type of person or unit in the target, do we have a meaningful chance of observing comparable units in the source data?
If the answer is no, weighting can become extreme, estimates unstable, or the target effect not identifiable without additional assumptions.
10. Positivity is not a technical footnote
Consider a tutoring intervention tested only among students already scoring above 70 per cent.
Now suppose the target population includes many students scoring below 30 per cent.
If baseline attainment strongly modifies the intervention effect, the source study provides no direct evidence about the lower-attainment region.
A transport model can still output a number.
That does not mean the data earned that number.
The honest conclusion may be:
“The target effect is only partly identified by the available evidence; a new study is needed in the unsupported region.”
11. A worked example: why the source average can differ from the target average
Imagine a hypothetical educational intervention with different effects by prior knowledge.
Among students with high prior knowledge, the intervention improves the outcome by 2 points.
Among students with lower prior knowledge, it improves the outcome by 8 points.
Now suppose the source trial contains:
- 80% high-prior-knowledge students;
- 20% lower-prior-knowledge students.
The source average effect is therefore:
(0.80 × 2) + (0.20 × 8) = 3.2 points.
But suppose the target population contains:
- 40% high-prior-knowledge students;
- 60% lower-prior-knowledge students.
The target average effect, if those subgroup effects transport conditionally, would be:
(0.40 × 2) + (0.60 × 8) = 5.6 points.
Nothing went wrong in the original trial.
The population composition changed.
This is the central intuition behind weighting and standardisation methods: preserve the conditional effect information from the source, but average it according to the target population rather than the source population.
12. The crucial assumption: conditional transportability
A common identification strategy assumes that once we condition on a sufficient set of measured variables, the treatment effect can be transported between source and target.
In plain language:
After accounting for the effect modifiers that matter, study membership should not carry additional information about the treatment effect.
This is a strong scientific statement.
It can fail when an important effect modifier is unmeasured, measured differently across datasets, or absent from one population.
That is why variable selection should be driven by causal reasoning, subject-matter knowledge and evidence—not simply by feeding every column into a prediction model.
13. Directed acyclic graphs can help identify what must be adjusted
Modern transportability work often uses causal diagrams, including directed acyclic graphs, to represent treatment, outcome, source-versus-target selection and variables that modify or confound effects.
The graph is not proof.
It is a disciplined way to state assumptions.
A useful diagram asks:
- what causes entry into the source study?
- what affects the outcome?
- what modifies the treatment effect?
- which variables differ between source and target?
- which variables are measured comparably in both?
A graphical model can expose a hidden assumption before the statistics make that assumption invisible.
14. Weighting changes who the source study represents
One major family of transportability methods uses weights.
The idea is intuitive.
If a kind of participant is underrepresented in the source relative to the target, give comparable source participants more weight. If another type is overrepresented, give them less.
In transport settings this is often implemented through inverse odds of sampling or participation weights, though exact formulations depend on design and notation.
The goal is to create a weighted source sample whose distribution of relevant covariates resembles the target population.
Then the treatment effect estimated in that weighted pseudo-population is intended to answer the target question.
15. Weighting can reveal when the data are being stretched too far
Suppose one rare target group appears only a handful of times in the source.
To make the weighted source resemble the target, those few observations may receive enormous weights.
That creates several warnings:
- large variance;
- sensitivity to a few observations;
- small effective sample size;
- possible model dependence;
- poor practical overlap.
Extreme weights are therefore not merely a computational nuisance. They may be telling you that the source study does not contain enough relevant evidence for the target you chose.
16. Outcome modelling takes a different route
Another approach models outcomes in the source study as a function of treatment and relevant covariates.
The fitted model is then used to predict potential outcomes for the covariate distribution found in the target population.
This is often described as outcome regression followed by standardisation.
Conceptually:
learn the conditional response pattern in the source → apply that conditional pattern across the target covariate distribution → average over the target.
This can be efficient when the outcome model is well specified.
But it can also extrapolate badly when the target contains combinations of variables rarely or never observed in the source.
17. Calibration weighting can use target summary information
Sometimes researchers have individual-level source data but only summary statistics for the target population.
Calibration approaches can choose weights so selected moments—such as means or proportions of relevant covariates—match the target summaries.
Methods such as entropy balancing or method-of-moments calibration live in this family.
This is useful when full target microdata are unavailable.
But the same scientific law remains:
matching measured summaries cannot correct an unmeasured effect modifier.
18. Doubly robust methods combine two models
Modern analyses often combine a model for source-versus-target membership with a model for outcomes.
These approaches are called doubly robust because, under specific conditions, the target effect estimator can remain consistent if one of the two nuisance models is correctly specified even when the other is not.
Examples include augmented weighting estimators and targeted maximum likelihood approaches.
“Doubly robust” does not mean assumption-free.
It does not repair missing effect modifiers, absent overlap, poor measurement or a wrongly defined target population.
19. Balance diagnostics should be shown, not assumed
After weighting or matching, check whether the relevant source covariate distribution actually resembles the target.
Useful diagnostics can include:
- standardised differences before and after weighting;
- overlap plots;
- weight distributions;
- effective sample size;
- covariate ranges;
- regions of weak or absent support.
A transported estimate without transport diagnostics is hard to audit.
20. Measurement comparability can break transport before modelling begins
Suppose baseline severity is measured with one instrument in the source and another instrument in the target.
Or school disadvantage is defined differently in two administrative systems.
Or an outcome score changed its scale between years.
The columns may have similar labels while representing different constructs.
Transportability requires more than common variable names. It requires comparable meaning.
This is especially important when target data come from routine administrative or real-world sources while source data come from a tightly controlled study.
21. Missing data and measurement error are transport problems too
Suppose an important effect modifier is measured in the study but missing for a large fraction of the target population.
Or it is measured accurately in the trial but noisily in routine data.
The transport model is now trying to align populations through an imperfect lens.
Researchers should therefore examine:
- which variables are missing;
- whether missingness differs by population;
- whether the same construct was measured the same way;
- how measurement error may alter weighting or outcome models;
- whether sensitivity analyses change the conclusion.
22. Transporting a randomised trial is not the same as transporting an observational study
A randomised trial gives strong protection against treatment–outcome confounding within the trial when randomisation and follow-up are handled appropriately.
An observational source study does not receive that protection automatically.
When transporting observational evidence, the analysis must handle two distinct problems:
- internal causal identification: was the treatment effect estimated without important confounding inside the source?
- external transport identification: can that source effect be extended to the target population?
An analysis cannot repair the second problem while ignoring the first.
The modern literature therefore treats transport from observational studies as requiring explicit causal structures and assumptions for both confounding and source–target selection.
See How Causal Inference Works for the broader causal-inference owner.
23. Implementation can be an effect modifier
In education and public programmes, an intervention rarely travels as a sealed object.
It is delivered by people, inside institutions, with varying time, skill, resources and adherence.
A programme tested with expert coaching, protected timetable space and unusually high implementation support may not produce the same effect when scaled to ordinary schools.
For this reason, implementation conditions can belong inside the transport model rather than being treated as administrative background.
Education research increasingly emphasises target schools, moderators, implementation conditions and scale when judging whether effects should generalise.
24. The counterfactual can change between source and target
An intervention effect is always relative to something.
In one trial the comparison group may receive minimal support. In another system, “business as usual” may already include strong tutoring, technology, screening or teacher training.
If the counterfactual changes, the effect can change.
This is particularly important in education, medicine and policy.
Transportability is not merely about whether the treated population looks similar. It is also about whether the treatment contrast remains scientifically comparable.
25. Time can turn yesterday’s target into today’s source mismatch
Populations drift.
Clinical practice changes.
curricula change.
technology changes.
background risks change.
AI systems and user behaviour change.
A transported estimate therefore has a temporal boundary.
The target population must be defined not only by who and where, but sometimes by when.
26. Scale can change the causal system
A school programme tested in twelve motivated schools may create different effects when deployed nationally.
Teacher training capacity may become a bottleneck. Timetables may change. Peer interactions may change. Costs may alter implementation. Programme fidelity may fall. Competing interventions may interact.
At large scale, units may also interfere with one another, violating simple assumptions that each unit’s outcome depends only on its own treatment.
Transportability must therefore ask whether the mechanism itself remains intact at the target scale.
27. A second worked example: when transport should stop
Suppose a trial of a mathematics support programme includes students with baseline scores from 45 to 90.
The target population includes scores from 10 to 90.
Researchers discover that baseline score modifies the treatment effect.
Within 45–90, there is good source–target overlap.
Below 45, there is none.
The correct response is not to hide this problem by fitting a smooth curve through the entire 10–90 range.
A stronger analysis might:
- estimate a transported effect for the supported 45–90 region;
- report that the 10–44 region requires extrapolation;
- run sensitivity analyses under several plausible effect patterns;
- recommend direct evidence for the unsupported group.
Sometimes the scientifically mature result is a boundary, not a single universal number.
28. Transportability in education: from a study school to the schools that matter
Education provides an unusually clear example because schools differ at several levels at once.
Student characteristics matter.
Teacher experience matters.
school organisation matters.
curriculum matters.
implementation matters.
the comparison condition matters.
The IES guidance on generalisability recommends identifying potential moderators, defining the target population of students and schools, assessing sample–target differences and reporting how those differences affect inference.
This is a more disciplined question than “Does this teaching method work?”
The better question is:
“For which students and schools, under which implementation conditions, relative to which alternative, should this estimated effect be expected?”
29. Learning transfer is an analogy—not the same statistical problem
A student learns a method in one worked example and later applies it to an unfamiliar question.
That resembles transportability because a capability learned under one set of conditions is being tested under another.
But do not collapse the two concepts.
Statistical transportability estimates causal effects across populations under explicit identification assumptions.
Learning transfer asks whether a learner can use knowledge beyond the original learning context.
The analogy is useful because both ask which structures must survive a change in surface conditions. The mathematics and evidence standards are different.
For the learning owner, see Learning for Transfer.
30. Transportability in medicine: trial patients are not automatically routine patients
Clinical trials often use eligibility criteria, intensive follow-up and sites able to meet demanding protocols.
Real-world target populations may be older, have more comorbidities, use different background treatments or have different adherence patterns.
The formal transportability literature is especially developed in this domain because decision-makers routinely need to translate trial evidence into target-population effects.
Modern frameworks emphasise target definition, effect modifiers, covariate overlap, weighting or outcome modelling, doubly robust approaches, feasibility checks and sensitivity analysis.
This article is educational, not clinical advice. The methodological lesson is that a strong local causal estimate is the beginning of target-population reasoning, not the end.
31. Transportability in engineering: laboratory proof is not field proof
Suppose a component performs reliably in a controlled laboratory.
The deployment target experiences tropical humidity, vibration, dust, temperature cycling and variable maintenance.
Engineering rarely calls the formal statistical problem “transportability” in exactly the same way epidemiology does, but the scientific logic is recognisable:
Which environmental variables modify performance?
Does the test envelope cover the deployment envelope?
Where are we interpolating?
Where are we extrapolating?
What direct field evidence is still missing?
The analogy helps, but the formal estimand and design must remain domain-specific.
32. Transportability in AI: benchmark success is source-population evidence
An AI system performs well on a benchmark.
That benchmark contains certain languages, tasks, prompt styles, difficulty ranges and scoring procedures.
Deployment contains different users, tools, time horizons, error costs and distributions.
Again, “transportability” in machine learning is used in several related ways, so terminology varies.
But the same discipline helps:
- define the deployment target;
- identify variables likely to modify performance;
- measure source–target shift;
- check support;
- evaluate under target-like conditions;
- report failure regions rather than only an average score.
One benchmark result should never silently become “the model can do this everywhere”.
33. Sensitivity analysis asks how fragile the transported result is
Every transport analysis rests on assumptions.
Some can be checked partly in observed data. Others cannot.
A good sensitivity analysis asks:
- What if an unmeasured effect modifier differs between source and target?
- What if a measured moderator is noisy?
- What if extreme weights are trimmed?
- What if a different outcome model is used?
- What if the target definition changes?
- What if implementation differs more than assumed?
If the conclusion reverses under small plausible changes, the transported claim should be weaker.
34. Validation is stronger than elegance
A beautiful transport model should eventually face new evidence.
If outcome data later become available in the target population, compare transported predictions with what actually occurred.
Replication in additional target-like settings is especially valuable.
Model agreement is not the final receipt.
New observations are.
35. Transportability can improve future study design
The best time to think about transport is often before the trial begins.
If decision-makers know the intended target population, researchers can:
- sample sites more deliberately;
- measure likely effect modifiers consistently;
- avoid unnecessary exclusions;
- ensure adequate representation of important subgroups;
- collect comparable target-population data;
- pre-specify transport analyses;
- design replication or scale-up studies that test the weakest boundaries.
Good transportability is partly a design problem, not just an after-the-fact statistical rescue.
36. What weighting cannot fix
Weighting is often visually persuasive because it makes source covariate distributions resemble the target.
But it cannot automatically fix:
- unmeasured effect modifiers;
- variables defined differently across datasets;
- outcome measurement drift;
- intervention versions that are not equivalent;
- target regions absent from the source;
- confounding in an observational source study;
- interference or system change at scale;
- incorrect causal assumptions.
A weighted table can look balanced while the causal question remains unresolved.
37. What a transported estimate should report
A responsible transportability analysis should make the following visible:
- Source population: who generated the original causal evidence?
- Target population: who or what is the new estimand about?
- Treatment and comparator: are they meaningfully the same across source and target?
- Outcome: is it defined and measured comparably?
- Effect modifiers: which variables are believed to change the effect?
- Overlap: where is the target well represented in the source and where is it not?
- Identification assumptions: what must be true for the target effect to be identified?
- Estimator: weighting, outcome modelling, calibration, doubly robust method or another approach?
- Diagnostics: balance, weight stability, effective sample size, model fit and support?
- Sensitivity: how much do reasonable alternative assumptions move the answer?
- Uncertainty: what is sampling uncertainty and what uncertainty comes from assumptions?
- Boundary: which target subgroups remain weakly supported?
38. Common transportability errors
- Using “representative” without naming the target population.
- Assuming random assignment creates a representative sample.
- Matching on every measured variable instead of identifying effect modifiers.
- Ignoring target regions absent from the source.
- Treating extreme weights as merely a software problem.
- Using the same variable name as proof of measurement comparability.
- Transporting an observational association before establishing internal causal validity.
- Ignoring changes in the comparator or implementation.
- Reporting one transported number without overlap or sensitivity diagnostics.
- Calling model extrapolation “evidence” without marking the unsupported region.
- Confusing general scientific external validity with formal statistical transportability.
39. A practical learner method: TRACE the transport claim
For students and readers, a compact way to interrogate a transport claim is:
T — Target: Exactly who or what is the new conclusion about?
R — Relevant modifiers: Which source–target differences could alter the effect?
A — Available overlap: Does the source contain comparable cases across the target range?
C — Causal assumptions: What must be true after conditioning on measured variables?
E — Estimate and evidence check: How was the target effect estimated, diagnosed, stress-tested and validated?
TRACE is not a statistical estimator. It is a reading discipline.
40. A parent-friendly example
A parent hears: “This study technique improved test scores by six points.”
The transportability questions are:
- Who was studied?
- What ages and prior-attainment levels?
- What subject?
- How was the technique implemented?
- What did the comparison group do instead?
- Was the outcome measured immediately or later?
- Does the child resemble the studied population on variables likely to change the effect?
- Are we applying the same intervention or merely something with the same label?
The right conclusion may be “this evidence is relevant and worth testing”, not “this guarantees the same six-point gain”.
41. Why this matters for examination and tuition claims
Educational marketing often transports claims silently.
A method worked with one group.
A testimonial becomes a general promise.
A school result becomes a claim about every learner.
A short-term practice improvement becomes a claim about durable capability.
Transportability reasoning blocks that shortcut.
Ask what the original evidence actually measured, who produced it, what support conditions existed, and whether the learner now being discussed lies inside the evidence envelope.
This is one reason eduKate’s wider diagnostic approach distinguishes observed performance from claims about the learner.
42. How transportability connects to neighbouring scientific methods
Transportability does not replace the other layers of scientific reasoning.
- Internal validity asks whether the causal claim is credible inside the source study.
- Confounding asks whether a third variable distorts the treatment–outcome comparison.
- Effect size asks how large the estimated effect is.
- Confidence intervals express sampling uncertainty around estimates under a specified analysis.
- Meta-analysis combines evidence across studies and can expose heterogeneity across settings.
- External validity retains the broader question of whether findings apply beyond the original conditions.
Transportability occupies one precise position in that system:
given a source causal effect and a defined target population, what extra information and assumptions allow us to estimate the effect for that target?
43. Frequently asked questions
Is transportability the same as external validity?
No. External validity is broader. Transportability usually refers to formal extension of causal inferences from a source study to a specified target population, often using statistical adjustment for source–target differences.
Is transportability the same as generalisability?
Not always. A common distinction treats generalisability as extending from a sample to a broader population of which it is a subset, while transportability concerns a partly distinct target population. Definitions vary, so good work states the source and target explicitly.
Does a large randomised trial automatically transport well?
No. Randomisation protects internal treatment comparisons; large sample size improves precision. Neither guarantees that the participants represent the target population on variables that modify the effect.
What is an effect modifier?
A variable across which the size or direction of the treatment effect changes. Differences in effect-modifier distributions between source and target populations can make their average effects differ.
What is positivity or overlap?
It is the requirement that relevant types of target units have comparable representation in the source data. Poor overlap forces unstable weighting or unsupported extrapolation.
What is inverse-odds weighting?
It is a transport weighting approach that gives source participants weights related to how common their covariate profiles are in the target relative to the source, so the weighted source better resembles the target on selected variables.
What is standardisation?
It models conditional outcomes or treatment effects in the source and averages the resulting predictions over the target population’s covariate distribution.
What does “doubly robust” mean?
It describes estimators combining weighting and outcome models that can remain consistent if one of those nuisance models is correctly specified under the required identification assumptions. It does not make transport assumption-free.
Can machine learning solve poor transportability?
Machine learning can improve flexible estimation, but it cannot create evidence where key target groups or effect modifiers are absent. Better prediction machinery does not remove the need for causal assumptions and overlap.
Can transportability be tested?
Some assumptions and diagnostics can be examined in observed data, and transported predictions can be compared with later target-population evidence when available. But important identification assumptions may remain untestable from the existing data alone.
44. How do we know? Evidence base
The modern transportability literature is mature enough that the broad workflow is not speculative.
A 2026 conceptual framework endorsed by the International Society for Pharmacoepidemiology describes transportability analyses in terms of target definition, causal diagrams, effect modifiers, feasibility, overlap, weighting, outcome regression, calibration and doubly robust methods.
A 2023 review of generalisability and transportability methods provides a step-by-step workflow covering whether transport is appropriate, data availability, identifiability, population similarity, missingness, sensitivity analysis and interpretation.
Methods guidance in the NCBI Bookshelf distinguishes generalisability from transportability and emphasises explicit target populations and study-participation mechanisms.
For education, the Institute of Education Sciences provides a dedicated guide on designing impact studies for stronger generalisability, including defining target students and schools, identifying moderators, recruiting sites, comparing the sample with the target and adjusting for differences.
Recent methodological work has also extended transport ideas beyond randomised trials to observational studies, where confounding and source-to-target selection must both be handled.
45. Sources and further reading
- Transportability Analyses in Comparative Effectiveness Research: A Conceptual Framework and Methodological Principles — recent framework covering target populations, effect modifiers, overlap, weighting, outcome regression, calibration and doubly robust methods.
- An Overview of Current Methods for Real-world Applications to Generalize or Transport Clinical Trial Findings to Target Populations of Interest — practical workflow for translating trial findings to target populations.
- Methods to Apply Results from Randomized Trials to Patients in Clinical Practice — methodological guidance on extending trial inferences to target populations.
- Generalizability of Randomized Trial Results to Target Populations: Design and Analysis Possibilities — foundational design and analysis overview.
- Methods for Extending Inferences from Observational Studies — causal structures, identification assumptions and estimators for target-population inference from observational sources.
- Enhancing the Generalizability of Impact Studies in Education — Institute of Education Sciences guidance on target populations, moderators, sampling, adjustment and reporting in education research.
Conclusion: evidence does not travel by itself
Maya asks whether the study was large.
Jia Jun asks whether it was randomised.
Hana asks whether the source resembles the target.
Ethan asks whether a model can bridge the difference.
All four questions matter.
But none is sufficient alone.
Scientific transportability begins by naming the target. Then it asks which variables change the effect, whether the source contains comparable cases, whether measurements and interventions mean the same thing, which causal assumptions are required, how those assumptions are encoded in weighting or outcome models, and where the evidence still runs out.
The mature conclusion is rarely “the study applies everywhere”.
It is more precise:
“Given this source evidence, this target population, these effect modifiers, this overlap, these measurements and these assumptions, this is the effect we can justify transporting—and this is where the claim must stop.”
