Psychology dissertation overlap weighting offers a principled way to compare exposed and unexposed groups when an observational dataset contains adequate, but imperfect, covariate overlap. Instead of trying to make every participant represent the whole sample, it emphasises people whose observed characteristics made either exposure plausible.
This guide explains the estimand, design decisions, calculations, diagnostics and reporting choices needed for a defensible dissertation. It is relevant to binary exposures such as receiving a school wellbeing programme, accessing therapy, experiencing a policy change, or choosing one of two support pathways. Weighting cannot create randomisation, remove unmeasured confounding or repair weak measurement, so causal language must remain conditional on explicit assumptions.
What overlap weighting estimates
A propensity score is the estimated probability of receiving the exposure given measured pre-exposure covariates. For participant i, write this probability as e(Xi). With a binary exposure A, overlap weighting assigns exposed participants weight 1 − e(X) and unexposed participants weight e(X). A person receives more influence when their observed profile is common to both groups and less influence when their exposure was nearly predictable.
The resulting contrast is the average treatment effect in the overlap population, commonly abbreviated ATO. This target is not automatically the average effect for everyone, for those exposed, or for those unexposed. It concerns a covariate-defined population in which there is the greatest empirical uncertainty about exposure assignment. The scientific question must fit that population before the method is selected.
The foundational balancing-weights paper by Li, Morgan and Zaslavsky defines overlap weights within a broader class of propensity-score weights. The authors show that the weights are bounded and, within that class under stated conditions, minimise asymptotic variance for the overlap-population contrast. A later JAMA methods overview explains how the target resembles participants for whom either exposure option was plausible.
| Method | Primary target | Participant influence | Main distinction |
|---|---|---|---|
| Overlap weighting | Average effect in the overlap population | Highest near moderate propensity scores | Bounded weights smoothly reduce influence at the extremes |
| Inverse probability weighting | Often the full sample ATE | Can be very high at extreme propensity scores | Targets a different population and may be unstable with poor overlap |
| Propensity score matching | Depends on the matching design | Selected matched participants contribute | Pairs or groups records and may discard unmatched cases |
| Entropy balancing | Depends on chosen calibration target | Optimised to satisfy balance constraints | Directly calibrates selected covariate moments |
These methods are not interchangeable corrections for confounding. The existing propensity score matching guide addresses matched designs, while the entropy balancing guide focuses on moment constraints. Overlap weighting deserves a separate plan because it changes the target population, the weights and the interpretation.
When psychology dissertation overlap weighting fits
The approach can suit a secondary-data dissertation comparing two naturally occurring groups when both groups contain participants with similar measured baseline profiles. Imagine comparing improvement in anxiety scores among students who used guided online support and those who used self-guided materials. Students with severe baseline symptoms, prior service use or easy clinical access may overwhelmingly select guided support. Overlap weighting would emphasise students for whom either pathway was reasonably plausible rather than extrapolating heavily to profiles seen almost exclusively in one pathway.
Another example is a school dataset comparing pupils exposed or unexposed to a new classroom programme. The propensity model might include baseline wellbeing, prior attainment, age, school characteristics and pre-programme attendance. The analysis would estimate a programme contrast for pupils in the empirical region shared by participating and non-participating settings. It would not establish an effect for schools whose characteristics effectively determined adoption.
Conditions that strengthen the design
A credible application requires a well-defined binary exposure, a defensible time order, a meaningful ATO question, measured pre-exposure confounders and observations in both groups across relevant covariate patterns. Outcome-free design is preferable: determine eligibility, covariates, propensity specification and diagnostics without searching outcome results for a favourable method.
Overlap weighting is less suitable when exposure versions are vague, treatment can change repeatedly, interference between participants is central, or most profiles appear in only one group. Longitudinal exposures with time-varying confounding may require the methods discussed in the marginal structural models guide. A small sample with sparse categories can also make propensity estimation and uncertainty assessment fragile even though the resulting weights are bounded.
Define the causal question before modelling
Write the question as a target contrast: among people represented by the overlap of measured baseline characteristics, what would be the average difference in outcome under exposure versus non-exposure? Define eligibility, exposure versions, assignment time, follow-up, outcome and contrast. The target trial emulation guide can help align these components even when a full emulation is not feasible.
Next, draw a causal diagram or provide an equally explicit subject-matter argument for covariate selection. Include variables that causally precede exposure and outcome. Do not select variables solely because their univariable p-values are small. Avoid conditioning on descendants of exposure, colliders or variables recorded after assignment, because doing so can introduce bias.
The site’s confounding variables guide explains why a long adjustment list is not necessarily a valid one. Baseline outcome measurements are often highly prognostic and useful, but they must be temporally prior. Demographic variables should be represented respectfully and only when theory, design or equity assessment justifies their use.
State the identification assumptions
A causal interpretation generally requires consistency, exchangeability conditional on measured covariates, positivity in the target region and correct handling of sampling and missingness. Consistency means that the compared exposure versions correspond to the intervention or condition being discussed. Conditional exchangeability means that no unmeasured common causes remain after adjustment. Positivity means that each relevant covariate profile has some chance of each exposure.
Overlap weighting softens the influence of observations near propensity scores zero and one, but it does not prove positivity or rescue structural absence of a comparison group. Nor does exact balance on modelled covariate means prove exchangeability. Measurement error, omitted nonlinear relationships and unmeasured constructs can remain.
Build the propensity model carefully
For a binary exposure, logistic regression is a transparent starting point. Model exposure using the pre-specified covariates needed for confounding control. Continuous variables should not be divided into arbitrary categories merely for convenience. Flexible terms, such as restricted splines, can represent plausible nonlinearity, and selected interactions can represent treatment-assignment mechanisms that differ across covariate combinations.
Propensity modelling is a design task, not an exposure-prediction competition. A high classification accuracy can signal that groups are easy to separate, which is the opposite of useful overlap. Judge specifications primarily by whether they support the estimand and achieve acceptable weighted balance without a misleadingly tiny effective sample.
The current WeightIt documentation identifies ATO as the estimand requesting an overlap sample for supported binary-treatment methods. Software makes calculations reproducible, but the dissertation must still explain the formula, covariate rationale, model form and versioned code rather than treating a package default as a methodological justification.
Calculate and inspect the weights
After estimating propensity scores, give an exposed participant the probability of the opposite condition, 1 − e(X), and an unexposed participant the probability of exposure, e(X). Both weights lie between zero and one. Normalisation may be useful for some summaries, but it does not change a simple weighted mean within each group when applied consistently.
Plot propensity-score distributions by exposure before weighting. Then summarise weight minima, maxima, central quantiles and effective sample size by group. Effective sample size is commonly calculated as the squared sum of weights divided by the sum of squared weights. It describes information loss from unequal weighting, not the number of independent people or a guarantee of adequate power.
| Checkpoint | What to inspect | Warning sign | Possible response |
|---|---|---|---|
| Time order | All covariates measured before exposure | Post-exposure adjustment | Rebuild the covariate set using a causal rationale |
| Overlap | Propensity distributions and covariate support | Profiles unique to one group | Narrow the target or reconsider the comparison |
| Balance | Weighted mean and distributional differences | Material residual imbalance | Revise justified functional forms without outcome fishing |
| Weights | Range, quantiles and group totals | Few observations dominate | Check coding, model and target population |
| Precision | Effective sample size and interval width | Very weak information | Report uncertainty and reduce claims |
Diagnose balance beyond a single threshold
Report balance before and after weighting for every pre-specified covariate. Standardised mean differences are useful because continuous variables appear on a common scale. The cobalt balance documentation describes adjusted and unadjusted differences and effective sample size outputs. A Love plot makes the pattern visible, but a table should provide exact values.
A threshold such as an absolute standardised mean difference below 0.10 is a heuristic, not proof of validity. Inspect the largest imbalance, important prognostic variables, variance ratios, categorical proportions and, where important, distributional plots or empirical cumulative distributions. If the propensity model contains squared terms or interactions, assess balance for those represented features as well as main variables.
With a logistic propensity model fitted by maximum likelihood, overlap weights have an exact mean-balance property for included covariates under the standard setup. Treat this as a mathematical property of the fitted design, not evidence that omitted variables are balanced. Rounding, missing-data processing, survey weights, penalisation or alternative learners can also change expected behaviour.
Do not use significance tests as balance diagnostics
Balance is a property of the analysed sample, whereas null-hypothesis tests conflate imbalance with sample size. Large samples can produce small p-values for negligible differences, and small samples can hide important ones. Descriptive diagnostics, effect-size scales and plots are more informative. Pre-specify how unacceptable balance will trigger model revision or a change in target.
Estimate the outcome contrast and uncertainty
For a continuous outcome, the basic ATO estimate is the difference between overlap-weighted outcome means. Binary outcomes may be expressed as weighted risks, risk differences or risk ratios when supported by the analysis. A weighted regression can improve precision or support additional adjustment, but its link function and estimand must match the reported measure.
Standard errors must reflect weighting and, where the method supports it, estimation of propensity scores. Naive model-based standard errors that treat weights as fixed frequency counts can be inappropriate. Use a method justified by the estimator, such as a robust sandwich estimator, M-estimation or a carefully designed bootstrap. The software documentation and primary method should support the chosen variance procedure.
Do not report only a coefficient and p-value. Provide the effect estimate, confidence interval, outcome scale, target population and an interpretation in substantive units. For an anxiety measure, explain the estimated mean difference and its uncertainty among students in the overlap population. Avoid claiming a population-wide benefit when people with near-deterministic exposure received little weight.
| Dissertation section | Essential content | Transparent wording |
|---|---|---|
| Methods | Eligibility, exposure, outcome, ATO, covariates, model, weight formula | “We estimated the average contrast in the overlap population.” |
| Results | Group flow, overlap, balance, weight summaries, effective sample size | “Weighting reduced measured baseline differences.” |
| Effect | Estimate, interval, scale and variance method | “The weighted mean difference was…, with a 95% confidence interval…” |
| Discussion | Target population, assumptions, measurement and missingness limits | “A causal interpretation depends on no important unmeasured confounding.” |
Sensitivity analyses that add information
A useful sensitivity plan examines decisions that are scientifically plausible rather than trying many models until one is significant. Compare a primary specification with justified nonlinear terms, alternative coding of key baseline variables and an outcome model that adds important prognostic covariates. Report whether conclusions change in magnitude, direction or precision.
Compare the overlap-population result with another estimand only when each answers a meaningful question. An ATE estimate from inverse probability weights is not a robustness check if the full-population target lacks support. If both are reported, clearly distinguish their populations. Large differences may reflect effect heterogeneity, different extrapolation or instability rather than an error.
Address missing data explicitly. Complete-case weighting estimates an effect for a selected complete-case population unless strong conditions hold. Multiple imputation should preserve exposure, outcome, covariates, interactions and auxiliary predictors, and weighting should be conducted within each imputed dataset with estimates combined using a justified procedure. Do not impute propensity scores as though they were raw variables.
Assess unmeasured confounding through a method appropriate to the effect scale and design, accompanied by subject-matter judgement. Quantitative sensitivity analysis can show how strong an omitted factor would need to be, but it cannot certify that no such factor exists. Negative controls may reveal certain biases when credible controls are available, as discussed in the site’s negative control analysis guide.
Common mistakes and how to repair them
Choosing overlap weights only because IPW is unstable
Bounded weights are attractive, but the estimand must drive the choice. State why the overlap population is substantively relevant before citing improved stability. If the research question truly concerns everyone eligible, limited positivity is a design limitation, not merely a computational inconvenience.
Calling weighted groups randomised
Overlap weighting can mimic selected balance features of a trial for measured baseline covariates, but exposure remains observational. Randomisation also addresses unmeasured factors probabilistically and defines assignment procedures. Write “weighted observational comparison” and state the identification assumptions.
Reporting exact balance as proof of no confounding
Exact mean balance covers included model terms under the relevant estimation setup. It does not cover omitted constructs, measurement error or every distributional feature. Pair the mathematical result with wider diagnostics and a limitation statement.
Hiding the changed target population
Readers need to know who contributes most. Describe weighted baseline characteristics, group weight totals and effective sample sizes. Explain that people whose profiles almost determine exposure contribute less to the ATO result.
Using post-exposure covariates
Variables measured after exposure can be mediators or consequences of both exposure and outcome-related processes. Adjusting for them may block part of an effect or create collider bias. Fix the time origin and defend each covariate before estimation.
Reproducible workflow for a dissertation
- Define the binary exposure, outcome time, eligibility criteria and ATO question.
- Use theory and evidence to identify pre-exposure confounders and draw a causal diagram.
- Document data provenance, exclusions, missingness and measurement quality.
- Inspect raw covariate distributions and empirical overlap without using outcomes to tune the design.
- Fit the pre-specified propensity model with justified nonlinearities and interactions.
- Calculate overlap weights and verify exposure coding and formulas.
- Inspect propensity distributions, weight summaries, effective sample sizes and weighted covariate balance.
- Revise only for defensible design reasons, recording every change.
- Estimate the outcome contrast with an uncertainty method appropriate to estimated weights.
- Run pre-specified sensitivity analyses and preserve versioned code, outputs and decisions.
A reproducible appendix should include software and package versions, a variable dictionary, propensity formula, weight definition, diagnostic code and the final analysis specification. Remove direct identifiers and follow the dataset’s governance terms. Code transparency does not override participant confidentiality or data-use agreements.
Frequently asked questions
What is the focus of psychology dissertation overlap weighting?
It estimates an average exposure contrast for the covariate-defined population with the greatest observed overlap between groups. It is not automatically an estimate for the full sample.
Are overlap weights the same as propensity score matching?
No. Overlap weighting retains observations with smoothly varying influence, while matching constructs matched sets and may exclude records. Their target populations can also differ.
Can overlap weighting fix extreme propensity scores?
It limits their influence because weights are bounded, but extreme scores still reveal limited empirical support. The analysis must describe that limitation and the narrower ATO target.
Does exact mean balance remove all confounding?
No. The property applies to included covariate terms under the relevant fitted model. Unmeasured variables, measurement error and omitted distributional features can remain imbalanced.
Should outcomes be used to choose the propensity model?
Prefer outcome-free design. Choose covariates and functional forms from temporal ordering, causal knowledge and balance diagnostics, not from whichever version gives a preferred outcome result.
What should I report about overlap weights?
Report the formula, propensity model, score distributions, weight summaries, effective sample size, pre- and post-weighting balance, outcome estimator, variance method and target-population interpretation.
Can I make causal claims after overlap weighting?
Only cautiously and only if the design, data and assumptions support them. State that exchangeability depends on measured covariates and discuss positivity, consistency, missingness and measurement limitations.
Conclusion
Psychology dissertation overlap weighting is most defensible when the overlap population answers a real scientific question, covariates are temporally and causally justified, and diagnostics are treated as evidence rather than decoration. Its bounded weights and covariate-balance properties can support stable observational comparisons, but they do not replace thoughtful design or eliminate unmeasured confounding.
If you need dissertation support, seek ethical guidance that helps you define the estimand, audit your code and understand your results while you retain authorship and comply with university rules. A responsible review should strengthen your reasoning, not manufacture data, conceal analytic decisions or guarantee a particular finding.
