Psychology researcher comparing matched participant groups and covariate balance

Psychology dissertation propensity score matching can make observational groups more comparable on measured baseline characteristics, but it cannot turn non-random data into a randomised experiment. A defensible analysis begins with a causal question, defines the population and treatment effect, chooses covariates using subject knowledge, checks overlap, and demonstrates balance before outcomes are compared.

This guide explains that workflow for psychology dissertations. It treats matching as a design stage, not a button to press after finding an interesting result. The examples use a hypothetical evaluation of whether voluntary access to an eight-week online cognitive behavioural programme reduces anxiety at follow-up. Students who enrol may differ from non-enrollees in baseline anxiety, prior treatment, age, motivation, digital access, and other characteristics. Propensity score methods try to construct a comparison that is more credible with respect to measured pre-treatment variables.

What propensity score matching estimates

A propensity score is a participant’s estimated probability of receiving the exposure or intervention, conditional on selected baseline covariates. Two participants can have similar scores even when their individual covariate values are not identical. Matching uses the score, sometimes alongside exact or distance-based constraints, to pair or group exposed and unexposed participants.

The method targets imbalance in observed baseline characteristics. It does not control an unmeasured cause of both programme enrolment and anxiety. It also cannot repair a badly measured confounder, reverse causation, a poorly defined intervention, or selection caused by restricting analysis to a collider. Read the site’s guides to confounding variables and selection bias before presenting matching as a causal solution.

Matching is part of study design

Outcome values should not drive the matching specification. Researchers can examine treatment assignment and baseline covariates while designing the matched sample, then analyse outcomes after balance and overlap decisions are fixed. This separation reduces opportunities to choose a design because it produces a preferred effect. Stuart’s review describes matching as a preprocessing stage that precedes outcome analysis.

Propensity scores are not outcome predictions

A propensity model predicts exposure, not the psychological outcome. High classification accuracy is not the objective. A model can predict treatment well yet leave important covariates imbalanced. The practical criterion is whether the resulting matched sample achieves acceptable covariate balance for the intended estimand without discarding so many participants that the target population changes beyond recognition.

Decide whether psychology dissertation propensity score matching fits

The method is most plausible when the dissertation compares a clearly defined binary exposure or intervention in observational data, includes enough exposed and unexposed participants, measures credible common causes before exposure, and shows meaningful overlap between groups. It may support a quasi-experimental design, but the strength of inference still depends on assumptions and data quality.

Research situation Suitability Reason
Programme enrolment versus non-enrolment with rich baseline data Potentially suitable Measured pre-treatment differences can be assessed and balanced
Exposure has three or more versions Needs an extension Standard binary matching does not represent multiple treatments
Important confounders were not measured Weak Matching cannot balance information that is absent
Almost every high-risk participant received treatment Often weak Limited overlap leaves few credible comparisons
Treatment changes over time Usually needs another method Time-varying exposure and confounding require longitudinal methods

A small dissertation dataset can make matching unstable. One-to-one matching may discard observations, while flexible propensity models require information across covariate patterns. Do not use a universal minimum sample rule. Instead, inspect group sizes, treatment prevalence, overlap, the number and distribution of covariates, missingness, expected attrition, and the precision needed for the outcome estimate.

Define the causal question and estimand first

Write the comparison in plain language before fitting a model. For example: among students who enrolled in the online programme, what would their mean follow-up anxiety have been if, contrary to fact, they had not enrolled? This question targets the average treatment effect among the treated, commonly abbreviated ATT. A different question about the whole eligible population targets the average treatment effect, or ATE.

The estimand guides which participants should represent whom, how weights are constructed, and how exclusions affect interpretation. Nearest-neighbour matching of treated participants to controls commonly targets the ATT. Other matching or weighting schemes may target the ATE or an overlap population. State the estimand rather than writing only that the groups were “matched.”

Estimand Plain-language question Typical interpretation
ATT What was the effect for participants who received the intervention? Applies to treated participants represented in the matched sample
ATE What was the average effect in the eligible study population? Requires support for both treatment conditions across that population
ATO What was the effect among participants with substantial treatment equipoise? Applies to the region where groups overlap most

State the assumptions without hiding them

Consistency and a well-defined intervention

Consistency assumes that a participant’s observed outcome under the treatment received corresponds to the relevant potential outcome. “Received support” may be too vague if support ranges from one automated message to weekly therapy. Define treatment timing, content, intensity, and versions. If versions are substantively different, a single binary indicator may conceal important variation.

Conditional exchangeability

This assumption says that, after conditioning on appropriate baseline covariates, treatment assignment is independent of the potential outcomes. It is often paraphrased as no unmeasured confounding. The data cannot prove it. Defend it using theory, prior evidence, temporal ordering, measurement quality, and a causal diagram. Phrase conclusions conditionally and discuss plausible unmeasured common causes.

Positivity and overlap

For relevant covariate patterns, participants must have a non-zero probability of receiving either condition. If every student with severe symptoms enters the programme and no comparable untreated students exist, the missing counterfactual cannot be recovered by a clever model. Empirical overlap is assessed from score distributions and covariate patterns, but substantive knowledge is needed to judge whether the comparison makes sense.

No interference

Standard analyses also assume one participant’s treatment does not affect another participant’s outcome. This can fail in group therapy, classrooms, couples, peer networks, or shared residences. If programme participants share coping techniques with controls, interference changes the causal question and may dilute the contrast.

Choose propensity-score covariates before treatment

Include credible pre-treatment causes of treatment and outcome, especially strong outcome predictors. Use psychological theory, clinical knowledge, prior studies, and a directed acyclic graph. For the anxiety-programme example, candidates could include baseline anxiety, previous therapy, medication, age, study workload, prior digital-service use, and socioeconomic access. Record why each variable is included.

Do not select covariates merely because their univariate p-values are below .05. Statistical significance depends on sample size and does not identify confounders. Avoid variables measured after enrolment, such as engagement, satisfaction, or therapeutic alliance, because they may be consequences of treatment. Adjusting for a mediator can remove part of the effect; adjusting for a collider can introduce bias.

Baseline versions of outcomes are often valuable because they predict follow-up outcomes and clarify change. Their timing must truly precede treatment. Document every operational definition using the site’s guide to operational definitions.

Estimate the propensity score transparently

Binary logistic regression is common, but it is not the only option. Specify treatment as the dependent variable and the selected baseline covariates as predictors. Continuous variables need not be categorised. Consider non-linear terms and defensible interactions where treatment assignment plausibly varies non-linearly. The site’s logistic regression guide explains coding, sparse data, and model specification.

Do not interpret propensity-model coefficients as the main scientific findings. Their role is design. Save predicted scores, examine their distributions by group, and assess whether extreme values reflect genuine lack of overlap, sparse combinations, coding errors, or an overly rigid model.

Select a matching strategy

Nearest-neighbour matching

Each treated participant is paired with the nearest control according to the propensity score or its logit. Matching may be one-to-one or use multiple controls. Matching without replacement gives each control at most once; matching with replacement can improve closeness when suitable controls are scarce but requires analysis that respects repeated use and weights.

Calipers and exact constraints

A caliper prevents poor matches by setting a maximum acceptable distance. A frequently studied choice is 0.2 standard deviations of the logit of the propensity score, but it is not a law. Report the scale, threshold, and rationale. Exact or coarsened-exact constraints can ensure matches share a crucial feature such as study site, diagnostic category, or baseline risk band.

Full matching and weighting alternatives

Full matching forms sets containing at least one treated and one control participant and can retain more data. Inverse probability weighting creates a weighted pseudopopulation rather than pairs. Overlap weighting emphasises participants with more comparable treatment probabilities. These approaches address related goals but produce different estimands, diagnostics, and variance requirements. A dissertation should choose the method for the question, not because one command is familiar.

Decision What to report What to inspect
Distance measure Propensity score, logit, or combined distance Whether close scores also improve covariate balance
Ratio Number of controls per treated participant Balance, retained sample, and precision
Replacement Whether controls can be reused Weights and repeated-control influence
Caliper Threshold, scale, and justification Unmatched cases and residual imbalance
Common support How non-overlap was handled Who is excluded and how the estimand changes

Diagnose overlap before accepting matches

Plot propensity score distributions by treatment group before and after matching. Histograms, density plots, empirical cumulative distributions, or mirrored plots can reveal unsupported regions that one summary statistic hides. Also inspect important covariates directly. Two participants with similar scores can still differ meaningfully on a key variable.

Trimming or restricting to common support can improve comparability, but it changes the population represented by the result. Record how many participants are removed from each group, their characteristics, and the resulting estimand. If many treated participants have no plausible controls, the honest conclusion may be that the available data cannot estimate the intended effect.

Check covariate balance, not propensity-model fit

Balance diagnostics ask whether the matched groups have similar distributions of baseline covariates. Report standardised mean differences before and after matching for every planned covariate. An absolute value below 0.10 is often used as a practical flag, but it is not proof of adequate balance. Examine variance ratios, distributional plots, higher moments, and interactions where relevant.

A Love plot can display standardised differences across covariates. Do not rely on post-match significance tests: their p-values mix imbalance with sample size and can look reassuring after matching reduces the sample. Austin’s balance-diagnostics paper recommends examining balance in baseline covariates rather than testing whether the treatment-assignment model predicts well.

If balance is poor, revise the design without consulting the outcome: change functional forms, add defensible interactions, strengthen exact constraints, alter the caliper, use a different matching method, or reconsider the target population. Pre-specify a decision process where possible and retain an audit trail of attempted specifications.

Analyse outcomes after matching

Matching does not itself estimate the effect. Compare outcomes using a model or estimator that reflects matched sets, weights, repeated controls, and outcome type. For a continuous anxiety score, this might be a weighted mean difference or regression with appropriate standard errors. Binary and time-to-event outcomes need methods aligned with their scales.

Report the effect estimate, confidence interval, units, and target population. A p-value alone is inadequate. Avoid automatically treating matched observations as independent. The correct variance estimator depends on the matching algorithm, replacement, weighting, and estimand. Seek statistical supervision when clustering or complex weights make the variance calculation uncertain.

Residual outcome adjustment for prognostic baseline variables can sometimes improve precision or address remaining imbalance, but explain the rationale and preserve the matched design. Do not return to an unrestricted outcome-model search after matching.

Handle missing data as part of the design

Complete-case matching can change the population and create selection bias when missingness is related to exposure or outcome risk. Describe missingness for treatment, covariates, and outcomes separately. Decide whether multiple imputation is defensible, and ensure its sequence is compatible with propensity estimation and effect calculation. Combining estimates across multiply imputed matched datasets requires more than selecting the single dataset with the neatest balance.

Use the guide to psychology dissertation missing data to plan assumptions, diagnostics, and sensitivity analyses. Never fill missing baseline values with a mean solely to make matching software run.

Run sensitivity and robustness analyses

Good measured balance does not establish balance on unmeasured variables. Identify plausible omitted confounders and discuss the direction of likely bias. Where suitable, use a formal sensitivity analysis to show how strong an unmeasured association would need to be to explain the estimate. Interpret such analysis as a structured “what if,” not evidence that no hidden bias exists.

Compare defensible specifications: alternative calipers, matching ratios, replacement choices, or an outcome regression using the same confounder set. Examine whether conclusions depend on a few heavily weighted or repeatedly used controls. Distinguish a planned primary analysis from robustness checks so readers can see which result answers the original question.

A psychology dissertation propensity score matching workflow

  1. Define the exposure, comparison, outcome, eligible population, time zero, and follow-up.
  2. Choose ATT, ATE, or another estimand and state it in plain language.
  3. Draw a causal diagram and select measured pre-treatment covariates.
  4. Audit timing, coding, missingness, group sizes, and measurement quality.
  5. Estimate propensity scores without using outcomes to tune the design.
  6. Inspect overlap and decide whether the target population is supported.
  7. Apply a documented matching specification.
  8. Assess balance for every covariate and revise the design if necessary.
  9. Lock the design, then estimate the outcome contrast with suitable uncertainty.
  10. Run sensitivity analyses and report exclusions, assumptions, and limitations.

Preregistering the estimand, covariate logic, balance criteria, and outcome analysis makes analytic flexibility visible. The preregistration guide shows how to distinguish planned decisions from later changes.

Software and reproducibility

R’s MatchIt documentation demonstrates matching and balance assessment across several methods. Whatever software you use, preserve code, package versions, random seeds where relevant, data exclusions, matching specifications, diagnostic outputs, and a participant-flow record. A menu click is not a reproducible method.

Do not copy code without understanding its default estimand, distance, replacement rule, or treatment coding. Confirm which category is treated, how missing values are handled, and whether exported weights must be used in the outcome model. Keep de-identified analysis files under the approved data management plan.

How to report the matched study

In Methods, report the causal question, estimand, eligibility criteria, treatment definition, time zero, covariates and their timing, propensity model, distance, ratio, replacement, caliper, exact constraints, common-support rule, balance criteria, missing-data approach, outcome model, and variance estimator. Explain all deviations from the plan.

In Results, provide a participant-flow account, unmatched and matched group sizes, exclusions, score overlap, covariate balance before and after matching, effective sample size where weighting applies, effect estimates with confidence intervals, and sensitivity analyses. A baseline table should contain descriptive summaries and balance measures rather than p-values alone.

In Discussion, limit the conclusion to the supported population and acknowledge unmeasured confounding, measurement error, interference, selection, missingness, and limited overlap. Use the official STROBE checklists to strengthen reporting of observational studies; STROBE is a reporting aid, not a guarantee that the design is unbiased.

Common mistakes to avoid

  • Matching every available variable: include covariates for a causal reason, not because they exist.
  • Including post-treatment measures: this can block mediation or create bias.
  • Optimising classification accuracy: treatment prediction is not the design goal.
  • Reporting only the matched sample size: show overlap, exclusions, and balance.
  • Using p-values as balance tests: report standardised and graphical diagnostics.
  • Ignoring the matching structure: outcome models and standard errors must respect it.
  • Claiming all confounding is removed: only measured, adequately modelled covariates can be balanced.
  • Generalising to discarded participants: trimming and unmatched cases can narrow the target population.

Frequently asked questions

Does propensity score matching prove causation?

No. It can improve comparability on measured baseline covariates under explicit assumptions. Unmeasured confounding, poor measurement, limited overlap, interference, and selection may still bias the estimate.

What balance threshold should I use?

An absolute standardised mean difference below 0.10 is a common diagnostic flag, not a universal pass mark. Consider the importance of each covariate, variance ratios, full distributions, interactions, and the consequences of residual imbalance.

Should the outcome be included in the propensity model?

No. The propensity model uses treatment and pre-treatment covariates. Outcome values should not guide matching choices. Strong baseline predictors of the outcome may be included as covariates when they precede treatment.

Can I match using SPSS?

Extensions and workflows exist, but availability varies by version and installation. The essential requirements remain the same: document defaults, examine overlap, report balance, use appropriate matched-sample analyses, and preserve a reproducible record.

Is one-to-one matching always best?

No. One-to-one matching is easy to explain, but full matching, variable ratios, replacement, or weighting may provide better balance or precision. Compare defensible strategies using the estimand and diagnostics, not the most favourable outcome.

What if many treated participants have no matches?

Do not force poor matches. Reconsider the target population, treatment definition, data source, or method. Report who was excluded and recognise that the effect may apply only to the supported subset.

Can matching fix a small sample?

No. Matching may reduce the analysed sample and cannot manufacture overlap or information. A small study needs a restrained design, realistic precision goals, and transparent acknowledgement of uncertainty.

Conclusion

Psychology dissertation propensity score matching is defensible when it is anchored to a clear estimand, credible pre-treatment covariates, adequate overlap, transparent matching rules, and convincing balance diagnostics. The matched design must then be carried into outcome analysis, uncertainty estimation, sensitivity checks, and cautious interpretation. It reduces imbalance in measured characteristics; it does not erase the observational nature of the evidence.

If you want ethical dissertation support, use feedback to clarify your causal question, review your assumptions, check code and diagnostics, or improve reporting. Keep ownership of the research decisions, analysis, and final writing, and follow your university’s rules on authorship, data protection, and permitted assistance.

Authoritative references

Leave a Reply

Your email address will not be published. Required fields are marked *