Psychology dissertation propensity score subclassification divides an observational sample into strata of similar exposure probabilities, then compares outcomes within those strata.
The method can make exposed and comparison groups more alike on measured baseline covariates without forcing every participant into a pair. It is useful when a dissertation examines a clearly defined binary exposure, such as voluntary therapy uptake, participation in a school wellbeing programme, or use of a digital mental health service.
Subclassification is not automatic proof of causality. Its credibility depends on the estimand, timeline, propensity-score model, overlap, number and composition of strata, covariate balance, weighting, outcome analysis, and assumptions about unmeasured confounding. This guide explains each decision and shows how to report it transparently.
Table of Contents
What propensity score subclassification does
The propensity score is the estimated probability of receiving the exposure given selected pre-exposure covariates. Instead of matching individual participants, subclassification orders participants by that score and places them into ranges, often called strata, subclasses, blocks, or bins.
Each usable stratum must contain exposed and comparison participants. The effect is estimated within strata and then pooled, or the subclasses are converted into weights for a single marginal analysis. Comparisons are therefore made among participants with broadly similar predicted exposure probabilities.
The official MatchIt subclassification documentation explains that subclass boundaries can be based on score quantiles in the exposed group, comparison group, or full sample, depending on the target estimand. It also documents how subclass proportions determine analysis weights.
A psychology example
Suppose a dissertation compares university students who chose an eight-week online anxiety programme with students who used standard support. Programme uptake may be associated with baseline anxiety, previous counselling, medication, age, workload, financial strain, study mode, and help-seeking attitudes.
A model estimates each student’s probability of programme uptake from those baseline variables. Students are then divided into score strata. Within a middle stratum, for example, programme and comparison students may all have predicted uptake probabilities between 0.35 and 0.48. Their outcomes are more defensibly comparable than outcomes from students at opposite ends of the score distribution.
| Feature | Subclassification | Pair matching | Direct weighting |
|---|---|---|---|
| Design unit | Score-based strata | Pairs or fixed-ratio sets | Individual weights |
| Retention | Often high when overlap is adequate | Can discard unmatched participants | Usually retains participants within support |
| Main tuning choice | Number and boundaries of strata | Distance, caliper, ratio, replacement | Weight formula and trimming rules |
| Main diagnostic risk | Residual imbalance within strata | Poor matches or changed target population | Extreme or unstable weights |
| Outcome analysis | Pooled stratum effects or stratum weights | Matching weights with dependence handled | Weighted marginal model |
When psychology dissertation propensity score subclassification fits
The method fits an observational study with a binary exposure, a defensible set of baseline confounders, sufficient overlap, and enough participants from both groups across the score distribution. It is particularly practical for moderate or large samples where pair matching would remove many records or where a simple stratified design is easier to explain.

Relevant psychology applications include optional counselling, naturally chosen learning support, workplace resilience training, parenting interventions, peer-support participation, and service referral pathways. The exposure must be defined consistently, with a clear start date and meaningful alternative.
Subclassification is a poor choice when one exposure group is absent across wide score regions, key confounders were not measured, covariates were recorded after exposure, or strata become too sparse for stable comparisons. It also does not solve interference, measurement error, informative attrition, or outcome missingness by itself.
Separate design from outcome analysis
Build and assess the subclasses without examining exposure-effect estimates. This separation reduces pressure to alter the propensity model until a preferred result appears. A design log should record every specification, its rationale, overlap, stratum counts, and balance diagnostics.
Define the causal question and estimand first
The average treatment effect in exposed participants asks about the people who received the programme or exposure. The average treatment effect in comparison participants targets those who did not. The average treatment effect targets the full eligible population. These are different questions and can yield different effects when responses vary across participants.
The estimand influences which score distribution defines the quantile boundaries and how stratum results are pooled. In MatchIt, ATT boundaries use quantiles among exposed participants, ATC boundaries use the comparison group, and ATE boundaries use the full sample.
For the anxiety-programme study, an ATT might answer whether participants who chose the programme benefited compared with what they would have experienced under standard support. An ATE might inform a policy about offering the programme to every eligible student. State the target population in the research question and analysis plan.
Set a defensible time zero
Align eligibility, exposure assignment, baseline covariates, follow-up start, outcome measurement, and censoring. Baseline anxiety and previous therapy must be measured before programme uptake. Attendance, adherence, post-enrolment medication changes, later distress, and dropout occur after exposure and should not enter the propensity model as baseline confounders.
Matching or stratifying on post-exposure variables can block part of the effect or create collider bias. A causal diagram and dated variable inventory can make this mistake visible before analysis.
Select baseline covariates by causal reasoning
Include credible pre-exposure causes of both exposure and outcome, plus strong outcome predictors that may improve precision. Use theory, earlier research, clinical or educational knowledge, and temporal ordering. Do not select covariates only because they predict exposure, are statistically significant, or improve model accuracy.
For the student example, plausible covariates include baseline anxiety, depressive symptoms, previous counselling, medication, age, disability accommodations, study level, workload, sleep quality, financial strain, referral route, and help-seeking attitudes. Site or university may require separate models or exact restriction if access systems differ substantially.
Avoid including instruments, colliders, mediators, or redundant transformations without a clear rationale. Variables that predict exposure but not outcome can increase score separation and harm precision. Totals entered alongside all their component items can create instability without improving confounding control.
Handle missing baseline data deliberately
Complete-case analysis can change the target population and weaken overlap. When multiple imputation is appropriate, repeat the propensity-score design and outcome analysis within each imputed dataset and use a justified pooling procedure. Matching or subclassifying only one completed dataset can conceal uncertainty.
Report missingness by exposure group, the imputation model, number of imputations, whether exposure and outcome informed imputation, how strata or weights were handled, and how estimates were pooled.
Estimate the propensity score for balance
Logistic regression is common for a binary exposure, but flexible models can represent nonlinearities and interactions. Model choice should be guided by post-design covariate balance rather than prediction statistics alone. A high classification accuracy or area under the receiver operating curve does not demonstrate successful confounding control.
Prespecify plausible nonlinear terms and interactions. Baseline anxiety may have a curved relationship with programme uptake, while the association between prior counselling and uptake may differ by referral route. Examine whether the score model creates adequate overlap and balanced strata.
The foundational insight is that a propensity score summarizes measured covariate information into a balancing score under strong assumptions. Those assumptions include no unmeasured confounding for the defined contrast, positivity, consistency, and correct handling of time. The score cannot balance variables that were omitted or measured poorly.
Inspect overlap before making strata
Plot score distributions by exposure group and examine the tails. Participants with probabilities near zero or one may have no credible counterparts. Report how common support was defined, who was excluded, and how exclusion changes the target population.
Do not hide poor overlap by widening strata until every participant fits. A stratum containing very low-score comparison participants and very high-score exposed participants may have both groups but still provide an implausible comparison.
Choose the number and boundaries of subclasses
Five strata are often mentioned in older guidance, but they are not a universal rule. The current MatchIt matching-methods guide notes that its default of six is arbitrary and that the best number can be much larger and varies by dataset.
Too few strata leave meaningful score and covariate differences within bins. Too many create sparse cells, unstable estimates, or strata containing only one exposure group. Choose the number using sample size, overlap, balance, minimum cell counts, and precision, without consulting outcomes.
Quantile boundaries create roughly similar numbers from the group used to define them, but they do not guarantee equal total size or balanced exposure groups. Document the requested number, actual number, cut points, and any reassignment used to ensure both groups appear in each stratum.
Examine every stratum
For each stratum, report exposed and comparison counts, propensity-score range, baseline summaries, and any extreme imbalance. A subclass with one exposed participant and many controls can be technically usable yet contribute unstable information. Minimum cell-size rules should be planned, not invented after viewing effects.
| Decision | Question to answer | Evidence to report |
|---|---|---|
| Estimand | Whose effect is targeted? | ATT, ATC, or ATE and population definition |
| Score model | Which baseline factors predict exposure? | Variables, transformations, interactions, link, software |
| Common support | Where can groups be compared? | Plots, rule, exclusions, revised population |
| Subclass count | How finely should scores be divided? | Requested and achieved strata, cut points, cell counts |
| Balance | Are measured covariates comparable? | Weighted and within-stratum diagnostics |
| Pooling | How are stratum effects combined? | Weights, model, estimand alignment, uncertainty method |
Assess balance on covariates, not only the score
A balanced propensity score does not guarantee balanced baseline covariates. The official MatchIt balance guide recommends treating score balance as supplementary and directly assessing the covariates used to construct it.
Calculate standardised mean differences before and after subclassification for every prespecified covariate. Add variance ratios, empirical cumulative distribution differences, density or quantile plots, and categorical proportions. Use the same standardisation denominator before and after design.
Do not rely on balance-test p-values. Their power changes with sample size, so a large p-value can reflect imprecision rather than similarity. A common heuristic is an absolute standardised mean difference below 0.10, with tighter expectations for strong outcome predictors, but no threshold proves exchangeability.
Check aggregate and within-stratum balance
Overall weighted balance can conceal poor comparisons within individual strata. Inspect covariate balance, score spread, and group counts inside each stratum. If important imbalance remains, revise the score model, add nonlinear terms or interactions, change the number of strata, restrict support, or choose a different method.
Assess prespecified functions that matter substantively. If age has a nonlinear relationship with outcome, balance on mean age is insufficient. Check relevant nonlinear terms and interactions without turning diagnostics into an endless search.
| Diagnostic | Useful display | Warning sign |
|---|---|---|
| Common support | Score density or histogram by group | Separated regions or unsupported tails |
| Mean balance | Absolute standardised mean differences | Important covariates remain different |
| Distribution balance | Variance ratios, eCDF and density plots | Similar means hide spread or tail differences |
| Local balance | Stratum-specific summaries and plots | Good aggregate balance masks weak strata |
| Precision | Cell counts, weights, effective sample size | Sparse cells or concentrated influence |
| Target population | Participant flow by exposure group | Exclusions silently change the estimand |
Estimate and pool treatment effects
One approach estimates an effect within each stratum and combines the estimates using weights aligned with the estimand. Another converts subclass membership into weights and fits one marginal outcome model. The official MatchIt effect-estimation guide explains both approaches and notes that properly implemented estimates should be similar.
The weighted approach is sometimes called marginal mean weighting through stratification. Do not treat subclass weights as simple frequencies. Use robust uncertainty estimation appropriate to probability-style weights and the chosen estimator.
For a continuous anxiety score, report an adjusted marginal mean difference with a confidence interval. For a binary recovery outcome, an absolute risk difference is often easier to interpret than an odds ratio. Relative measures can be reported alongside absolute effects when useful.
Align uncertainty with the analysis
Standard errors must reflect weighting, sampling, repeated measurements, and any higher-level clustering such as therapists, schools, clinics, or families. The MatchIt guide advises robust rather than subclass-clustered standard errors for marginal mean weighting when there are few subclasses.
If bootstrapping is used, state whether the complete propensity-score design was repeated in every resample. A bootstrap that treats estimated scores and fixed strata as known answers a different uncertainty question.
Compare subclassification with related methods
Propensity score matching usually forms pairs or fixed-ratio matched sets. Subclassification instead groups score ranges and preserves more participants when overlap is reasonable, but it may leave residual within-stratum imbalance.
Optimal full matching also creates subclasses, but it optimises within-subclass distances and allows variable group ratios. Propensity-score subclassification uses predetermined score cut points. Full matching can often balance better, while simple subclassification may be faster and easier to communicate.
Direct propensity-score weighting gives continuous weights derived from individual scores. It can achieve fine adjustment but may be sensitive to extreme probabilities. Coarsened exact matching forms strata from bins of original covariates rather than one estimated propensity score.
Sensitivity analyses for a credible dissertation
Compare a small set of prespecified, defensible designs without viewing effect estimates. Options include different numbers of strata, a revised nonlinear score model, alternative support restrictions, and subclassification versus full matching or weighting.
For every design, compare balance, participant retention, stratum cell counts, weight distributions, effective sample size, and effect precision. Do not select a specification because it produces the largest or most significant effect.
Conduct an appropriate sensitivity analysis for unmeasured confounding. It cannot prove that hidden bias is absent. It shows how strong an omitted factor would need to be to change the conclusion under stated assumptions.
Document analytic decisions
Save the variable dictionary, causal diagram, score formula, transformations, software versions, random seeds where relevant, support rule, cut points, subclass assignments, weights, balance output, and outcome code. A reproducible archive lets supervisors and examiners distinguish principled design refinement from result-driven analysis.
Common mistakes and practical repairs
Using five strata automatically: compare defensible counts and choose by balance, cell size, and precision.
Checking only propensity-score overlap: inspect balance on every baseline covariate and important function.
Ignoring sparse strata: report cell counts and avoid unstable within-stratum estimates.
Adding post-exposure variables: rebuild the timeline and restrict the score model to genuine baseline information.
Choosing covariates from p-values: use causal reasoning and substantive evidence.
Fitting an unweighted outcome model: use estimand-aligned pooling or subclass weights with suitable robust uncertainty.
Claiming all confounding is controlled: state that the design addresses measured baseline covariates under assumptions.
A reproducible subclassification workflow
- Define eligibility, exposure strategies, time zero, follow-up, outcome, and estimand.
- Select pre-exposure covariates using theory, evidence, and a causal diagram.
- Audit missingness, measurement quality, nonlinearities, outliers, and collinearity.
- Estimate the propensity score without examining exposure-effect estimates.
- Plot overlap and apply a prespecified common-support rule if necessary.
- Create candidate subclasses and record boundaries and group counts.
- Assess aggregate and within-stratum covariate balance.
- Freeze the design after choosing a defensible specification.
- Estimate stratum effects or use subclass weights in a marginal model.
- Run prespecified sensitivity analyses and report remaining limitations.
Frequently asked questions
How many propensity-score subclasses should I use?
There is no universal number. Choose using sample size, overlap, cell counts, balance, and precision. Five or six may be starting points, not automatic recommendations.
Must every subclass contain both exposure groups?
Yes, if it is used to estimate a within-stratum contrast. A stratum containing only one group provides no direct comparison and signals weak overlap or overly fine division.
Is subclassification the same as full matching?
No. Subclassification uses score cut points, usually quantiles. Full matching optimises within-subclass distances and allows variable ratios under a different assignment objective.
Can I use machine learning to estimate the propensity score?
Yes, if the method is documented and yields adequate covariate balance. Predictive performance alone is not sufficient, and complex models do not repair missing confounders.
Should I test balance with p-values?
No. Use descriptive balance measures and plots because hypothesis-test power depends on sample size and does not directly measure design quality.
Can subclassification prove a causal effect?
No. It can improve comparability on measured baseline covariates. Causal interpretation still requires consistency, positivity, no unmeasured confounding, correct timing, and appropriate analysis.
What should I report in the methods chapter?
Report the estimand, timeline, covariates, missing-data plan, score model, support rule, subclass count and boundaries, cell sizes, balance, weights or pooling method, outcome model, uncertainty estimator, software, and sensitivity analyses.
Conclusion
Psychology dissertation propensity score subclassification can create an understandable design for comparing observational groups with similar predicted exposure probabilities. Its strength comes from explicit estimands, credible baseline covariates, adequate overlap, well-composed strata, direct balance assessment, and estimand-aligned outcome analysis.
If you need support, seek ethical methodological guidance that helps you justify and reproduce each decision while keeping data ownership, interpretation, and authorship with you. Responsible support should strengthen your independent understanding rather than promise a preferred result.
