Psychology dissertation marginal structural models help researchers estimate causal effects when treatment, exposure, and confounding variables all change over time. They are especially useful when a variable predicts later treatment and is itself changed by earlier treatment, a pattern known as treatment-confounder feedback.
This guide explains when a marginal structural model (MSM) is appropriate, how inverse probability weights are constructed, which assumptions support interpretation, and what evidence a dissertation should report. The emphasis is on defensible decisions rather than a mechanical recipe.
What psychology dissertation marginal structural models estimate
An MSM describes a marginal contrast between potential outcomes under different exposure or treatment histories. “Marginal” means the target is averaged over a population rather than conditional on a particular covariate pattern. “Structural” signals that the model represents a causal contrast under explicitly stated identification assumptions.
Inverse probability weighting is one way to estimate an MSM. It is not the model itself. Weighting creates a pseudo-population in which, if the models and assumptions are adequate, measured confounders no longer predict treatment at each decision point. Researchers then fit a comparatively simple weighted outcome model for the treatment history of interest.
The original methodological paper by Robins, Hernán and Brumback introduced MSMs for settings in which time-dependent confounders are affected by previous treatment. Their distinction matters because standard adjustment can be biased when the same variable is both a later confounder and a mediator of an earlier effect. See the indexed original article for the foundational argument.
When treatment-confounder feedback arises in psychology
Consider a longitudinal study of a digital stress-management programme. Current distress influences whether a participant uses the programme next week. Earlier programme use may also reduce current distress. Distress is therefore a confounder of later use and outcome, but it also lies on a pathway from earlier use to the final outcome.
Simply adding every distress measurement to an ordinary regression may block part of the total effect of earlier programme use. Omitting distress leaves later exposure confounded. An appropriately specified MSM can address this measured feedback by weighting each observed treatment decision according to its conditional probability.
Other plausible applications include changing psychotherapy attendance, adherence to a sleep intervention, exposure to workplace stress, repeated social-media use, and access to university counselling. The method is not justified merely because data are longitudinal. It becomes relevant when the causal question concerns a treatment or exposure history and the data-generating structure includes time-varying confounding.
| Research pattern | Why ordinary adjustment can fail | MSM target |
|---|---|---|
| Weekly intervention use and distress | Distress predicts later use but responds to earlier use | Mean outcome under sustained use regimes |
| Changing counselling attendance and symptoms | Symptoms affect attendance and are affected by previous sessions | Risk or mean contrast across attendance histories |
| Repeated exposure with informative dropout | Dropout depends on evolving prognosis | Effect under specified exposure regimes among a represented target population |
Distinguish MSMs from nearby methods
A mixed-effects model addresses correlated repeated observations and heterogeneity across people. It does not automatically solve time-varying confounding. A cross-lagged panel model estimates temporal associations among repeatedly measured variables, but its coefficients are not automatically equivalent to intervention effects.
Baseline propensity score matching can improve comparability for a single exposure decision. Longitudinal MSMs instead address sequences of decisions and multiply period-specific weight components. The psychology primer by Thoemmes and Ong illustrates both point-treatment and longitudinal weighting, including binary and continuous exposures, in an emerging-adulthood context.
MSMs also differ from marginal models fitted with generalised estimating equations. GEE describes population-average associations and handles within-person correlation, but it supplies a causal answer only when the design, adjustment and assumptions identify that answer. “Marginal” in the name is not sufficient.
Define the causal question before modelling
Begin with a target trial-style statement. Identify the eligible population, time zero, treatment options at each decision point, follow-up, outcome, causal contrast, censoring rules and analysis strategy. A vague question such as “Does app use predict wellbeing?” leaves too many choices unresolved.
A clearer estimand might compare mean end-of-semester wellbeing if every eligible student used the programme in at least four of six weeks with mean wellbeing if no eligible student used it. Static regimes like these are easy to explain but may be unrealistic. Dynamic regimes, such as offering extra support when distress exceeds a threshold, can be more relevant but require enough observations compatible with each rule.
Specify whether the target is an intention-to-treat-like assignment effect, an effect of sustained adherence, a cumulative dose effect, or the effect of a dynamic strategy. The exposure definition must be consistent across the protocol, code, results and interpretation.
State the identification assumptions
Consistency and well-defined interventions
Consistency links a participant’s observed outcome to the potential outcome under the treatment history they actually followed. It becomes doubtful when “treatment” includes materially different versions. For example, counselling attendance may combine different therapeutic approaches, session lengths and delivery modes. Define versions or justify treating them as equivalent.
Sequential exchangeability
At each time, treatment must be independent of future potential outcomes after conditioning on the measured history used in the weight model. This is often described as no unmeasured time-varying confounding. It cannot be verified from the data alone. A careful dissertation uses theory, temporal ordering and subject knowledge to defend the confounder set, while acknowledging plausible omissions.
Positivity over observed histories
Every treatment option being compared must have positive probability within relevant covariate histories. If all students with severe current risk receive immediate support, the data cannot identify what would happen under “no support” for that history without extrapolation. Extremely small estimated probabilities are warning signs, not merely computational inconveniences.
Correctly specified models
Weight models must represent the treatment mechanism adequately, including nonlinearity and interactions where needed. The weighted outcome model also requires a functional form that matches the estimand. Cole and Hernán describe consistency, exchangeability, positivity and correct weight-model specification as central conditions and emphasise diagnostics and sensitivity analysis in their weight-construction article.
Build the longitudinal dataset carefully
A person-period dataset usually contains one row per participant per decision time. For every row, distinguish baseline variables, covariate history measured before the current treatment, current treatment, subsequent censoring and the eventual outcome. A timeline is more useful than a long variable list because it exposes impossible ordering.
Do not use a measurement collected after the current treatment decision to predict that same decision. Align windows explicitly when using ecological momentary assessment. If distress is summarised over Monday to Sunday but programme use occurs throughout that week, the direction of influence is ambiguous. A lagged definition may be required.
Record reasons for missing waves and dropout. Missing covariates, treatment nonresponse and loss to follow-up are different processes. Multiple imputation can address some missing-data assumptions, while inverse probability of censoring weights can represent informative loss under its own conditional exchangeability and positivity conditions.
Construct stabilised inverse probability weights
For a binary treatment at time t, estimate the probability of the treatment actually observed. The denominator model conditions on baseline covariates, previous treatment and time-varying confounder history. A stabilised numerator usually conditions on a smaller history, commonly baseline variables and previous treatment, chosen to match the intended marginal model.
The period-specific stabilised weight is the numerator probability divided by the denominator probability. Multiply these components across decision times to obtain each participant’s cumulative treatment weight. For continuous or multinomial exposures, use an appropriate conditional density or category probability rather than forcing a binary model.
If censoring is informative, construct stabilised inverse probability of censoring weights for remaining observed, then multiply treatment and censoring weights. State each numerator and denominator formula. A methods chapter that reports only “IPTW was applied” is not reproducible.
| Component | Typical information | Purpose |
|---|---|---|
| Treatment denominator | Baseline, past treatment, prior time-varying covariates | Remove measured confounding of each decision |
| Treatment numerator | Past treatment and selected baseline terms | Stabilise variance while preserving the target model |
| Censoring component | Predictors of remaining observed | Represent participants with similar histories who were lost |
| Cumulative weight | Product across eligible times | Create the longitudinal pseudo-population |
The R package ipw supports point and longitudinal exposure settings and informative censoring. The WeightIt longitudinal documentation shows stabilised products, balance assessment and effective sample size. Software can calculate weights, but it cannot choose a defensible causal structure for you.
Diagnose the weighted pseudo-population
Inspect weight histograms, quantiles, minima, maxima and the proportion of total weight carried by a small number of observations. Stabilised weights often have a mean near one when models and implementation are coherent, although that is not a pass-fail theorem. Report the effective sample size, which may be far below the raw sample when weights are variable.
Assess covariate balance after weighting at every treatment time and across relevant histories. Use standardised differences and graphical summaries rather than treatment-model discrimination alone. A high area under the curve for the treatment model can actually signal poor overlap.
Investigate extreme weights. They may reflect coding errors, sparse histories, structural non-positivity or model misspecification. Weight truncation or winsorisation can reduce variance but may reintroduce bias. Pre-specify thresholds where possible and present sensitivity analyses using reasonable alternatives rather than selecting the threshold that produces a desired result.
| Diagnostic | Question | Possible response |
|---|---|---|
| Weight distribution | Are a few observations dominating? | Check coding, overlap and functional form |
| Time-specific balance | Are prior confounders balanced at each decision? | Revise models using principled nonlinearities or interactions |
| Effective sample size | How much information remains after weighting? | Temper precision claims and reconsider the estimand |
| Truncation sensitivity | Do conclusions depend on tail handling? | Report multiple justified thresholds transparently |
Fit and interpret the marginal outcome model
Fit the chosen outcome model with the cumulative weights. The link function should match the outcome and estimand: an identity link for a mean difference, log link for a risk ratio, or suitable survival model for a time-to-event outcome. Include treatment-history terms required by the causal contrast and variables used for stabilisation when appropriate.
Use uncertainty estimates that respect weighting and repeated observations. Robust sandwich standard errors are common; a participant-level bootstrap may also be defensible if the entire estimation procedure, including weight estimation, is repeated within each resample. Explain the approach rather than relying on default software output.
Translate coefficients into interpretable contrasts such as predicted mean differences or risks under specified regimes, with confidence intervals. Avoid claiming that weighting “proves causality.” The estimate has a causal interpretation only relative to the stated estimand, measured variables, model choices and assumptions.
A worked psychology example
Imagine 600 university students measured weekly for six weeks. Each week records prior-week distress, sleep difficulty, academic workload and use of an optional digital coping programme. The primary outcome is wellbeing at week seven. Earlier use may improve distress, while current distress influences later use, creating treatment-confounder feedback.
The dissertation compares two hypothetical regimes: use in at least four weeks versus no use. At each week, a denominator model predicts observed use from baseline variables, previous use, and earlier distress, sleep and workload. The numerator uses baseline variables and previous use. Period weights are multiplied, then combined with censoring weights based on predictors of remaining in follow-up.
Suppose the raw sample is 600 but the weighted effective sample size is 410. That illustrative loss of information would prompt careful precision reporting. If balance improves for most variables but current distress remains imbalanced at week five, the analysis is not ready. The researcher should revisit timing, functional form and overlap before interpreting the outcome model.
The final results would report regime-specific predicted wellbeing, their contrast and uncertainty, alongside weight distributions, balance plots, effective sample size and truncation sensitivity. The numbers in this example illustrate workflow only and are not empirical findings.
Alternatives and sensitivity analyses
The parametric g-formula models the outcome process and simulates outcomes under intervention strategies. G-estimation targets parameters of structural nested models. Doubly robust approaches combine treatment and outcome models, while targeted learning methods can use flexible nuisance estimation. These approaches have different estimands, assumptions and implementation demands; they are not interchangeable labels.
Compare substantive conclusions across reasonable specifications. Consider alternative exposure definitions, lag structures, confounder sets, treatment models, censoring models and truncation thresholds. Negative-control or quantitative bias analyses may probe particular threats, but no sensitivity analysis recovers information about an unmeasured confounder without additional assumptions.
A transparent MSM can be more credible than a superficially sophisticated model whose causal structure is hidden. If overlap is poor, narrow the target population or redefine the regime rather than forcing a population-wide claim.
Reporting checklist for a psychology dissertation
- State the population, time zero, decision schedule, regimes, follow-up, outcome and causal contrast.
- Show a timeline or directed acyclic graph and justify temporal ordering.
- List every treatment and censoring numerator and denominator model.
- Explain why each covariate is included using subject knowledge, not significance testing.
- Report weight summaries, time-specific balance, overlap and effective sample size.
- Describe truncation, missing-data handling, outcome model and uncertainty estimator.
- Separate identification assumptions from testable diagnostics and acknowledge residual confounding.
- Provide reproducible code, variable definitions and a decision log, subject to ethical data-sharing limits.
Common errors to avoid
Do not treat every time-varying predictor as a confounder. Some may be mediators, colliders or consequences of selection. Do not select covariates using automated p-values. Do not interpret a good treatment-model fit as evidence of balance. Do not hide extreme weights or report only truncated results.
Avoid changing the estimand after seeing the results. Do not combine treatment and censoring weights without verifying their time ordering. Finally, do not generalise beyond participants and histories supported by the data. These problems are design and interpretation failures, not cosmetic reporting issues.
Frequently asked questions
What is a marginal structural model in simple terms?
It is a model for comparing population-average potential outcomes under different treatment or exposure histories. Weighting is commonly used to estimate it when measured confounders change over time.
When should a psychology dissertation use an MSM?
Use one when the question concerns longitudinal treatment or exposure histories and a time-varying covariate both predicts later treatment and is affected by earlier treatment. Longitudinal data alone are not enough.
Does inverse probability weighting remove all confounding?
No. It addresses confounding represented by the measured history in correctly specified models, subject to consistency, exchangeability and positivity. Unmeasured or poorly measured confounding can remain.
How large should an MSM sample be?
There is no universal minimum. Requirements depend on treatment frequency, number of time points, outcome prevalence, overlap, model complexity and weight variability. Plan using simulations that mirror the proposed design and report the weighted effective sample size.
Should extreme weights always be truncated?
No. First investigate their cause. Truncation may improve precision but can add bias. Use pre-specified, justified thresholds and show sensitivity analyses.
Can MSMs be used with continuous exposures?
Yes, but treatment models estimate conditional densities rather than binary probabilities. Continuous exposures often make positivity and model specification more demanding, so diagnostics are essential.
Are MSM results automatically causal?
No. Causal interpretation depends on a well-defined estimand, correct temporal ordering, adequate measurement, consistency, sequential exchangeability, positivity and appropriate models. The dissertation should state these conditions explicitly.
Conclusion
Psychology dissertation marginal structural models are most valuable when a clear longitudinal causal question meets treatment-confounder feedback. Good practice begins with the estimand and timeline, not the software. It continues through transparent weight construction, balance and positivity checks, defensible uncertainty estimation and honest sensitivity analysis.
If you need methodological support, seek ethical guidance that helps you justify the design, understand the analysis and retain authorship of every decision. A consultation should strengthen your independent research rather than supply unverifiable results or replace academic supervision.
