Psychology postgraduate reviewing nested participant groups and multilevel trajectories

Psychology dissertation multilevel modelling lets you analyse outcomes when observations are nested, repeated, or grouped rather than statistically independent. Common examples include daily reports within participants, pupils within classrooms, clients within therapists, and responses to several stimuli within the same person. A well-specified model separates variation at the relevant levels and estimates relationships without pretending that every row of data supplies entirely new information.

This guide explains when a multilevel model is justified, how to translate a psychology design into levels, and how to plan fixed effects, random effects, centring, power, diagnostics, and reporting. It focuses on defensible decisions rather than software menus. Use it with the broader psychology dissertation regression analysis guide when reviewing coding, residuals, interactions, and coefficient interpretation.

Why nested psychology data need special treatment

Conventional regression assumes that residuals from different observations are independent after accounting for the predictors. That assumption is doubtful when observations share a context. Two mood reports from one participant tend to resemble each other. Pupils taught in the same classroom may share a teacher, timetable, peer environment, and recruitment process. Responses to the same stimulus may resemble one another across participants.

Ignoring that dependence can produce standard errors that are too small and conclusions that are too confident. Aggregating all pupil scores to classroom averages avoids pseudo-replication but discards individual variation and usually reduces information. A multilevel model represents the dependence explicitly through level-specific random components while retaining observations at their measured level.

The University of Bristol Centre for Multilevel Modelling lists mixed models, hierarchical models, random-coefficient models, and variance-component models among the related names. Terminology varies across disciplines and software. In a dissertation, define the model structure instead of relying on a label alone.

Identify levels from the design, not the spreadsheet

A level represents a source of dependence or sampling. It is not simply a column or the order in which data are entered. In a two-level school study, pupils may be Level 1 and schools Level 2. In a diary study, days may be Level 1 and participants Level 2. In an experiment where every participant responds to multiple items, responses are crossed by participants and items rather than items being neatly nested within participants.

Psychology design Lower level Higher level or grouping Possible model focus
Daily stress diary Daily observations Participants Within-person stress change and between-person differences
School intervention Pupils Classrooms or schools Intervention effect with classroom or school variation
Therapy service study Clients Therapists or clinics Client predictors and between-therapist outcome variation
Repeated laboratory task Trials Participants Trial-level condition effects and participant heterogeneity
Participants rate shared stimuli Responses Participants crossed with items Generalisation across both people and stimuli
Longitudinal cohort Measurement occasions Participants, possibly within sites Growth, time-varying predictors, and site variation

Draw the structure before modelling. Mark which variables change within a person, which vary only between people, and which belong to a school, therapist, clinic, item, or site. Include crossed classifications and multiple memberships when they genuinely exist. For example, a pupil who changes school is not perfectly nested in one school across the study.

What fixed and random effects mean

Fixed effects describe average relationships

Fixed effects are the regression coefficients ordinarily reported as intercepts, slopes, group differences, and interactions. A fixed condition coefficient can represent the average difference between intervention and comparison conditions across the populations represented by the random effects. The intercept is the expected outcome when predictors equal zero, which is only useful when zero has a defensible meaning.

A coefficient is conditional on every other term in the model. In a longitudinal model with time, treatment, and their interaction, the treatment coefficient may represent the group difference at the time value coded zero, while the interaction represents a difference in rate of change. Centre or recode time so that the reference occasion answers a useful question.

Random intercepts represent baseline heterogeneity

A random-intercept model allows clusters to have different expected outcome levels. In a therapy study, clinics may differ in average symptom scores after accounting for measured predictors. The fixed intercept is the overall expected value, while the random-intercept variance describes remaining between-clinic heterogeneity.

Random effects are usually treated as draws from a distribution rather than estimated as unrelated fixed indicators for every cluster. This partial-pooling approach draws imprecise cluster estimates toward the overall pattern, with the extent of shrinkage depending on cluster size and estimated variance. Avoid ranking individual clinics or teachers as though their conditional predictions were error-free performance scores.

Random slopes represent heterogeneous relationships

A random slope allows a within-cluster effect to vary. The association between daily sleep and next-day mood may be stronger for some participants than others. A condition effect may vary across participants or experimental items. The model can also estimate covariance between random intercepts and slopes, although that covariance can be difficult to estimate with limited clusters.

Random slopes should follow the design and intended generalisation, not be added mechanically. Barr’s methodological simulations show why confirmatory analyses of replicated participant or item data can become anti-conservative when relevant variation in experimental effects is omitted. Later methodological work emphasises balancing design justification, convergence, precision, and error control rather than treating either the most complex or simplest structure as universally correct.

The intraclass correlation coefficient

For a simple two-level continuous-outcome model, the intraclass correlation coefficient, or ICC, is the proportion of residual outcome variance located between clusters. It is calculated from the empty model as between-cluster variance divided by total between-cluster plus within-cluster variance. It can also be understood as the expected correlation between two randomly selected observations from the same cluster under that model.

An ICC of .10 would mean that 10% of the modelled outcome variance lies between clusters and 90% within clusters. It does not mean that the grouping variable explains 10% of each individual’s score, nor does it establish a causal contextual effect. For binary, count, or ordinal models, the ICC is scale-dependent and may require a latent-variable or simulation-based definition.

Do not use a non-significant random-intercept test or a small sample ICC as the sole permission slip for single-level analysis. Dependence is created by the design, and even a modest ICC can matter when clusters are large. The number of clusters also controls information for higher-level effects and variance estimates.

Separate within-person and between-person effects

A time-varying predictor can carry two distinct questions. In a diary study, people who usually sleep more than other people may report better average mood. Separately, a person may report better mood on nights when they sleep more than their own usual amount. A single uncentred sleep coefficient can blend these between-person and within-person associations.

To separate them, calculate each participant’s mean sleep and the deviation of each occasion from that participant mean. Include both in the model. The person mean estimates a between-person association, while the deviation estimates a within-person association. This is often called a within-between or contextual-effects specification.

Grand-mean centring subtracts the overall sample mean. It makes zero more interpretable but does not automatically separate within-cluster and between-cluster information. Group-mean centring subtracts a cluster-specific mean and can isolate a within-cluster contrast when the cluster mean is also represented appropriately. Describe exactly which mean was used, whether it came from observed occasions, and how missing occasions affect it.

Choose the outcome distribution correctly

Linear mixed models suit a continuous outcome when the conditional residual and variance assumptions are reasonable. Binary outcomes need a generalized mixed model such as mixed-effects logistic regression. Counts may need Poisson or negative-binomial models. Ordered ratings may be better represented by ordinal mixed models than by treating a single bounded item as continuous.

The distribution applies to the outcome conditional on predictors and random effects, not to every raw variable. Inspect residuals at the appropriate levels, fitted-versus-residual plots, influential clusters, predicted probabilities, and the observed range. The normality testing guide explains why raw-score normality is not the right diagnostic for a linear model.

Outcome Typical multilevel family Interpretation focus Important check
Approximately continuous scale score Linear mixed model Conditional mean differences Residual form, variance, influence, random effects
Binary response Logistic mixed model Conditional log odds, odds, or predicted probability Separation, sparse clusters, calibration
Count Poisson or negative-binomial mixed model Conditional rate ratio Exposure, overdispersion, excess zeros
Ordered category Ordinal mixed model Cumulative category probabilities Threshold and proportional-odds assumptions
Time to event Multilevel survival model Hazard or survival difference Censoring and time-scale assumptions

Plan sample size across all levels

Multilevel power depends on more than total sample size. It depends on the number of higher-level units, observations per unit, imbalance, ICCs, effect locations, random-slope variance, outcome distribution, missingness, and the exact hypothesis. A study with 600 pupils in six schools does not provide the same information for a school-level intervention as 600 pupils spread across 60 schools.

Simulation and empirical work on multilevel sample-size planning shows that information at higher levels is often a major constraint, especially for variance components, cluster-level effects, and their standard errors. No finding supplies a universal minimum. Rules such as “30 clusters” cannot replace design-specific planning because acceptable performance changes with model complexity and inferential target.

Simulation is often the most transparent method. Specify plausible fixed effects, variance components, cluster sizes, missingness, and the planned model; repeatedly generate data; fit the model; and estimate the probability of meeting the stated decision or precision criterion. Vary uncertain inputs rather than presenting one optimistic scenario. The power analysis guide explains how to justify effect inputs and sensitivity scenarios.

Build the model from the research question

Specify the analysis before fitting alternatives

State the outcome, link function, levels, fixed predictors, contrasts, random structure, centring, interactions, estimation method, missing-data assumptions, and primary hypothesis. Explain which population of participants, items, schools, therapists, or sites the random effects are intended to represent. Record this in the preregistration where feasible.

Use an empty model for description, not ritual

An intercept-only model partitions baseline variance and provides an ICC for simple Gaussian structures. It can help reveal whether the proposed levels are estimable. It is not always the correct comparator for longitudinal trends, crossed random effects, or generalized outcomes, and it should not become a mechanical gatekeeping test.

Add theoretically justified fixed effects

Enter predictors in blocks that reflect the design and hypotheses, not a search for significance. Code categorical variables so the reference group and contrasts are clear. Include lower-order terms for interactions unless a substantively justified parameterisation says otherwise. Report how covariates were chosen and measured.

Define random effects from replication

Ask which intercepts and effects could vary across sampled units. Within-participant manipulations often motivate participant-level random slopes; shared stimuli can motivate item random intercepts and slopes. A slope cannot be estimated at a level where its predictor never varies. Keep the structure consistent with design and planned generalisation.

Compare models carefully

Likelihood-ratio tests apply to nested models under specific estimation and boundary conditions. Information criteria compare a trade-off between fit and complexity but do not prove a model true. When comparing fixed effects in Gaussian mixed models, maximum-likelihood rather than restricted maximum-likelihood fits are commonly required. Refit the selected fixed structure using the prespecified estimation approach for final inference.

Convergence and singular fits are evidence, not nuisances

A convergence warning means the optimiser did not establish a reliable optimum under the requested settings. A singular fit usually means at least one random-effect dimension is estimated at or near a boundary, or correlations imply a lower-dimensional structure. Neither problem should be hidden by copying coefficients from the output.

First verify coding, identifiers, cluster counts, predictor variation within clusters, scaling, sparse cells, and data errors. Then inspect the random-effect estimates and optimiser diagnostics. Try a documented alternative optimiser or sensible rescaling as a diagnostic, not as a hunt for preferred p values. If simplification is needed, apply a principled rule connected to the design and report the attempted and final structures.

Do not delete a theoretically important random slope only because its p value is large. Variance testing occurs at a boundary and is often imprecise. Conversely, retaining a structure that the available data cannot identify does not create information. State the limitation and narrow the generalisation when necessary.

Missing data and unbalanced observations

Multilevel models can use participants with different numbers or timings of observations under a likelihood framework, so they do not require every person to have a complete rectangular record. This flexibility does not make missingness harmless. Validity still depends on the model, included predictors, and a defensible missing-at-random assumption or sensitivity analysis.

Describe missingness at each level. A missing diary occasion differs from complete participant dropout or a clinic leaving the study. Include variables related to missingness when justified, consider multiple imputation that preserves clustering, and compare plausible assumptions. Consult the missing-data guide before defaulting to complete cases.

Diagnostics and sensitivity checks

Diagnostics must reflect the model. For a linear mixed model, examine conditional residuals, residual variance patterns, cluster-level random effects, fitted trajectories, influential units, and whether the chosen covariance structure captures repeated observations. For generalized models, examine calibration, overdispersion, zero inflation, residual simulations, and sparse outcomes.

Useful sensitivity checks might compare reasonable time functions, justified random structures, robust standard errors, alternative outcome families, or influential-cluster exclusions. Each check should answer a stated concern. Do not present dozens of undisclosed models and report only the one that crossed a significance threshold.

Problem signal Question to investigate Transparent response
Convergence warning Are coding, scaling, optimiser, or model complexity responsible? Diagnose, document attempts, and avoid unverified estimates
Singular fit Which variance or correlation is unsupported? Use a design-led simplification or qualify generalisation
Nonlinear time pattern Does a straight slope misrepresent change? Use planned polynomial, spline, or piecewise terms
Influential cluster Is the result dominated by one school, clinic, person, or item? Verify data and report influence sensitivity
Heteroscedastic residuals Does variance differ by time, condition, or cluster? Model variance or use a justified robust approach
Sparse binary outcomes Are estimates separated or weakly identified? Consider penalised or Bayesian methods and cautious inference

How to interpret cross-level interactions

A cross-level interaction asks whether a lower-level relationship varies with a higher-level predictor. For example, does the within-pupil association between daily academic stress and mood vary with classroom climate? Include the relevant main effects, use centring that separates within- and between-pupil information, and define the values at which the interaction is probed.

Report predicted outcomes or simple slopes at meaningful moderator values with uncertainty. Avoid saying that one variable “causes the relationship” unless the design supports causal inference. An interaction can reflect scale choice, residual confounding, measurement error, or limited overlap. The moderation analysis guide provides plotting and interpretation principles that also apply here.

Report the model so another researcher could reproduce it

Describe the data structure before the results. Give the number of units at every level, typical and range of cluster sizes, repeated occasions, outcome distribution, coding, centring, software, package and version, estimator, and missing-data handling. State the complete fixed and random formula, including crossed effects and covariance restrictions.

For results, report fixed-effect estimates, standard errors or intervals, test method, exact p values where used, variance components, random-effect correlations when stable, ICCs where meaningful, and model-fit information. Translate log odds or log rates into odds ratios, rate ratios, or predicted probabilities while retaining uncertainty. Include a table that separates fixed effects from random components.

Do not report only a likelihood-ratio test or a software screenshot. Explain the effect in the outcome’s units and show predicted values for important interactions or nonlinear change. Save syntax and model objects so every value can be traced. The results-section guide helps organise tables and narrative without duplicating output.

Common multilevel modelling mistakes

  • Treating all rows as independent: map the sampling and repeated-measure structure before analysis.
  • Calling every grouping variable a level: distinguish nested, crossed, and multiple-membership structures.
  • Using one mixed coefficient for two questions: separate within-unit and between-unit effects when needed.
  • Choosing random effects by significance alone: connect them to replication and generalisation.
  • Relying on an ICC threshold: consider design, cluster size, target effect, and precision.
  • Using total sample size for power: plan the numbers of clusters and observations per cluster.
  • Ignoring warnings: diagnose convergence and singularity before interpreting coefficients.
  • Reporting software defaults as method: name the family, link, estimator, random structure, and tests.

Frequently asked questions

When does a psychology dissertation need multilevel modelling?

Consider it when observations are repeated within people, people are grouped within settings, responses share items, or another design feature creates dependence. The final choice depends on the outcome, hypothesis, sampling, cluster count, and intended generalisation.

Is multilevel modelling the same as repeated-measures ANOVA?

They overlap for some balanced designs, but mixed models can represent irregular timing, unequal observations, continuous time, several random effects, and different outcome distributions. Greater flexibility also requires more explicit assumptions and diagnostics.

How many clusters do I need?

There is no universal number. Requirements depend on whether the target is a fixed effect, variance component, cross-level interaction, or prediction, plus effect size, ICC, balance, and complexity. Use design-specific simulation and show sensitivity to uncertain inputs.

Should I centre predictors?

Centre when it makes the intercept or interaction meaningful or separates within-cluster from between-cluster effects. Grand-mean and group-mean centring answer different questions. State the calculation and include the relevant cluster mean when decomposing effects.

Can SPSS run multilevel models?

Yes. SPSS Mixed Models and Generalized Linear Mixed Models cover many designs. R, Stata, SAS, MLwiN, and Bayesian platforms offer further options. Defaults and inferential tests differ, so preserve syntax and report the exact implementation.

What if my model is singular?

Check data, coding, within-cluster variation, scaling, cluster counts, and the random-effect estimates. If the structure is not identifiable, simplify using a prespecified or design-led rule and explain how that changes the population of generalisation.

Conclusion

Strong psychology dissertation multilevel modelling begins with the dependency structure and the scientific estimand. Translate the design into levels, separate within-unit and between-unit effects, specify fixed and random components from theory and replication, and plan power at every level. Treat convergence, missingness, and diagnostics as substantive evidence rather than technical obstacles.

For ethical academic support, request feedback on your model rationale, code, diagnostics, or reporting. A responsible reviewer can identify weaknesses and explain alternatives, but should not invent observations, hide failed models, or produce results that you cannot reproduce and defend.

Authoritative references

Leave a Reply

Your email address will not be published. Required fields are marked *