Psychology dissertation multiverse analysis examines how conclusions change across multiple defensible decisions about data processing, measurement, exclusions, covariates, and statistical models. Instead of presenting one selected route as inevitable, it maps the results produced by a prespecified set of reasonable analytic paths.
This approach can strengthen a dissertation when the same research question permits several credible analyses. It does not mean running every imaginable model or searching for a preferred result. A useful multiverse is theory-led, computationally reproducible, and transparent about why each decision is admissible.
What is psychology dissertation multiverse analysis?
A conventional analysis often follows one path: choose an outcome score, apply an exclusion rule, select covariates, fit a model, and report its estimate. Yet another competent researcher might make different defensible choices. A multiverse analysis executes a defined set of those choices and shows how the focal result varies across them.
Steegen, Tuerlinckx, Gelman, and Vanpaemel introduced the method as a way to disclose how conclusions depend on arbitrary data-construction choices. The Association for Psychological Science record for the original multiverse paper identifies transparency and selective reporting as central concerns.
The unit of interest is usually a specification. A specification combines one permitted option from each decision point. If a researcher considers two exclusion rules, three outcome constructions, and two model forms, the full Cartesian set contains 12 paths. Some combinations may be invalid or redundant and should be removed using rules stated before inspecting the focal estimates.
When does a multiverse answer a useful question?
Use a multiverse when several analytically reasonable choices could materially affect the same estimand. Examples include alternative cut-offs for an attention check, plausible scoring rules for a psychological scale, defensible covariate sets, transformations for a skewed outcome, or models that express the same substantive question under different assumptions.
A multiverse is less suitable when one option is clearly required by theory, design, or measurement evidence. Including a known-invalid model does not demonstrate robustness. It merely adds noise. Likewise, models that estimate different causal or descriptive quantities should not be pooled as though they answer one question.
Begin with a precise psychology dissertation hypothesis. Identify the population, exposure or predictor, outcome, comparison, time frame, and target effect. Then ask which decisions preserve that target and which decisions silently change it.
Multiverse analysis is not specification shopping
Specification shopping means trying alternatives and foregrounding the one that supports a desired story. A multiverse has the opposite purpose. It makes the whole justified decision set visible. The primary summary must not depend on which path gives the smallest p-value, largest effect, or most attractive graph.
Distinguish related robustness approaches
Multiverse analysis overlaps with sensitivity analysis, specification-curve analysis, and many-analyst studies, but the designs are not identical. Using the correct label helps readers understand what was varied and what the evidence supports.
| Approach | Main question | Typical output |
|---|---|---|
| Multiverse analysis | How does one result vary across a justified set of analytic decisions? | Distribution or grid of estimates linked to decision paths |
| Specification-curve analysis | What pattern appears across theoretically justified, valid, non-redundant specifications? | Ordered effect curve, decision panel, and sometimes joint inference |
| Targeted sensitivity analysis | Does a result change under one identified assumption or alternative? | Small set of comparative estimates or bounds |
| Many-analyst study | How do independent analysts answer the same question using the same data? | Variation across teams, workflows, and conclusions |
Simonsohn, Simmons, and Nelson describe specification-curve analysis as identifying theoretically justified, statistically valid, non-redundant specifications, displaying them, and evaluating the overall pattern. The accessible PubMed record for specification-curve analysis summarises applications in which some findings appeared robust and others did not.
A many-analyst project delegates the question to independent teams. Silberzahn and colleagues found substantial variation when 29 teams analysed the same dataset and question. The University of Minnesota research record for Many Analysts, One Data Set documents how analytic choices and results differed. A student dissertation will usually conduct one systematic multiverse rather than recruit many teams.
Define the estimand before defining the universe
The estimand is the quantity the analysis aims to learn about. It might be a mean difference in depressive-symptom scores, an association between sleep duration and anxiety, or the change in a treatment contrast after adjustment. Two models may look similar while estimating different quantities.
For example, a coefficient from an unadjusted regression describes an association in the observed sample. A coefficient adjusted for a mediator does not generally represent the same total association. Placing both in one curve without explanation can turn a conceptual change into apparent instability.
Write a short estimand statement before listing decisions. Explain why every branch remains relevant. If branches answer related but distinct questions, group and label them rather than calculating one undifferentiated robustness percentage.
Build a defensible decision tree
Audit the workflow from raw data to final estimate. At each stage, list alternatives a knowledgeable researcher could justify without seeing the focal result. Consult the protocol, instrument manuals, disciplinary norms, diagnostic evidence, and substantive theory.
| Decision domain | Psychology example | Admissibility question |
|---|---|---|
| Sample construction | Exclude responses faster than a prespecified threshold or use an attention-check rule | Does each rule have a documented quality rationale? |
| Outcome scoring | Validated total score or justified subscale score | Do both operationalise the stated construct? |
| Missing data | Complete cases or a compatible imputation strategy | Are assumptions and estimators comparable? |
| Covariates | Age only or age plus a prespecified baseline measure | Does adjustment preserve the target effect? |
| Model form | Linear model or robust regression for the same conditional mean question | Is each model statistically suitable? |
| Influential observations | All valid cases or a prespecified influence sensitivity set | Are rules independent of desired results? |
Exclude options that violate the design or cannot estimate the target. Do not include a dichotomised outcome merely because it is common if the cut-off lacks clinical or theoretical meaning. Do not add every possible covariate combination when some adjust for consequences of the exposure or create collider bias.
Separate arbitrary choices from uncertain assumptions
An arbitrary choice is one of several substantively equivalent options, such as two accepted scoring conventions. An uncertain assumption concerns how the data-generating process works, such as whether missingness is ignorable. Both can be explored, but they require different interpretation. Label the reason for each branch in the decision table.
Control the size of the multiverse
Decision combinations can grow quickly. Five binary decisions produce 32 paths, while ten produce 1,024. A large number is not automatically rigorous. Redundant or incoherent specifications can dominate the display and make conclusions depend on how the analyst counted options.
Use a minimal set that represents genuine uncertainty. Collapse options that yield identical data or models. Remove logically incompatible combinations before execution. If a scoring rule requires complete item responses, it should not be combined with an imputation approach defined only for a composite score.
Create a machine-readable specification table with a unique identifier for every path. Record each decision value, analysis status, sample size, estimate, standard error or interval, and reason for failure. Never silently omit models that did not converge.
Preregister decisions without pretending certainty
A multiverse can be planned before data collection or before the analyst sees outcome-related results. Preregistration should define the research question, admissible options, exclusions, model requirements, primary summaries, and rules for failed specifications. It may also identify a conventional primary analysis alongside the multiverse.
Preregistration does not require claiming that every issue is foreseeable. Mark additions made after data access and explain them. A transparent amendment is more credible than retroactively describing an exploratory branch as planned. Use the site’s psychology dissertation preregistration guide to document timing and deviations.
Plan sample size and precision
A multiverse does not solve low power. Each specification uses the available information differently, and exclusions or complete-case analysis may reduce the sample. Some paths may have unstable estimates even when others appear precise.
Base sample planning on the main estimand and plausible least-efficient admissible analyses. Simulate the intended decision set when practical. Examine bias, interval width, convergence, and the ability to detect practically meaningful variation, not only whether any path reaches a significance threshold.
Explain these goals in the psychology dissertation power analysis. If the dissertation uses existing data, distinguish a precision assessment from prospective power and state the limitations imposed by the fixed sample.
Handle missing data and exclusions consistently
Missing-data choices can change both the sample and the estimand. Complete-case, available-case, maximum-likelihood, and multiple-imputation approaches rely on different assumptions and may not be compatible with every model. Describe the assumed missingness mechanism and ensure the analysis software propagates uncertainty correctly.
Do not call several ad hoc deletion rules a robustness analysis. Each branch needs an ethical and methodological justification. Report how sample composition changes across paths, including whether excluded participants differ on observed variables. The guide to psychology dissertation missing data provides a framework for diagnostics and reporting.
Similarly, distinguish invalid records from valid but influential observations. Correct errors when evidence supports correction. Retain valid cases in the primary data and use transparent sensitivity branches where influence is a concern. Follow a prespecified psychology dissertation outlier analysis.
Execute the multiverse reproducibly
Write modular code that separates data preparation, specification generation, model fitting, extraction, and visualisation. Each result row should be traceable to one complete decision record. Set random seeds where algorithms or imputations are stochastic and save software versions.
Add automated checks. Confirm that outcome coding is consistent, reference groups are correct, model matrices have expected dimensions, sample sizes match the decision rules, and confidence intervals use the intended level. Flag warnings, non-convergence, singular fits, and impossible estimates.
Store results in a tidy table rather than copying values from console output. Re-run the complete pipeline from a clean session. When permissions allow, share code, a data dictionary, and de-identified or synthetic data. This extends the broader psychology dissertation data analysis workflow.
Summarise results without a vote count
Do not reduce the multiverse to the percentage of p-values below .05. Statistical significance mixes estimate size, uncertainty, and sample size, and a threshold discards that information. A majority of significant paths is not a probability that the hypothesis is true.
Show the distribution of effect estimates and intervals. Report a robust range, median or another prespecified descriptive summary when appropriate. Identify which decisions account for major shifts. Make clear when specifications are dependent because they use the same participants and overlapping variables.
| Result feature | Question to answer | Cautious interpretation |
|---|---|---|
| Direction | Do credible paths agree on the sign? | Direction was consistent or changed across defined decisions |
| Magnitude | Does the estimate remain practically similar? | Effect sizes occupied a narrow or broad range |
| Precision | How wide are intervals across paths? | Some conclusions remain uncertain despite similar point estimates |
| Decision influence | Which choices produce the largest changes? | Variation was associated with specific scoring or adjustment decisions |
| Failures | Which paths did not estimate successfully? | Failures limit the represented universe and require explanation |
A specification curve often orders focal estimates from smallest to largest and aligns them with a panel showing the decisions in each path. Preserve effect direction and scale, display uncertainty where legible, and label the null and any practical-equivalence region. Avoid rainbow colour schemes and tiny labels that make the figure inaccessible.
Worked psychology example
Imagine a dissertation asking whether nightly social-media use is associated with anxiety among university students. Reasonable uncertainty exists about whether to use a validated anxiety total or mean score, whether to adjust for age and baseline stress, how to handle a preregistered attention check, and whether to model social-media time linearly or with a justified nonlinear term.
The student defines 16 admissible paths after removing combinations that change the target question. Every path uses the same coding direction and reports the association on a comparable scale. The specification table stores decisions, sample size, coefficient, interval, diagnostics, and convergence status.
Most estimates are positive, but their magnitudes vary. The largest change occurs when baseline stress is included, while scoring choice makes little difference. The student does not claim that the majority “proves” an effect. The discussion explains that conditional and minimally adjusted associations answer different questions and that cross-sectional data cannot establish whether social-media use increases anxiety.
How to report a multiverse in a dissertation
The method section should explain why analytic multiplicity matters, define the estimand, list every decision and option, justify admissibility, describe excluded combinations, and state how results will be summarised. Provide the full specification table in the main text, appendix, or repository.
The results should report the number of planned, completed, failed, and excluded specifications. Give effect ranges and uncertainty, identify consequential decisions, and show a readable figure. If a primary model was preregistered, report it without giving it privileged credibility merely because it was first.
The discussion should separate robust features from fragile ones. Explain whether variation reflects measurement, sample construction, adjustment, or model assumptions. Do not suggest that agreement across a narrow and researcher-selected universe guarantees truth. Conclusions remain limited by design, measurement quality, sampling, and untested assumptions.
Common errors to avoid
- Defining branches after seeing which estimates support the preferred conclusion.
- Including invalid or substantively different models simply to enlarge the universe.
- Counting dependent specifications as independent replications.
- Summarising only significant versus non-significant results.
- Hiding failed models or changing convergence rules between paths.
- Mixing total, direct, conditional, and subgroup effects without labelling estimand changes.
- Presenting every possible covariate subset without causal or theoretical justification.
- Using a complex figure without a downloadable specification table.
Frequently asked questions
How many specifications should a psychology dissertation include?
There is no target number. Include the smallest complete set of theoretically defensible, statistically valid, non-redundant paths needed to represent genuine analytic uncertainty.
Must I run every possible combination?
No. Remove incoherent, invalid, redundant, and estimand-changing combinations using documented rules. Explain why exclusions were made before interpreting results.
Is multiverse analysis the same as p-hacking?
No. P-hacking selectively searches and reports favourable analyses. A well-designed multiverse discloses the full justified set and evaluates how the focal result varies across it.
Can a multiverse replace a primary analysis?
It can be the primary robustness framework, but many dissertations also specify one theory-led primary model. The roles of confirmation, robustness, and exploration should be stated clearly.
Should I count significant specifications?
A count may be descriptive, but it should not be the main conclusion. Prioritise effect magnitude, direction, uncertainty, practical importance, and the decisions associated with change.
What happens when some models fail?
Record every failure, warning, and reason. Report how failures alter coverage of the planned universe. Do not silently substitute a different model or remove inconvenient paths.
Which files should accompany the dissertation?
Provide analysis code, the decision table, specification-level results, software details, and a data dictionary. Share protected data only when consent, governance, and confidentiality permit.
Conclusion
A strong psychology dissertation multiverse analysis turns hidden researcher choices into inspectable evidence. Define one clear estimand, build a justified decision set, execute every valid path reproducibly, and interpret the distribution of effects rather than hunting for thresholds. The method is most informative when it shows both stability and the exact sources of fragility.
If you seek dissertation support, choose ethical guidance that helps you understand, reproduce, and defend your own work. Keep responsibility for all decisions, protect participant information, follow institutional requirements, and acknowledge permitted assistance transparently.
