Psychology researcher organising multiple outcome comparisons into a controlled analysis plan

psychology dissertation multiple comparisons arise whenever one project evaluates several outcomes, groups, time points, contrasts, subgroups, models, or questionnaire subscales. Each additional test creates another opportunity for a chance result to look convincing. The solution is not to avoid rich research questions, but to define the test family, distinguish confirmatory from exploratory work, choose an error rate deliberately, and report every analysis transparently.

This guide explains how to identify multiplicity, decide which hypotheses belong together, select a defensible adjustment, implement common procedures, and interpret adjusted results without reducing a dissertation to a hunt for statistical significance. Examples cover experiments, surveys, repeated-measures studies, factorial designs, and secondary-data projects.

Essentials to Psychology Dissertation Multiple Comparisons

A conventional significance threshold controls the probability of a Type I error for one test under its null model. It does not automatically control the chance of at least one false rejection across a collection of tests. If many independent null hypotheses are tested at .05, the chance of obtaining at least one result below .05 rises with the number of tests. Correlation among tests changes the exact probability, but it does not make multiplicity disappear.

The practical problem is broader than post-hoc comparisons after ANOVA. A dissertation may test depression, anxiety, stress, wellbeing, and sleep; compare three intervention groups at four occasions; inspect several subscales; add unplanned covariates; and repeat models within demographic subgroups. These decisions can create multiple testing families even when the software presents each result in a separate table.

The American Statistical Association statement on p-values emphasises that conclusions should not depend only on whether a p-value crosses a threshold, that transparent reporting is essential, and that statistical significance does not measure effect size or importance. Multiplicity control therefore belongs within a wider analysis strategy that includes estimation, uncertainty, theory, and disclosure.

Table of Contents

A psychology example

Suppose a student evaluates a mindfulness programme using four outcomes measured immediately after treatment and at three-month follow-up. Separate treatment-control tests for every outcome and occasion produce eight tests. If the student also compares results by gender and prior therapy, the number grows further. Reporting only the smallest unadjusted p-value would conceal the search process and exaggerate the apparent evidence.

A stronger approach specifies one primary outcome and occasion, treats related secondary outcomes as a defined family, labels subgroup analyses exploratory unless prespecified, and reports effect estimates with confidence intervals. This preserves the substantive value of the study while making the strength of each claim proportionate to the analysis.

Planning psychology dissertation multiple comparisons

Multiplicity is best managed before data analysis. Start by translating the research aims into explicit hypotheses and planned comparisons. Then create a test inventory showing every outcome, predictor, contrast, time point, subgroup, interaction, mediator, and alternative model that might support a claim.

Define the unit of inference

Ask what statement the dissertation is trying to protect from false-positive conclusions. A family might include all pairwise comparisons among treatment groups for one outcome, all secondary outcomes addressing one broad claim, or all coefficients used to evaluate one theoretical mechanism. Tests do not belong in the same family merely because they appear in the same software output, and they should not be separated merely because they appear in different chapters.

There is no universal mechanical rule for defining a family. The decision should follow the claims a reader could make from the results. If success on any one of six wellbeing outcomes would be used to claim that an intervention improves wellbeing, those outcomes usually belong together. If two outcomes address clearly distinct, independently motivated research questions, separate families may be defensible, provided that distinction was not invented after seeing the results.

Separate confirmatory and exploratory analyses

Confirmatory hypotheses are specified before examining the relevant outcomes and are tied to a planned inferential procedure. Exploratory analyses can be valuable, but their evidential status differs. Label them clearly, present all tests conducted within the exploration, and frame findings as patterns for future investigation rather than settled confirmation.

Psychology Dissertation Multiple Comparisons

Preregistration can document primary outcomes, model specifications, exclusion rules, contrasts, and adjustment methods. The site’s guide to psychology dissertation preregistration explains how to record those decisions while retaining a transparent route for justified deviations.

Potential family Psychology example Planning question
Pairwise group contrasts CBT, supportive counselling, and wait-list groups Are all pairs relevant, or only comparisons with the control?
Multiple outcomes Anxiety, depression, sleep, and functioning Would any favourable outcome support the same overall claim?
Repeated occasions Post-test, three months, and six months Which occasion is primary, and how are trajectories modelled?
Subscales Cognitive, somatic, and affective symptom scores Are subscales theoretically distinct and psychometrically supported?
Interactions Treatment by age, gender, and baseline severity Which moderators were hypothesised before analysis?
Alternative models Several covariate sets or exclusion thresholds Which is primary, and which are sensitivity analyses?

Understand the error rate you want to control

An adjustment method is meaningful only in relation to its target error rate. Familywise error rate and false discovery rate answer different questions. Neither is automatically correct for every dissertation.

Familywise error rate

The familywise error rate, or FWER, is the probability of making at least one Type I error within a defined family. Controlling FWER is useful when even one false positive would seriously distort the principal conclusion, when the number of prespecified hypotheses is modest, or when decisions are high stakes.

Bonferroni controls FWER by comparing each p-value with alpha divided by the number of tests, or equivalently multiplying each p-value by the family size. It is simple and works without requiring independence, but it can be conservative. The NIST guidance on Bonferroni comparisons describes its use for a finite, prespecified set of contrasts and simultaneous confidence intervals.

False discovery rate

The false discovery rate, or FDR, concerns the expected proportion of false rejections among the rejected hypotheses over repeated applications of a procedure. It is often considered when a study examines a larger, coherently defined set of hypotheses and tolerates some false discoveries in exchange for greater power.

FDR control is not a promise that a fixed percentage of the significant results in one dissertation are false. Benjamini’s retrospective account of the false discovery rate explains the distinction between controlling an error rate and estimating how many results in one realised set are wrong.

Per-comparison error

Leaving every test at .05 controls the error probability for each individual test, not the collective claim. This may be acceptable for a single genuinely primary test or for explicitly descriptive work that avoids dichotomous claims. It is not an adequate defence when a researcher searches a large family and highlights whichever result happens to cross .05.

Choose an adjustment method in Psychology dissertation multiple comparisons

Method selection should consider the hypothesis family, dependence among tests, direction of comparisons, whether hypotheses were prespecified, and whether the objective is strong protection against any false rejection or a balance between discoveries and errors.

Method Controls Useful setting Important caution
Holm FWER Small or moderate families of planned tests Report the whole family and adjusted p-values, not only retained results
Bonferroni FWER Few prespecified contrasts or simultaneous intervals Can lose power as the family grows
Tukey FWER for pairwise mean comparisons All pairwise comparisons following a suitable group model Not a general solution for outcomes, subgroups, and models
Dunnett FWER for comparisons with a control Several interventions compared with one control Does not cover every intervention-to-intervention contrast
Benjamini-Hochberg FDR Larger families where discovery is the central objective Its assumptions and interpretation must be stated accurately
Model-based contrasts Depends on procedure Repeated measures, mixed models, or estimated marginal means The software option still requires a declared family and method

Why Holm is often a sensible FWER default in Psychology dissertation multiple comparisons

Holm’s sequential procedure orders p-values and compares them with progressively less stringent thresholds. It controls FWER and is at least as powerful as the simple Bonferroni procedure. The official R documentation for adjusted p-values lists Holm, Bonferroni, Hochberg, Hommel, Benjamini-Hochberg, and Benjamini-Yekutieli procedures and notes their different targets and assumptions.

When specialised contrast procedures help

If the research question concerns every pair of group means, Tukey-type simultaneous comparisons may be more efficient than treating the contrasts as an arbitrary list. If several treatments are compared only with one control, Dunnett-type comparisons fit that restricted family. Planned contrasts grounded in theory can answer more focused questions than an indiscriminate set of pairwise tests.

The choice must follow the research aim. Running several procedures and reporting the one that yields the most favourable result simply creates another layer of multiplicity.

Do not rely on an omnibus test alone

A significant omnibus ANOVA answers whether the model contains some group difference, but it does not identify which groups differ. Follow-up contrasts form their own inferential family. Conversely, a non-significant omnibus result does not automatically license or forbid every planned contrast; the logic depends on the analysis plan and procedure.

Multiplicity can also be hidden inside factorial designs. An open-access psychology study on hidden multiplicity in exploratory multiway ANOVA shows that main effects and interactions create several hypotheses even before post-hoc comparisons begin. The authors discuss FWER control, FDR control, omnibus approaches, and preregistration as possible responses.

Use the site’s psychology dissertation ANOVA guide for model assumptions and contrast structure, then document multiplicity separately rather than assuming that the omnibus F test resolves every issue.

Multiple outcomes and repeated measurements

Psychology dissertations often measure several related constructs because no single outcome captures the phenomenon. Correlation among outcomes may reduce the inflation relative to independent testing, but it does not justify ignoring the family. The analysis plan should identify a primary outcome, define secondary outcomes, and explain whether they support one shared claim or distinct hypotheses.

Primary and secondary outcomes

A primary outcome should be chosen for substantive reasons, not because it produced the smallest p-value. Secondary outcomes can deepen interpretation, test mechanisms, or reveal unintended effects. Their results remain valuable when estimates are imprecise or do not survive adjustment.

Report effect sizes and intervals for every outcome. The guide to psychology dissertation effect sizes explains why magnitude and precision remain central regardless of the testing decision.

Repeated time points

Testing each group difference separately at baseline, post-test, and several follow-ups often produces a fragmented family. A longitudinal or mixed model can estimate trajectories and group-by-time interactions more coherently. Planned contrasts can then target clinically or theoretically meaningful occasions with simultaneous inference where appropriate.

Model-based analysis does not erase multiplicity when many outcomes, contrasts, or interactions are examined. Review the psychology dissertation longitudinal study guide when observations are repeated over time.

Subgroups, moderators, and researcher degrees of freedom

Subgroup analysis is tempting because an overall effect may differ across ages, genders, cultures, baseline severity levels, or diagnostic groups. Yet testing the effect separately in many subgroups and comparing which subgroup is significant is misleading. A significant result in one group and a non-significant result in another does not itself prove that effects differ.

Test the interaction directly when moderation is the research question. Prespecify a small set of theoretically motivated moderators, retain continuous variables where defensible, and avoid repeatedly changing cut-points. The psychology dissertation moderation analysis guide covers interaction interpretation and probing.

Other researcher choices can expand the effective analysis set: excluding different outliers, transforming outcomes several ways, adding or removing covariates, changing scoring rules, or switching between one-tailed and two-tailed tests. Sensitivity analyses should test robustness, not provide a menu from which the most favourable result is selected.

Worked Psychology dissertation multiple comparisons example

Consider a dissertation comparing online cognitive behavioural therapy, guided self-help, and wait-list control for student anxiety. Anxiety at post-test is the primary outcome. Depression and functioning are secondary outcomes. Anxiety is also measured at three-month follow-up.

Step 1: state the claims

The primary claim is that each active intervention reduces post-test anxiety compared with wait-list control. This produces two planned control comparisons. A secondary claim is that the interventions improve depression and functioning. Follow-up anxiety addresses maintenance rather than immediate efficacy.

Step 2: define families

The two primary control contrasts can form one FWER-controlled family. The four secondary treatment-control comparisons across depression and functioning can form a second family. The two follow-up anxiety contrasts may form a third family if maintenance is treated as a separate research question. The dissertation must justify this structure before inspecting outcomes.

Step 3: choose procedures

A Dunnett-type procedure can compare both interventions with the same control for the primary outcome. Holm adjustment can provide a simple alternative for the two planned contrasts. For the broader secondary family, Holm offers strict FWER control, while Benjamini-Hochberg may be defensible if the stated aim is to identify promising secondary signals and all results are reported.

Step 4: report estimates

For each contrast, report the estimated mean difference, confidence interval, standardised effect where useful, raw p-value, adjusted p-value, adjustment method, and family definition. Do not change the substantive interpretation solely because an adjusted p-value moves slightly above or below .05.

Result Estimate Interval Raw p Adjusted p Interpretation
Online CBT versus control -4.2 anxiety points -6.8 to -1.6 .002 .004 Data support a lower post-test mean, with magnitude and uncertainty reported
Guided self-help versus control -2.1 anxiety points -4.6 to 0.4 .098 .098 Estimate favours treatment, but the interval includes little or no difference
Online CBT versus self-help -2.1 anxiety points -4.5 to 0.3 .086 Not in primary family Exploratory comparison, clearly labelled and not used to redefine success

The numbers are illustrative, not evidence from a real trial. Their purpose is to show that transparent interpretation incorporates the estimate, interval, adjustment, and status of the comparison.

Software implementation without mechanical reporting

Most statistical packages can adjust p-values, estimate simultaneous intervals, or construct model-based contrasts. The difficult part is not clicking the option; it is deciding which tests form the family and which error rate matches the claim.

SPSS

SPSS provides procedures such as Bonferroni, Sidak, Tukey, and selected control-group comparisons in relevant analysis menus. For general collections of p-values, researchers may need syntax, another validated tool, or a carefully documented manual procedure. Save the unadjusted results, adjustment settings, family size, and software version. The SPSS psychology dissertation workflow can help organise reproducible syntax and output.

R and other reproducible workflows

R’s p.adjust() can apply several standard procedures to a declared vector of p-values. Packages for estimated marginal means and general linear hypotheses can produce adjusted contrasts and simultaneous intervals. Whatever software is used, verify whether the named option controls FWER or FDR, what dependence assumptions apply, and which comparisons the software treated as one family.

How to report Psychology dissertation multiple comparisons

Readers should be able to reconstruct what was planned, what was tested, and how the adjustment affected interpretation. Report multiplicity decisions in the method and the complete results in the results section. Discuss unexpected patterns cautiously in the discussion.

Method section

  • Identify primary, secondary, and exploratory hypotheses.
  • Define each test family and explain the rationale.
  • Name the error rate and adjustment procedure.
  • State the significance or discovery-rate level.
  • Describe planned contrasts, tails, outcomes, occasions, and subgroup tests.
  • Record software, packages, versions, and relevant options.

Results section

  • Report every prespecified test, not only favourable results.
  • Give effect estimates and confidence intervals.
  • Provide raw and adjusted p-values where this aids transparency.
  • Label post-hoc and exploratory analyses.
  • Explain deviations from the analysis plan.
  • Use tables that make the family structure visible.

Discussion section

Interpret the pattern rather than sorting results into a simplistic significant or non-significant list. A precise estimate near zero may argue against a meaningful effect. A promising but imprecise estimate may justify further research. An exploratory adjusted finding can generate a hypothesis without being presented as definitive confirmation.

Use the guide to the psychology dissertation discussion section to connect statistical uncertainty with theory, limitations, and practical meaning.

Common mistakes to avoid

Adjusting only the tests that were significant

The family must be defined before filtering by results. Applying a correction only to p-values already below .05 omits the unsuccessful opportunities that created the multiplicity problem.

Using Bonferroni automatically

Bonferroni is valid in many settings, but validity does not mean optimality. Holm may offer greater power for a general FWER-controlled family. Specialised procedures may better match all-pairs or control-group contrasts. FDR may suit a larger discovery-focused family.

Creating tiny families after seeing the data

Dividing one broad family into several convenient groups can weaken error control. Separate families require substantive, design-based justification rather than a post-hoc attempt to preserve small p-values.

Ignoring interactions and model searches

Multiplicity includes more than pairwise comparisons. Factorial main effects, interactions, outcomes, subgroups, covariate sets, exclusion rules, and alternative scoring choices can all contribute to the opportunity for selective inference.

Equating adjustment with research quality

A corrected p-value cannot repair biased sampling, poor measurement, confounding, model misspecification, low power, or selective outcome definition. Multiple-comparison control is one part of a defensible design, not a certificate of validity.

Frequently asked questions

Do I always need to correct for Psychology dissertation multiple comparisons?

No single rule fits every analysis. A genuinely single primary test does not create a multiple-test family. Several tests that support one collective claim usually require a multiplicity strategy, which may include prespecification, hierarchical testing, simultaneous modelling, adjustment, or explicitly exploratory interpretation.

Is Holm better than Bonferroni?

For a general family where strong FWER control is required, Holm is often preferable because it controls FWER while being no less powerful than simple Bonferroni. A specialised contrast procedure may be more suitable when the comparisons have a particular structure.

Should I adjust p-values for several questionnaire subscales?

Often yes when favourable evidence on any subscale would support the same overall claim. First confirm that the subscales are theoretically and psychometrically defensible. Then define whether they form one family, separate prespecified questions, or exploratory outcomes.

Does a significant omnibus ANOVA remove the need for adjustment?

No. The omnibus test and follow-up comparisons answer different questions. Pairwise contrasts still require an appropriate simultaneous procedure when they are interpreted as a family. Factorial ANOVA can also contain multiplicity across main effects and interactions.

Should I report raw or adjusted p-values?

Reporting both can improve transparency when the table clearly identifies the method and family. At minimum, report adjusted values or simultaneous intervals used for inference, along with effect estimates. Never present raw values as if they were the basis of an adjusted decision.

Can I use Benjamini-Hochberg with correlated outcomes?

Dependence matters. The standard procedure has established control under independence and certain positive-dependence conditions, while other variants address broader dependence more conservatively. Consult a statistician when the dependence structure is complex, and state the software method and assumptions.

Conclusion

Psychology dissertation multiple comparisons should be handled by design, not patched after a long search for small p-values. Define each inferential family from the research claims, distinguish confirmatory from exploratory work, choose FWER or FDR control deliberately, and use a method that matches the comparison structure. Then report all planned tests, effect estimates, intervals, raw and adjusted evidence, deviations, and limitations.

If you need support aligning hypotheses, analysis families, software output, and APA-style reporting, seek ethical dissertation guidance that helps you understand and defend your own decisions. A useful review should strengthen your reasoning and transparency while leaving authorship, interpretation, and academic responsibility with you.

Leave a Reply

Your email address will not be published. Required fields are marked *