Psychology dissertation effect size reporting shows how large an observed difference, relationship, or model contribution is, not merely whether a statistical test crossed a threshold. A useful effect-size analysis identifies the estimand, chooses a measure that matches the design, reports uncertainty, and interprets magnitude in the study’s psychological context.
This guide explains raw and standardised effects, Cohen’s d and Hedges’ g, correlations, risk measures, eta-squared statistics, regression effects, confidence intervals, repeated-measures choices, practical importance, and APA-style reporting. It is designed for quantitative psychology dissertations and supports, rather than replaces, your approved analysis plan and supervisor’s methodological guidance.
What is an effect size?
An effect size is a numerical description of the magnitude and often the direction of a phenomenon. It may describe a mean difference in scale points, a standardised difference, the strength of an association, a proportion of variance, an odds ratio, or a model parameter. The most useful measure directly answers the research question.
For example, a wellbeing intervention could produce a three-point mean difference on a validated scale. That raw difference preserves the instrument’s units. A standardised mean difference expresses the same contrast relative to a chosen standard deviation. Neither representation is automatically better. Raw effects are often easier to interpret, while standardised effects can support comparison across measures and studies.
The American Statistical Association states that a p value does not measure effect magnitude or importance. A small p value can accompany a trivial effect in a large sample, while an important but imprecisely estimated effect can fail to reach conventional significance in a small sample. Effect estimates and confidence intervals therefore answer questions that a p value cannot.
Why psychology dissertations should report effect sizes
APA’s Journal Article Reporting Standards for quantitative research call for effect sizes and confidence intervals or other statistical information for the main findings. The APA Task Force on Statistical Inference also emphasised that effect sizes support future power analyses and meta-analyses. Reporting them helps readers judge magnitude, precision, theoretical relevance, and compatibility with earlier research.
Effect sizes also make your dissertation more cumulative. A future researcher may use your estimate to plan a replication, combine evidence, or compare settings. Lakens explains that this cumulative purpose requires precise labels and enough design information to reconstruct the calculation, especially for within-participant studies where several versions of a standardised mean difference exist.
| Question | Useful quantity | Why it helps |
|---|---|---|
| How many scale points separate groups? | Raw mean difference | Retains clinically or practically meaningful units |
| How large is a difference relative to variability? | Cohen’s d or Hedges’ g | Provides a standardised comparison |
| How strongly do two variables vary together? | Pearson’s r or rank correlation | Describes direction and association strength |
| How different are binary outcome probabilities? | Risk difference, risk ratio, odds ratio | Represents absolute or relative contrast |
| How much variance is associated with an ANOVA effect? | Eta squared, partial eta squared, or omega squared | Summarises variance attribution under stated model |
| How much does an outcome change with a predictor? | Unstandardised regression coefficient | Preserves predictor and outcome units |
Define the estimand before choosing a statistic
The estimand is the population quantity your analysis aims to estimate. Before opening software, write it in words. Are you estimating the average score difference between assigned conditions, the within-person change after an intervention, the association between two constructs, or the unique contribution of a predictor after adjustment?
That statement determines the effect-size family. It also prevents a common mistake: selecting whichever statistic appears automatically in software. Your psychology dissertation hypotheses should identify the comparison or association clearly enough to guide this choice.
Raw effects and standardised effects
A raw effect remains in meaningful units. A six-minute sleep difference, a four-point anxiety-score difference, or a 0.8-second response-time change can be evaluated against theory, measurement error, prior interventions, costs, or stakeholder priorities. Always report raw descriptive statistics when possible, even if a standardised effect is also needed.
A standardised effect divides or otherwise scales the raw effect so that units no longer dominate. This can help when instruments differ, but the number depends on the standardiser and population variability. The same raw difference can yield different standardised effects in homogeneous and heterogeneous samples. Standardisation does not create universal meaning.
Effect sizes for mean differences
Independent groups
For two independent groups, Cohen’s d commonly divides the mean difference by a pooled within-group standard deviation. The sign reflects the order of subtraction, so define which group comes first. Hedges’ g applies a small-sample bias correction and is often preferable when groups are modest in size or when estimates will inform evidence synthesis.
If group variances differ materially, a pooled standard deviation may not represent either group well. A control-group standard deviation can be appropriate when the control provides the meaningful reference, but label that choice. Report group means, standard deviations, sample sizes, the effect-size formula or software, and a confidence interval.
Paired and repeated measurements
Within-participant designs need special care because the pre-post correlation affects precision and standardisation. One version of d divides the mean change by the standard deviation of change scores. Another uses a standard deviation that does not incorporate the within-person correlation. These answer related but not identical comparison needs.
Lakens recommends using informative subscripts and stating the standardiser. If you want to compare a repeated-measures effect with an independent-groups literature, choose a compatible effect measure. If you want to describe within-person change, a change-score standardiser may be more direct. Never write only “Cohen’s d” for a paired design without explaining the calculation.
More than two groups
An omnibus ANOVA effect indicates whether a factor accounts for variability across several means, but it does not reveal which contrasts matter. Report the omnibus effect and planned contrast effects. Pairwise standardised differences may be useful for focused comparisons, with multiplicity handled consistently with the analysis plan.
Effect sizes for correlation and regression
Pearson’s r is both a test statistic and an effect-size measure for a linear association. It ranges from minus one to one, with the sign indicating direction. Its square is sometimes described as shared variance, but this interpretation must be used carefully because it does not establish causation or unique explanatory contribution.
Use a rank correlation when the estimand concerns monotonic association or when the planned robust approach requires it. Do not switch from Pearson to Spearman solely because one yields a more favourable result. Plot the data and inspect influential observations using the outlier-analysis guide.
In regression, the unstandardised coefficient is often the clearest effect: it estimates the expected outcome change for a one-unit predictor change, holding specified covariates constant. A standardised beta can help compare predictors measured on different scales, but it depends on sample variability and does not automatically establish relative importance.
Report coefficients with intervals and explain coding. A coefficient for a binary predictor changes sign and meaning when the reference group changes. Interaction effects require conditional estimates or simple slopes; the main coefficient alone does not summarise moderation. See the regression and moderation guides for model-specific reporting.
Eta squared, partial eta squared, and omega squared
ANOVA software often reports partial eta squared. Eta squared describes the proportion of total observed variance attributable to an effect in the fitted design. Partial eta squared describes an effect relative to that effect plus its error term, excluding other modeled effects from the denominator. These values are not interchangeable.
Partial eta squared can appear larger than eta squared and may be difficult to compare across designs with different factors. Generalised eta squared was developed to improve comparability across some repeated-measures and between-participant designs. Omega squared adjusts for upward bias and may better estimate a population variance proportion, although variants and assumptions still require clear labels.
Choose the measure before inspecting results, follow disciplinary and supervisory expectations, and report its exact symbol. Lakens’ practical primer provides formulas for t tests and ANOVAs and explains why researchers should state which effect-size variant they calculated.
| Design or model | Primary effect options | Essential clarification |
|---|---|---|
| Independent two-group comparison | Raw difference, d, Hedges’ g | Group order and standard deviation used |
| Paired comparison | Mean change, repeated-measures d | Change-score or alternative standardiser |
| Correlation | r, rank correlation | Linear or monotonic estimand |
| Factorial ANOVA | Partial eta squared, omega squared, contrast effects | Denominator and factor structure |
| Linear regression | b, standardised beta, incremental R squared | Coding, units, adjustment set |
| Binary outcome | Risk difference, risk ratio, odds ratio | Reference event, baseline risk, model scale |
| Mediation | Indirect effect with interval | Scale, model, and inferential method |
Confidence intervals are part of the effect
A point estimate alone hides uncertainty. A confidence interval shows the range of parameter values compatible with the data and model under the stated procedure. Wide intervals indicate that meaningfully different effects remain plausible. Narrow intervals provide greater resolution, though neither width nor coverage repairs bias, poor measurement, or misspecification.
Do not interpret a 95% confidence interval as a 95% probability that the fixed parameter lies inside the observed interval. Instead, describe the values the data are compatible with and their practical implications. Consider whether the interval includes effects that are harmful, negligible, beneficial, or large enough to change theory or practice.
For standardised effects, use software that provides appropriate interval calculations. Simple normal approximations may perform poorly in small samples or for bounded measures. Record the software version, package, procedure, and any bootstrap settings. If bootstrapping is used, state the resampling unit so clustered or repeated observations are not treated as independent.
Do not treat generic benchmarks as universal truths
Labels such as small, medium, and large can be useful teaching shorthand, but they are not universal psychological laws. Funder and Ozer argue that even effects described as small can be consequential when they accumulate across people, occasions, or systems. Conversely, a numerically large effect may have little practical value if the outcome is unreliable, rare, or costly to change.
Interpret magnitude using the measurement scale, previous estimates, plausible mechanisms, intervention costs, duration, base rates, harms, and affected population. A two-point improvement may matter greatly on one clinical scale and be negligible on another. A correlation may be theoretically revealing without supporting individual prediction.
Use a smallest effect size of interest where possible
A smallest effect size of interest, or SESOI, defines the minimum magnitude that would matter for your specific question. It can be justified from theory, measurement resolution, prior high-quality evidence, stakeholder judgement, practical costs, or resource constraints. Specify it before results if it will guide decisions.
Lakens and colleagues show how equivalence tests can evaluate whether effects at least as extreme as prespecified bounds can be rejected. A nonsignificant conventional test does not demonstrate absence. An equivalence analysis can provide information about whether the data rule out effects large enough to matter, provided the bounds are justified and the study is sufficiently precise.
Effect size and power analysis are connected but different
Power analysis uses a target effect size during planning. The observed effect size is an uncertain estimate after data collection. Do not use the observed effect to calculate “post hoc power”; for standard tests, it largely restates the p value and adds little information. Report the estimate and interval instead.
For planning, avoid assuming that a previous study’s point estimate is the true effect. Published estimates may be imprecise or selected for significance. Use a justified SESOI, a conservative synthesis, or a range of plausible values. The power-analysis guide explains these choices in detail.
A step-by-step dissertation workflow
Step 1: State the research contrast
Write the exact population comparison or association. Identify outcome, predictor, groups, time points, adjustment variables, and direction. Clarify whether the primary effect is raw, standardised, absolute, or relative.
Step 2: Choose the primary effect before testing
Select a measure that maps to the estimand and design. Define the numerator, denominator, reference group, sign, and any standardiser. Add secondary representations only when they answer a useful additional question.
Step 3: Verify data and model assumptions
Check coding, scale ranges, group labels, missingness, distribution shape, dependence, and model fit. Effect-size calculations inherit the assumptions and errors of the underlying analysis. Use the missing-data guide when incomplete observations affect denominators or analytic samples.
Step 4: Calculate estimates and intervals reproducibly
Use syntax or scripted analysis, retain output, and verify that the software’s definition matches your plan. Cross-check one calculation from descriptive statistics or a second trusted method when feasible. This catches reversed group order, incorrect pooled variance, and confusion among eta-squared variants.
Step 5: Interpret against context
Compare the point estimate and interval with the SESOI, prior credible evidence, measurement units, and real consequences. Separate statistical compatibility from practical judgement. Avoid claiming clinical importance unless the design, outcome, and evidence support that claim.
Step 6: Report enough to reproduce the effect
Give descriptive statistics, sample sizes, effect label, formula or standardiser where ambiguous, confidence level, interval, analysis population, and software. State exploratory status and multiplicity decisions. If results change under defensible alternatives, report the sensitivity analysis rather than selecting one version silently.
Psychology-specific examples
Independent-groups intervention
A student compares a brief sleep-hygiene intervention with an information-only control. The primary estimand is the adjusted mean difference in sleep-quality score after treatment. The student reports adjusted means, the raw difference with a 95% confidence interval, and Hedges’ g using the pooled residual standard deviation. The raw difference is compared with a prespecified two-point threshold considered worthwhile.
The p value is presented as one part of the model output. The discussion focuses on whether the interval contains trivial and worthwhile effects, the short follow-up, adherence, and scale reliability.
Within-participant attention experiment
A student measures response time under neutral and threat conditions in the same participants. The primary effect is the mean within-person difference in milliseconds. A repeated-measures standardised difference is also reported, with its standardiser named. Trial exclusions are prespecified and a robust sensitivity analysis checks whether unusually slow valid trials drive the estimate.
This reporting makes the practically interpretable raw change visible and prevents readers from assuming the reported d is equivalent to an independent-groups d.
Common effect-size mistakes
- Reporting only p values and calling significance the size of an effect.
- Writing “Cohen’s d” without specifying group order or the standard deviation used.
- Using independent-groups formulas for paired data.
- Confusing eta squared with partial eta squared.
- Calling a standardised beta the unique importance of a predictor.
- Applying generic benchmarks without theory or practical context.
- Reporting a point estimate without a confidence interval.
- Using the observed effect for post hoc power.
- Converting among effect sizes without checking design assumptions.
- Interpreting an odds ratio without baseline risk or reference event.
Effect-size reporting checklist
| Check | What to record |
|---|---|
| Estimand | Population contrast or association in words |
| Measure | Exact statistic and variant |
| Direction | Group order, reference category, or sign meaning |
| Scale | Raw units and standardiser where relevant |
| Precision | Confidence level, limits, and calculation method |
| Context | SESOI, prior evidence, practical or theoretical meaning |
| Reproducibility | Descriptives, analytic sample, software, syntax, sensitivity checks |
Frequently asked questions
Which effect size should I use for a t test?
Report the raw mean difference and interval. For independent groups, Cohen’s d or bias-corrected Hedges’ g may add a standardised comparison. For paired data, choose a repeated-measures version and name its standardiser.
Is partial eta squared the same as eta squared?
No. Their denominators differ. Partial eta squared removes variance associated with other modeled effects from its denominator, while eta squared relates the effect to total observed variance. Label the chosen measure exactly.
Does a large effect size prove practical importance?
No. Practical importance depends on outcome meaning, uncertainty, duration, costs, harms, base rates, population, and design quality. Interpret the estimate and interval in context.
Can I compare Cohen’s d across studies?
Only cautiously. Check whether studies used compatible designs, outcomes, populations, standardisers, and time points. A standardised metric does not erase substantive differences.
What if my confidence interval is very wide?
Report it honestly. A wide interval means the study cannot distinguish among substantially different magnitudes. Discuss limited precision and avoid treating the point estimate as settled.
Conclusion
Strong psychology dissertation effect size reporting starts with the research question, not a software menu. Define the estimand, preserve raw units where useful, select a design-compatible standardised measure, and report a confidence interval. Name every variant and reference category clearly.
Interpret magnitude against theory, measurement, prior evidence, and practical consequences rather than generic labels alone. If you need ethical academic support, seek guidance that teaches you to choose, calculate, and explain the statistic for your own approved project while keeping the analysis decisions and authorship entirely yours.
Authoritative references
- American Psychological Association: JARS-Quant reporting standards
- Wilkinson and the APA Task Force: Statistical methods in psychology journals
- Lakens: Calculating and reporting effect sizes
- Fritz, Morris, and Richler: Effect-size estimates
- Pek and Flora: Reporting effect sizes in psychological research
- American Statistical Association statement on p values
- Lakens and colleagues: Equivalence tests for psychological research
