Confidence intervals compared with equivalence bounds for a psychology dissertation

Psychology dissertation equivalence testing evaluates whether an effect is small enough to be considered practically equivalent to zero or another reference value. It addresses a question that an ordinary non-significant test cannot answer: are effects large enough to matter statistically incompatible with the data?

The method is useful for replication studies, validation work, intervention comparisons, manipulation checks, and theory tests that predict negligible differences. Its credibility depends on justified equivalence bounds, suitable power, accurate effect-size units, and conclusions that distinguish equivalence from mere uncertainty.

What is psychology dissertation equivalence testing?

Traditional null-hypothesis significance testing usually asks whether the data are inconsistent with an exact null effect. When the p-value is above the chosen alpha level, the researcher fails to reject that null. This outcome does not demonstrate that the effect is absent, trivial, or practically unimportant. It may simply reflect imprecise data.

Equivalence testing reverses the practical question. The researcher defines lower and upper bounds representing effects considered meaningfully different from the reference. The test asks whether effects at or beyond those bounds can be rejected. When both sides are rejected, the result supports statistical equivalence within the prespecified region.

Lakens’ open practical primer explains the two one-sided tests procedure for t tests, correlations, and meta-analyses. The accessible equivalence-tests primer in PubMed Central emphasises that a non-significant conventional test does not by itself support absence of a meaningful effect.

Understand the two one-sided tests procedure

The two one-sided tests procedure, commonly called TOST, specifies a lower equivalence bound and an upper equivalence bound. One test evaluates whether the population effect is greater than the lower bound. The other evaluates whether it is smaller than the upper bound. Equivalence is concluded only when both tests reject their respective null hypotheses.

For a symmetric region of minus 0.30 to 0.30 in standardised mean-difference units, the hypotheses are not simply “effect equals zero” versus “effect does not equal zero.” Instead, the equivalence null includes effects at or below minus 0.30 and effects at or above 0.30. The alternative places the effect inside the interval.

At alpha .05, TOST can be understood through a 90% confidence interval for the effect. If the entire interval lies within the equivalence bounds, both one-sided tests are significant. This 90% interval serves the equivalence decision; a 95% interval can still be reported for ordinary estimation and null testing.

Conventional test Equivalence test Combined interpretation
Significant Not equivalent Effect differs from zero and meaningful effects remain plausible
Significant Equivalent Effect differs from zero but is contained within the defined trivial region
Not significant Equivalent Data support an effect smaller than the equivalence bounds
Not significant Not equivalent Result is inconclusive because meaningful and null effects remain plausible

The second row is possible with a precise estimate that differs slightly from zero but remains practically negligible. This demonstrates why statistical significance and practical importance are separate. The fourth row is common in small studies and must not be labelled evidence of no effect.

Define the smallest effect size of interest

The smallest effect size of interest, or SESOI, is the smallest effect considered theoretically or practically meaningful for the research question. Equivalence bounds are commonly placed at the negative and positive SESOI when effects in either direction matter equally. Asymmetric bounds are also legitimate when costs, benefits, or theory differ by direction.

Choose the SESOI before inspecting the focal result. The tutorial by Lakens, Scheel, and Isager explains that bounds should answer a meaningful question and may be based on theoretical predictions, just-noticeable differences, or cost-benefit considerations. The open university record for the peer-reviewed equivalence-testing tutorial gives detailed guidance on this decision.

Defensible sources for an equivalence bound

  • A quantitative theoretical prediction below which a mechanism has no substantive consequence.
  • A validated minimally important difference on the raw outcome scale.
  • A perceptual or behavioural threshold such as a just-noticeable difference.
  • A cost-benefit analysis incorporating intervention burden, risk, and expected benefit.
  • A clearly justified effect that a well-powered prior study could reliably detect.
  • A discipline-informed consensus developed independently of the observed estimate.

A conventional label such as “small effect” is not a complete justification. Cohen’s benchmarks may lack context for the population, outcome, and decision. Explain why effects inside the bounds would not change the theoretical or practical conclusion.

Raw-unit bounds are often easier to defend. Five points on a validated anxiety scale, ten minutes of sleep, or a specified accuracy difference may have interpretable consequences. Standardised bounds are useful across measures but depend on the standard deviation used. State whether the standardiser is pooled, control-group, baseline, or another quantity.

Match bounds to the estimand and design

The bounds must use the same estimand and scale as the reported effect. A bound for a standardised mean difference cannot be entered as a raw mean difference. A correlation bound cannot be applied directly to a regression slope. Transformations can be used only when their assumptions and reference quantities are clear.

For paired data, define whether the standardised effect uses the standard deviation of change scores or another convention. For correlations, set bounds in correlation units and consider measurement reliability and range restriction. For regression, clarify whether the target is an unstandardised slope, standardised coefficient, incremental variance, or model prediction.

Use the site’s psychology dissertation effect-size guide to select and interpret the metric. Changing metrics after viewing results can change the equivalence conclusion and should not be treated as a harmless formatting choice.

Plan sample size for equivalence

Equivalence requires precision. Narrower bounds demand more information because the confidence interval must fit inside a smaller region. Sample size therefore depends on the bounds, alpha, desired power, design, variance, allocation, and expected true effect.

Power the study for its primary inferential goal. If the goal is to demonstrate equivalence when the true effect is zero, calculate the probability that both one-sided tests will reject at the chosen bounds. If the study also aims to detect meaningful effects, report power for that goal separately. One calculation does not automatically satisfy both.

Use realistic attrition, missingness, clustering, reliability, and unequal group sizes. Simulation may be necessary for multilevel, repeated-measures, or adjusted models. Document assumptions in the psychology dissertation power-analysis guide.

A result cannot be rescued by a post hoc claim that the observed confidence interval looks narrow. Prespecify the bounds and planning assumptions. When an existing dataset is too imprecise, report an inconclusive result rather than widening the bounds until equivalence appears.

Choose the correct analysis

Research design Possible equivalence target Important detail
Independent groups Raw or standardised mean difference Check variance assumptions and group allocation
Paired or repeated observations Mean change or paired standardised effect Use the within-pair dependence correctly
Correlation study Population correlation Bounds must be in correlation units
Regression model Specified coefficient or incremental contribution Covariate selection defines the estimand
Replication Effect smaller than a theoretically meaningful bound Do not derive convenient bounds from the replication result
Scale validation Difference, bias, or association within a justified tolerance Equivalence is not the same as reliability or agreement

The ordinary t-test article on this site explains assumptions and effect estimates for common mean comparisons. Consult the psychology dissertation t-test guide before adding TOST. Equivalence testing does not remove the need to check independence, distributional conditions, outliers, or the suitability of the estimator.

Run and document TOST correctly

The workflow begins with the estimand, bounds, alpha, and planned sample size. Clean and score data using prespecified rules. Estimate the effect and uncertainty, run both one-sided tests, and retain numerical outputs. Report software, version, function, options, and effect-size convention.

The official CRAN record for the TOSTER package documents procedures for t tests, correlations, differences between proportions, meta-analyses, and related power calculations. Software availability does not determine whether the chosen bounds are defensible.

A complete result reports:

  • the lower and upper bounds, their units, and their justification;
  • the estimated effect and confidence interval;
  • both one-sided test statistics, degrees of freedom where applicable, and p-values;
  • the conventional null test if it answers a relevant secondary question;
  • the alpha level, analysis software, and exact effect-size definition;
  • assumption checks, exclusions, missing-data handling, and sensitivity analyses.

Interpret confidence intervals graphically

A simple figure can show the point estimate, confidence interval, zero line, and equivalence bounds. This makes the four possible outcomes visible. Keep the scale consistent across related analyses and label whether the interval is 90% for TOST or 95% for ordinary estimation.

Do not truncate the axis in a way that exaggerates a negligible estimate. If multiple outcomes are tested, show each bound and adjust the inferential plan for multiplicity when necessary. Different outcomes may require different raw-unit bounds because their practical meanings differ.

Interval position Conclusion What not to claim
Entirely inside bounds Statistically equivalent within the defined region The effect is exactly zero
Overlaps a bound and zero Inconclusive There is no effect
Excludes zero but crosses a bound Different from zero, not equivalent The effect must be practically large
Inside bounds but excludes zero Statistically detectable yet practically equivalent Significance proves importance

Worked psychology example

Imagine a dissertation comparing a five-minute and a fifteen-minute mindfulness exercise on immediate anxiety. The shorter exercise would be preferable if any disadvantage is smaller than three points on a validated anxiety scale. Prior evidence and stakeholder consultation justify bounds of minus three and plus three points for a two-sided equivalence question.

The student preregisters the outcome, bounds, exclusions, alpha, allocation, and power calculation. After data collection, the adjusted mean difference favouring the longer exercise is 0.8 points. The 90% confidence interval ranges from minus 1.1 to 2.7 points. Because the entire interval lies within the bounds, both one-sided tests support equivalence.

The conclusion is not that the interventions are identical. It is that differences of three points or more in either direction are incompatible with the data at the specified error rate and under the model. Smaller differences remain possible. The interpretation also depends on the validity of the three-point threshold, outcome measurement, randomisation, missingness, and model assumptions.

Distinguish equivalence from noninferiority

Equivalence testing usually asks whether a difference lies within both a lower and upper tolerance. Noninferiority testing asks whether a new intervention is not worse than a comparator by more than a prespecified margin. It is normally one-sided and often applies when the new option offers another advantage such as lower cost, reduced burden, or fewer adverse effects.

Do not switch from a failed superiority test to a post hoc noninferiority claim. The margin, direction, analysis population, missing-data strategy, and design protections must be planned. A symmetric TOST is not automatically appropriate for a directional noninferiority question.

Use equivalence tests in replication research

A replication can combine a conventional test with an equivalence test. A non-significant replication that is also equivalent under a justified SESOI supports the conclusion that effects of meaningful size are unlikely. A non-significant but non-equivalent replication remains inconclusive.

Do not automatically set bounds from the original study’s observed effect. Published estimates can be inflated, and an original effect is not necessarily the smallest meaningful value. Derive the SESOI from theory, practical consequences, or a clearly defended design criterion. The psychology dissertation replication-study guide explains how to separate direct replication, extension, and sensitivity goals.

Handle multiple outcomes and exploratory bounds

Several outcomes create several equivalence claims. Prespecify a primary outcome and its bounds. If all outcomes must be equivalent for a decision, state that joint rule. If any one outcome can establish success, address the resulting multiplicity.

Exploratory plots across several bounds can be informative when no consensus SESOI exists. Label them as sensitivity analyses rather than selecting the narrowest or widest bound that produces the desired conclusion. Show how inference changes and explain which values stakeholders might reasonably defend.

Check robustness without changing the question

Equivalence conclusions can depend on exclusions, variance assumptions, missing-data handling, covariate adjustment, and the effect-size denominator. Plan a small set of sensitivity analyses that preserve the same estimand. Show whether the confidence interval remains within the original bounds rather than inventing new bounds for each result.

When an independent-groups analysis assumes equal variances, compare a justified unequal-variance approach if diagnostics or design make equality doubtful. For paired measurements, examine verified influential change scores and measurement floor or ceiling effects. Robust or bootstrap methods may be useful, but their compatibility with the intended equivalence test and error rate must be established.

Missing data require special care because complete-case results may describe a selected group. State the missingness assumptions, use an analysis suitable for the design, and examine plausible alternatives. Do not declare equivalence simply because removing incomplete or inconvenient cases narrows the interval. A smaller analysed sample normally reduces precision, but selection can also shift the estimate.

Covariate adjustment can improve precision in a randomised study when prespecified, yet different adjustment sets may target different conditional quantities in observational research. Explain why adjusted and unadjusted effects are comparable before presenting both as robustness checks. If they answer different questions, report them separately.

Finally, distinguish sensitivity of the estimate from sensitivity of the practical threshold. It is legitimate to display conclusions across several stakeholder-supported SESOIs, provided no single exploratory value is promoted as confirmatory. This analysis reveals where the evidence becomes decisive and where substantive judgement remains necessary.

Common mistakes to avoid

  • Interpreting p greater than .05 as proof of no effect.
  • Choosing equivalence bounds after inspecting the estimate.
  • Using generic small-effect benchmarks without context.
  • Entering standardised bounds into an analysis of raw differences.
  • Reporting only the TOST decision without the effect and interval.
  • Widening bounds because the planned study is underpowered.
  • Confusing equivalence, agreement, reliability, noninferiority, and absence.
  • Calling a statistically equivalent effect exactly zero.

Frequently asked questions

Does a non-significant t test show equivalence?

No. It only indicates insufficient evidence against the exact null under that test. Equivalence requires prespecified bounds and enough precision to reject effects at or beyond them.

What confidence interval is used for TOST?

With two one-sided tests at alpha .05, equivalence corresponds to a 90% confidence interval lying fully within the bounds. A 95% interval may also be reported for estimation.

Must equivalence bounds be symmetric?

No. Asymmetric bounds are valid when positive and negative effects have different theoretical or practical consequences. Justify both limits independently.

Can I use Cohen’s d of 0.20 as the bound?

You can only if 0.20 is substantively justified for the specific outcome and decision. A generic benchmark alone does not explain why smaller effects are unimportant.

Can a result be significant and equivalent?

Yes. A precise, very small estimate can differ statistically from zero while remaining entirely inside the prespecified equivalence region.

What if the result is neither significant nor equivalent?

The study is inconclusive for both questions. Meaningful effects and effects near zero remain compatible with the data. Report the uncertainty and avoid an absence claim.

Should equivalence testing be preregistered?

Yes whenever it supports a confirmatory claim. Preregister the estimand, bounds and rationale, alpha, analysis, exclusions, missing-data method, sample size, and decision rule.

Conclusion

Psychology dissertation equivalence testing can support a carefully bounded absence claim that ordinary non-significance cannot. Define the smallest effect of interest before seeing results, align bounds with the estimand, plan for precision, run both one-sided tests, and report the estimate and intervals alongside the decision.

If you need dissertation support, choose ethical guidance that helps you understand and reproduce your own analysis. Retain responsibility for the bounds, code, and interpretation, follow institutional requirements, protect participant data, and acknowledge permitted assistance transparently.