Psychology postgraduate reviewing Q-Q plots, a histogram, and residual diagnostics on a laptop

Psychology dissertation normality testing is not a ritual in which every variable must pass one test before analysis can begin. It is a model-specific investigation of whether distributional features could distort your estimates, standard errors, confidence intervals, or conclusions. This guide shows you what to examine, which data to test, how to combine graphs with formal tests, and what to do when normality is doubtful.

The central lesson is simple: check the assumption your chosen model actually makes. In many t tests, analyses of variance, and regression models, the relevant assumption concerns model errors or residuals, not the marginal distribution of every predictor and outcome. A careful workflow therefore begins with the research design and model, not with a software menu.

What normality means in a psychology dissertation

A normal distribution is a continuous, symmetric probability distribution described by a mean and standard deviation. Its familiar bell shape is useful, but visual resemblance to a bell is not the final criterion. Statistical models use normality in different places and for different purposes.

In a simple linear model, an observed score is represented as a systematic component plus an error component. Normality usually refers to the conditional errors: after the model accounts for predictors, the remaining deviations are assumed to follow an approximately normal distribution at each relevant predictor pattern. This is why checking raw depression scores does not, by itself, test the normality assumption of a regression predicting depression from sleep and stress.

The distinction matters because a skewed outcome can produce reasonably well-behaved residuals after appropriate predictors are included. Conversely, a roughly symmetric outcome can conceal non-normal residuals, heteroscedasticity, nonlinearity, clustering, or influential cases. The NIST normal probability plot guidance explains that the most frequent model-based application is a probability plot of fitted residuals.

Normality is not the same as measurement quality

A normal-looking variable is not automatically valid, reliable, unbiased, or measured at an appropriate level. A symptom scale may be skewed because most community participants report few symptoms. That shape could be substantively expected rather than evidence of an error. Likewise, a perfectly symmetric set of scores can still contain miscoding, careless responses, or a poor operational definition.

Normality checks should sit inside a wider psychology dissertation data analysis plan that addresses design, data quality, missingness, influential observations, model specification, and transparent reporting.

Why normality can matter, and why it is often overstated

Normality supports exact small-sample reference distributions for some familiar tests. When model errors are severely non-normal, especially in small or unbalanced samples, standard errors, confidence intervals, and p values may not have their advertised properties. Heavy tails and influential observations can be more consequential than mild, symmetric departures.

Yet many estimators do not require raw variables to be normal, and some procedures are reasonably robust to moderate departures under suitable conditions. Robustness depends on sample size, group balance, tail behaviour, heteroscedasticity, dependence, and the statistic being estimated. It should not be reduced to a universal sample-size threshold.

Research on practice in clinical psychology found that authors sometimes treated normality of variables, rather than errors, as a regression requirement. The abstract of Ernst and Albers’ study of regression assumptions highlights this misconception. A separate methodological discussion of linear regression and normality explains that non-normal errors primarily affect inference rather than the ordinary least-squares point estimates themselves, while the practical impact often diminishes in larger samples.

Analysis What to examine What not to demand automatically
Independent-samples t test Outcome distribution within each group; unusual cases; variance pattern A pooled outcome that is perfectly normal
One-way or factorial ANOVA Model residuals, group balance, variances, influential observations Every cell passing a formal test
Linear regression Residual distribution conditional on predictors, functional form, variance, influence Normally distributed predictors
Repeated-measures model Model-specific residuals and covariance assumptions Normality of each raw occasion considered alone
Correlation Scatterplot, linearity, unusual pairs, and joint distribution for Pearson inference Two separate Shapiro-Wilk passes as a complete check
Generalised linear model Distribution and link specified by the model; suitable diagnostics Normally distributed binary or count outcomes

Begin with the design and planned model

Before generating a histogram, write down the outcome, predictors, grouping structure, repeated observations, and planned estimand. Then identify where the normality assumption enters the proposed analysis. Your software documentation, statistical text, supervisor, or methods adviser can help when the model is complex.

For an independent-samples t test, inspect the continuous outcome within each group because the test compares group means. For an ANOVA dissertation analysis, fit the intended model and examine residuals while also considering balance and variance equality. For psychology dissertation regression, check residuals alongside linearity, constant variance, independence, and influence.

If your outcome is inherently binary, ordinal, a count, a proportion with many boundary values, or time to an event, a normal-error model may be conceptually wrong. The answer is usually a model designed for that outcome type, not transforming the data until a normality test becomes non-significant.

Use graphical checks before formal tests

Graphs show the nature and location of a departure. A single p value cannot tell you whether the problem is skewness, heavy tails, several modes, rounding, a ceiling, a floor, or one influential observation. Begin with plots that preserve context and then use numerical summaries as supporting evidence.

Histograms and density displays

A histogram gives an accessible overview of shape, spread, gaps, and potential multimodality. Its appearance depends on bin width, so inspect more than one sensible binning if the sample is modest. A density curve can help with shape but may smooth away important features. Always retain awareness of actual values and the scale’s possible range.

For example, reaction times are often positively skewed. The graph may show a long right tail caused by genuine slow responses, lapses, technical delays, or coding mistakes. Those explanations lead to different decisions. Investigate observations rather than deleting them merely to improve the picture.

Q-Q plots

A normal Q-Q plot compares ordered observed values with theoretical normal quantiles. Points close to a straight reference line indicate approximate agreement. Systematic curvature suggests skewness, an S-shaped pattern can indicate tail differences, and isolated end points can flag unusual observations. NIST describes how normal probability plots and residual histograms complement each other.

Do not expect a perfect line. Sampling variation produces small departures even when data arise from a normal population. Interpret magnitude and pattern in relation to sample size, research design, and the inferential method.

Residual plots are part of the same diagnosis

A Q-Q plot is not enough if the model is misspecified. Plot residuals against fitted values and relevant predictors to examine nonlinearity and changing variance. Check observations over collection order when drift or time dependence is plausible. NIST cautions that a normal probability plot is hard to interpret if functional form, constant variance, or independence is already wrong.

This broader approach prevents a common mistake: diagnosing a curved residual pattern as non-normality when the real problem is an omitted nonlinear term or an unsuitable link function.

How to interpret formal normality tests

Formal tests evaluate a null hypothesis that the sampled distribution matches a specified normal distribution. A small p value provides evidence against that model. A large p value does not prove normality; it may reflect limited power, a small sample, or departures the test detects poorly.

The Shapiro-Wilk test is commonly used and is often a reasonable formal check. Kolmogorov-Smirnov procedures require care when distribution parameters are estimated from the same data, and software may apply a Lilliefors correction. Anderson-Darling gives greater weight to tails than some alternatives. The accessible review by Ghasemi and Zahediasl summarises common graphical and formal normality assessments.

Evidence Useful contribution Important limitation
Histogram Shows shape, gaps, modes, boundaries, and possible outliers Appearance depends on bins and sample size
Q-Q plot Reveals systematic skew, tail behaviour, and unusual endpoints Requires judgement; small deviations are expected
Shapiro-Wilk test Provides a reproducible test of a normal model Can miss departures in small samples and flag trivial ones in large samples
Skewness and kurtosis Summarise asymmetry and tail or peak characteristics Do not describe multimodality or local anomalies
Residual versus fitted plot Checks functional form and variance pattern Does not by itself establish normality
Sensitivity analysis Shows whether conclusions depend on modelling choices Must be planned and interpreted, not used to hunt for significance

Why sample size changes the interpretation

With a small sample, a formal test can return a large p value even when a Q-Q plot shows a meaningful tail problem. With a very large sample, it can identify minor deviations that have negligible effect on the analysis. The question is therefore not simply, “Is p above .05?” It is, “What departure is present, how severe is it, and is the planned inference sensitive to it?”

A 2012 practical guide notes that visual methods and formal tests should be considered together. More recent methodological work also argues that the risks of non-normality are often manageable and that complicated alternatives can introduce their own errors. The abstract of Knief and Forstmeier’s analysis supports a proportionate response rather than automatic model replacement.

A defensible psychology dissertation normality testing workflow

  1. Specify the analysis before checking. Link each research question and hypothesis to a planned model.
  2. Verify the data. Check ranges, labels, duplicate records, missing-value codes, and scoring before interpreting distributions.
  3. Describe the relevant units. Produce appropriate sample sizes, centres, spreads, ranges, and group-specific summaries using the descriptive statistics guide.
  4. Plot the data in context. Use histograms, boxplots, dot plots, or group-specific displays to understand observable patterns.
  5. Fit the planned model. Save raw, standardised, or studentised residuals as appropriate to the analysis.
  6. Inspect residual diagnostics. Use a Q-Q plot, residual histogram, residual-versus-fitted plot, and influence measures.
  7. Add a formal test if useful. Record the test, statistic, sample size, and p value, but do not let it overrule all other evidence.
  8. Assess consequences. Consider sample size, balance, tail severity, outliers, heteroscedasticity, and inferential target.
  9. Run a justified sensitivity analysis. Compare a robust, resampling, transformed, or alternative-distribution analysis when it answers the same substantive question.
  10. Report the decision trail. State what was checked, what was found, what action followed, and whether conclusions changed.

What to do when normality is doubtful

Do not treat a failed test as an instruction to press a transformation button. First check data integrity, plot the departure, identify the relevant observations, and ask whether another assumption or model specification is responsible.

Investigate data and influence

Confirm that extreme values are not impossible codes, unit errors, duplicate cases, or scale-scoring mistakes. Genuine observations should not be removed simply because they are inconvenient. Use the outlier analysis guide to distinguish unusual values from influential cases and to document sensitivity checks.

Consider transformation carefully

A logarithm, square root, or other transformation can sometimes align a model with multiplicative processes, stabilise variance, and reduce right skew. It also changes the scale and interpretation of coefficients. Transformations require a substantive rationale, rules for zero or negative values, and clear back-translated interpretation where appropriate.

Do not choose transformations solely by trying alternatives until the preferred hypothesis becomes significant. If the decision depends on observed data, disclose it and consider presenting results on both the original and transformed scales.

Use a model that matches the outcome

Binary, count, ordinal, bounded, or zero-inflated outcomes often call for a generalised linear, ordinal, beta, hurdle, or related model. This can preserve the outcome’s meaning better than forcing normality. Complex models should be chosen with suitable expertise and adequate sample size.

Consider robust or resampling methods

Heteroscedasticity-consistent standard errors, trimmed-mean procedures, robust regression, permutation tests, and bootstrap confidence intervals can reduce dependence on particular distributional assumptions. They are not interchangeable. Each method has conditions, and clustered, paired, or repeated data require resampling that respects the design.

A nonparametric rank test may be useful for some questions, but it does not automatically test the same estimand as a mean comparison. Explain what changed and why. The goal is a method aligned with the psychological question, not a method chosen merely because its p value is smaller.

Observed issue Questions to ask Possible response
One extreme residual Data error, rare valid case, or high influence? Verify, report influence, and compare analyses with a justified sensitivity check
Moderate right skew Expected from the construct? Does variance also increase with the mean? Retain if inference is robust, transform with rationale, or use a suitable outcome model
Heavy tails Are standard errors unstable? Are several cases influential? Use robust inference or an alternative error distribution
Ceiling or floor effect Is the measure bounded or poorly targeted? Acknowledge measurement limits and consider an appropriate bounded or ordinal model
Multimodality Are groups, conditions, or latent subpopulations being mixed? Model known structure; avoid claiming one homogeneous normal population
Nonlinear residual pattern Is normality really the problem? Reconsider functional form, interactions, link, or omitted structure

Psychology-specific examples

Reaction-time experiment

A student compares reaction times after neutral and emotional primes. Raw scores in both conditions are right-skewed. A paired analysis concerns within-person differences, so separate raw-score tests do not directly assess its relevant distributional condition. The student checks data-quality rules, plots participant differences, examines influential pairs, and fits the planned model.

The difference scores show mild skew but no isolated influential case. The paired t-test and a design-respecting permutation analysis lead to the same substantive conclusion. The dissertation reports both as a sensitivity check instead of claiming that a non-significant Shapiro-Wilk test proves normality.

Symptom outcome in a regression

Another student predicts symptom severity from perceived stress, sleep quality, and age. The outcome is moderately skewed, but predictors do not need to be normal. After fitting the model, the residual Q-Q plot shows heavy tails and the residual-versus-fitted plot shows increasing spread. The student investigates scoring, examines influence, and uses heteroscedasticity-consistent standard errors alongside the conventional analysis.

The coefficients have similar directions and magnitudes, but one borderline interval widens. Reporting this difference provides more information than deleting high-symptom participants or presenting a pass-fail normality claim.

How to report normality checks transparently

Report only diagnostics that informed the analysis, but give enough detail for a reader to understand the decision. Name the data or residuals examined, the graph and test used, the pattern observed, the action taken, and the sensitivity result. Avoid empty phrases such as “all assumptions were met” without evidence.

A concise quantitative report might state that standardised residuals were inspected using a Q-Q plot and residual histogram, that mild right-tail departure was observed without isolated influential cases, and that the Shapiro-Wilk statistic and exact p value were recorded. It can then explain that the prespecified model was retained because the sample was balanced and estimates were stable under robust standard errors. Report your own values and adapt the reasoning to your model.

Keep diagnostic output in an appendix when space is limited, and direct readers to it from the results section. Present the final estimates, uncertainty, and effect sizes using the psychology dissertation results guide. If you used SPSS, save syntax and decisions with the SPSS dissertation workflow.

Common normality-testing mistakes

  • Testing every column. This confuses variable distributions with model assumptions.
  • Using p > .05 as proof. Failure to reject is not evidence that a distribution is exactly normal.
  • Ignoring plots. A p value does not reveal the form or source of departure.
  • Checking only after results are known. Data-dependent switching can inflate flexibility and should be disclosed.
  • Deleting cases to pass. Exclusion needs a defensible quality or influence rationale, not cosmetic improvement.
  • Transforming without interpretation. A transformed model answers questions on a changed scale.
  • Ignoring other assumptions. Nonlinearity, unequal variance, dependence, and poor specification can matter more.
  • Calling nonparametric tests assumption-free. Every method has conditions and a specific estimand.

Frequently asked questions

Do all variables need to be normally distributed?

No. The relevant assumption depends on the model. Linear regression, for example, does not require normally distributed predictors. Normality is usually evaluated for conditional errors when exact parametric inference depends on that condition.

Should I use Shapiro-Wilk or Kolmogorov-Smirnov?

Shapiro-Wilk is commonly preferred for a general normality check, but neither test should be used alone. Confirm what your software implements, especially when parameters are estimated, and interpret the result with Q-Q plots, model diagnostics, sample size, and the planned inference.

What if Shapiro-Wilk is significant?

Inspect the plot and identify the departure. Check data quality and model specification, then judge whether the deviation threatens your target inference. A significant result can reflect a trivial departure in a large sample. Consider a justified robust or alternative model and report a sensitivity analysis.

Can I proceed if the data are non-normal?

Often, yes, but not automatically. The answer depends on which quantity is non-normal, severity, sample size, group balance, influential observations, heteroscedasticity, dependence, and the method’s robustness. Document that reasoning rather than relying on a universal rule.

Should normality tests go in the methods or results chapter?

Describe the planned diagnostic approach in methods. Report observed diagnostics, deviations, decisions, and sensitivity results in the results chapter. Detailed plots or output can go in an appendix if your institutional guidance permits.

Is a histogram enough?

No single diagnostic is enough. A histogram explains overall shape, while a Q-Q plot is more sensitive to systematic quantile departures. Residual-versus-fitted and influence plots assess different assumptions and should accompany a model-based check.

Conclusion

Good psychology dissertation normality testing identifies the assumption made by the planned model, examines the right observations or residuals, combines visual and formal evidence, and evaluates practical consequences. It does not force every variable through a pass-fail test.

Make the decision trail reproducible: verify the data, fit the planned model, inspect Q-Q and residual plots, document departures, justify any alternative, and show whether conclusions are stable. If you need ethical academic support, request feedback on your diagnostic reasoning, reporting, or analysis plan while retaining responsibility for your data and final work.

Authoritative references

Leave a Reply

Your email address will not be published. Required fields are marked *