Psychology postgraduate researcher classifying psychology dissertation confounding variables in a simple diagram

Psychology Dissertation Confounding Variables Guide can make an observed relationship look stronger, weaker, absent, or even reversed. A confounder is not simply another predictor that reaches statistical significance. It is a causal source of bias that creates a non-causal path between the exposure or focal predictor and the outcome.

This guide explains how to identify plausible confounders before analysis, distinguish them from mediators and colliders, select defensible control strategies, report adjusted results, and discuss residual confounding honestly. The examples use psychology questions that can be adapted to surveys, cohort studies, secondary data, and quasi-experimental dissertations.

What is a Psychology Dissertation Confounding Variables confounding variable?

Suppose a dissertation asks whether social media use affects depressive symptoms. Sleep quality may influence how much social media a person uses and may also affect depressive symptoms. If sleep quality precedes both variables and is not caused by the focal exposure, it may open a non-causal route between social media use and depression. The crude association can then mix the effect of interest with differences in sleep quality.

In causal diagrams, a confounder is commonly represented as a shared cause of the exposure and outcome. Hernán and Robins explain the broader causal-inference framework in their freely available book, Causal Inference: What If. The book is relevant across psychology, epidemiology, statistics, economics, and other disciplines.

A variable does not become a confounder merely because it correlates with both the exposure and outcome in one sample. Confounding depends on the target causal question, temporal order, and assumed data-generating process. The same variable can have a different role when the exposure, outcome, or estimand changes.

Confounding is tied to a specific effect

Start by naming the exposure, outcome, population, and target effect. “Factors associated with anxiety” is too vague for rigorous confounder selection. A clearer question is: “What is the total effect of weekly workplace incivility on anxiety symptoms after three months among early-career nurses?” This identifies the focal exposure, later outcome, timeframe, and intended total effect.

If the dissertation is predictive rather than causal, the aim changes. Prediction models seek accurate forecasts and may include variables because they improve out-of-sample performance. A causal study selects covariates to block biasing paths without blocking part of the effect being estimated. Do not use causal language for a purely predictive model.

Confounder, mediator, collider, or effect modifier?

Many dissertation errors arise from treating every background variable as a confounder. Causal role matters more than the label used in a dataset.

Variable role Causal position Typical implication Psychology example
Confounder Common cause of exposure and outcome Adjust when needed to close a biasing path Prior stress affects both help-seeking and later wellbeing
Mediator Caused by exposure and then affects outcome Do not adjust when estimating the total effect Sleep loss mediates an effect of night-time phone use on attention
Collider Common effect of two variables Conditioning can create an association and introduce bias Study participation is affected by both distress and access to support
Effect modifier Changes the size or direction of an effect Model and report interaction or subgroup effects when prespecified A coping intervention works differently by baseline social support
Precision variable Predicts the outcome but does not create a biasing path May improve precision if the model and sample support inclusion Baseline score predicts follow-up score

The DAGitty covariate-role guide illustrates confounders, mediators, proxy confounders, and competing exposures with causal diagrams. Its central lesson is that covariates cannot be classified without a defined exposure and outcome.

Why adjusting for everything can fail

Adding every available variable to a regression model may adjust for a mediator, condition on a collider, amplify bias through an instrument-like variable, or use too many parameters for the sample. It can also make the focal coefficient difficult to interpret.

VanderWeele’s open-access paper on principles of confounder selection explains why statistical associations alone cannot establish temporal or causal role. The paper proposes a theory-informed approach when full causal knowledge is unavailable and discusses qualifications involving instruments, proxies, and variable timing.

Build a causal model before choosing controls

Confounder selection should begin while refining the research question, not after looking at a correlation matrix. Use theory, prior evidence, measurement timing, and knowledge of recruitment to describe how variables may influence one another.

Step 1: define the estimand

An estimand is the precise effect the analysis aims to estimate. Common choices include a total effect, a direct effect that does not operate through a named mediator, or an association without causal interpretation. Most student projects claiming that an exposure “affects” an outcome intend a total effect.

Write the target in plain language before drawing a diagram. For example: “the total effect of perceived academic workload during the first month of term on burnout at the end of term.” The target determines whether later coping behaviour is a mediator rather than a baseline confounder.

Step 2: list plausible causes

List variables that may cause the exposure, the outcome, or both. Sources may include the dissertation’s theoretical framework, longitudinal evidence, validated temporal models, and contextual knowledge. Include unmeasured concepts because their absence matters to the bias assessment.

For academic workload and burnout, plausible prior causes might include programme demands, employment hours, caregiving duties, baseline mental health, personality-related vulnerability, and financial strain. Avoid selecting variables merely because they are routinely collected demographics.

Step 3: establish temporal order

A confounder usually needs to exist before the focal exposure for a total-effect analysis. A variable measured at the same time may have existed earlier, may result from the exposure, or may reflect both. Measurement time and causal time are not always identical.

Use the operational definitions to state reference periods. “Current stress” and “workload during the previous four weeks” leave ambiguous order. Baseline measures, dated events, or repeated assessments can make the assumed sequence more credible.

Step 4: draw a directed acyclic graph

A directed acyclic graph, or DAG, uses arrows to represent assumed causal directions. It must not contain a loop that returns to the same variable. Mark the exposure and outcome, add shared causes, mediators, selection processes, and important unmeasured variables, then identify open backdoor paths.

DAGitty is a free environment for drawing and analysing causal diagrams. It can identify adjustment sets implied by a graph, but it cannot decide whether the graph is scientifically correct. That judgement must come from substantive knowledge and transparent assumptions.

A review of DAG use in applied research found considerable variation in how diagrams, estimands, and implied adjustment sets were reported. The authors recommend clearer, more transparent use of DAGs in the International Journal of Epidemiology. For a dissertation, include the final DAG in the methods or appendix and explain the reasoning behind its arrows.

Psychology-specific confounding examples

Social media use and depressive symptoms

Age, prior depressive symptoms, loneliness, sleep patterns, and stressful life events may be considered. Their roles depend on timing. Loneliness measured before the exposure period could influence both platform use and later symptoms. Loneliness measured after sustained platform use could instead be a mediator. Contemporaneous measures may not distinguish these possibilities.

A defensible design might measure baseline depressive symptoms and relevant prior causes before tracking social media exposure, then assess the outcome later. The methods should state which variables form the primary adjustment set and why.

Exercise and cognitive performance

Physical health, socioeconomic resources, sleep, age, and previous cognitive ability may affect both exercise behaviour and test performance. Motivation could influence both, but a motivation scale measured after an exercise programme may partly capture an intervention effect.

Do not interpret every adjusted coefficient as a separate causal effect. The model may be designed only to estimate the effect of exercise. Estimating the causal effect of sleep would require its own question, causal diagram, and adjustment set.

Therapy attendance and symptom change

Initial symptom severity can affect treatment attendance and subsequent outcome. Referral route, service access, comorbidity, readiness for change, and social support may also matter. In routine service data, attendance is not randomly assigned, so a simple comparison of attendees and non-attendees may be strongly confounded.

Variables recorded after therapy begins, such as therapeutic alliance or early improvement, may be mediators or time-varying confounders. Standard baseline adjustment may not resolve those more complex structures. Keep claims proportionate to the design and available methods.

Classroom intervention and achievement

When entire classes or schools receive an intervention, school resources, teacher experience, baseline achievement, and selection into participating schools can affect both exposure and outcome. Clustered data also violate ordinary independence assumptions. The causal question, allocation process, and multilevel structure must be considered together.

Control confounding through study design

Analysis cannot repair a confounder that was never measured well. Design-stage prevention is usually stronger than adding covariates after data collection.

Strategy How it helps Main limitation
Random assignment Balances measured and unmeasured causes in expectation May be infeasible or unethical; chance imbalance remains in small samples
Restriction Holds a confounder constant through eligibility Reduces generalisability and prevents studying the restricted variable
Matching Creates comparison groups similar on selected variables Cannot balance unmeasured causes and requires appropriate matched analysis
Repeated measurement Clarifies timing and supports baseline control Increases participant burden and attrition risk
Standardised procedures Reduces procedural differences across exposure groups Does not address background causes of exposure and outcome

True random assignment is central to many experimental designs. Natural exposure groups, self-selection, clinician decisions, and researcher-assigned conditions without randomisation require different assumptions. A quasi-experimental design may strengthen causal inference through comparison groups, interrupted time series, discontinuities, or other design features, but each has its own confounding threats.

Control confounding during analysis

Common analytic strategies include stratification, standardisation, multivariable regression, propensity scores, weighting, matching, and g-methods for time-varying settings. The appropriate method depends on the estimand, design, variable types, overlap between groups, sample size, and assumptions.

Multivariable regression

A regression model can adjust for a prespecified set of measured confounders. Correct variable selection is only one requirement. Functional form, interactions, measurement error, missing data, clustering, and model fit also matter.

Compare a minimally adjusted model with the primary adjusted model when that comparison serves a clear purpose, but do not choose the final set because it produces the preferred result. Report effect estimates with uncertainty, not only p-values. For continuous confounders, arbitrary categorisation can leave residual confounding and waste information.

Stratification and standardisation

Stratification estimates the exposure-outcome relationship within levels of a confounder. It is transparent for one or two categorical variables but becomes sparse with many strata. Standardisation combines stratum-specific estimates for a defined target population.

Propensity-score methods

A propensity score models the probability of exposure given measured covariates. It can support matching, weighting, stratification, or adjustment. Balance diagnostics are essential. A good exposure-prediction model is not automatically a good confounding-control model, and unmeasured confounding remains possible.

Sensitivity analysis

Sensitivity analyses examine how conclusions change under alternative plausible decisions. Examples include different functional forms, alternative defensible adjustment sets, complete-case versus missing-data methods, or quantitative assessment of an unmeasured confounder. These analyses should be planned where possible and interpreted as robustness checks, not proof that no bias exists.

A practical confounder-selection workflow

  1. State whether the aim is causal, descriptive, or predictive.
  2. Define the exposure, outcome, population, timeframe, and target effect.
  3. List plausible causes using theory, prior evidence, and contextual knowledge.
  4. Map temporal order and measurement occasions.
  5. Draw a DAG that includes measured and important unmeasured variables.
  6. Identify a sufficient adjustment set for the target effect.
  7. Check whether each selected variable can be measured validly and consistently.
  8. Choose a design and analysis compatible with the sample and data structure.
  9. Prespecify the primary adjustment set and sensitivity analyses.
  10. Report assumptions, crude and adjusted estimates, uncertainty, and remaining limitations.

Use the preregistration or protocol to distinguish confirmatory decisions from later exploration. If the final model changes, explain what changed and why.

How to report confounding in a dissertation

The methods chapter should name potential confounders, define how and when each was measured, explain the causal rationale, state the adjustment set, and describe the analytic method. The results should report descriptive information, missingness, crude and adjusted estimates, precision, and planned sensitivity analyses.

The STROBE cohort checklist asks observational researchers to define potential confounders, describe methods used to control confounding, report adjusted estimates with precision, and explain which confounders were included and why. Use the reporting guideline that matches the actual design.

Psychology Dissertation Confounding Variables

Methods wording template

The primary estimand was the total effect of [exposure] on [outcome] in [population and timeframe]. Potential confounders were identified before analysis from [theory, prior evidence, temporal knowledge, and causal diagram]. The prespecified adjustment set included [variables] because [brief causal rationale]. Variables measured after the exposure and considered potential mediators were not included in the total-effect model. The primary analysis used [method], with [planned sensitivity analysis] to examine [assumption].

Results wording template

The unadjusted association was [estimate and interval]. After adjustment for the prespecified set of [variables], the estimate was [estimate and interval]. The change is consistent with confounding by the measured covariates, but the observational design does not rule out residual or unmeasured confounding. Results were [similar or different] under [sensitivity analysis].

A changed coefficient does not prove that the adjusted estimate is unbiased. Discuss the direction and likely importance of remaining bias in the limitations.

Common mistakes and how to repair them

Mistake Why it is weak Better practice
Selecting confounders only by p-value A small sample can hide a real cause, and significance does not establish causal role Select from causal reasoning before outcome analysis
Adjusting for every measured variable May include mediators, colliders, or too many parameters Use a defensible adjustment set tied to one estimand
Calling demographics automatic confounders Age or gender may not create a biasing path for every question Explain the role of each variable in the causal model
Measuring a confounder after exposure The measure may capture a mediator or exposure consequence Use prior measurement or state the temporal ambiguity
Interpreting every regression coefficient causally Each exposure would require its own causal question and adjustment set Interpret the focal coefficient according to the specified estimand
Claiming adjustment removes all confounding Measurement error and unmeasured causes can leave bias Use cautious language and sensitivity analysis

Frequently asked questions

How do I know whether a variable is a confounder?

Define the exposure and outcome, then ask whether the variable plausibly causes both and precedes the exposure for the target total effect. Use theory, temporal evidence, and a causal diagram. Correlation or statistical significance alone is not enough.

Can a mediator also be a confounder?

Its role depends on the exposure time and target effect. A prior value may confound a later exposure-outcome relationship, while a value caused by the exposure may mediate its effect. Repeated measures can clarify the sequence.

Should I control for baseline outcome?

Baseline outcome often predicts exposure and follow-up outcome, so adjustment may be useful. However, the decision depends on the design, estimand, measurement quality, and causal model. Explain the rationale rather than treating baseline adjustment as automatic.

Does randomisation eliminate confounding?

Random assignment balances causes in expectation and is powerful protection against confounding. Small samples can still show chance imbalance, and nonadherence, attrition, missing data, or post-randomisation selection can complicate the effect being estimated.

Can regression fix unmeasured confounding?

No. Standard regression adjusts only for included, adequately measured variables under model assumptions. Design improvements, proxies in justified settings, natural experiments, instrumental-variable methods, or sensitivity analysis may help, but each adds assumptions.

Where should I put a DAG in my dissertation?

Place the main diagram in the methods chapter if it is central to variable selection. A more detailed diagram, alternative versions, and supporting evidence can go in an appendix. Ensure the text states the estimand and adjustment set implied by the graph.

Conclusion

Confounding control begins with a precise causal question, not a software menu. Define the target effect, map plausible causes and temporal order, make assumptions visible in a causal diagram, collect the necessary variables well, and select an adjustment method suited to the design. Then report both what the analysis addressed and what it could not.

Get ethical support with your causal model

If you need help reviewing a psychology dissertation confounding plan, seek feedback that preserves your authorship and complies with your institution’s research ethics and academic integrity rules. A responsible review can test whether your question, DAG, measurements, adjustment set, and interpretation agree without inventing evidence, analysing undisclosed data, or promising that statistical adjustment proves causation.

Leave a Reply

Your email address will not be published. Required fields are marked *