Psychology dissertation network analysis maps conditional relationships among symptoms, behaviours, beliefs, or other measured variables. It can reveal a pattern that a single total score or a collection of separate regressions may hide. A defensible network, however, requires more than an attractive diagram. The nodes must answer a clear research question, the estimator must match the data, and uncertainty must remain visible in the conclusions.
This guide explains how to design, estimate, check, and report a psychological network in a dissertation. It focuses on psychometric networks, where nodes are variables rather than people. It also shows why cross-sectional edges are associations, not evidence that changing one node will change another.
What does psychology dissertation network analysis estimate?
A network contains nodes and edges. Nodes may represent individual questionnaire items, carefully chosen subscale scores, behaviours, or momentary states. An edge represents a statistical relationship between two nodes. In a commonly used Gaussian graphical model, an edge is usually a partial correlation: the association between two variables after accounting for every other variable in the network.
This meaning differs from a zero-order correlation. Suppose sleep difficulty and concentration problems are correlated. Their network edge may become small after worry and fatigue are included. That result suggests the direct conditional association is limited within the specified set of variables. It does not prove that worry or fatigue causes the original correlation.
Network analysis also differs from structural equation modelling. A factor model often treats observed responses as effects of an unobserved construct. A psychometric network models conditional relations among the observed variables. These approaches answer different questions and can sometimes be combined, but one should not be presented as an automatic replacement for the other.
Decide whether a network answers the research question
Use a network when the pattern of interdependence is substantively important. Examples include connections among depression symptoms, links between academic stress and coping behaviours, or relations among sleep, pain, and mood. A network is less useful when the dissertation only needs one prespecified effect, a validated total score, or a causal estimate supported by an appropriate design.
State the estimand before selecting software. A between-person cross-sectional network asks which variables are conditionally associated across participants at one assessment. A temporal network from intensive longitudinal data asks how earlier states predict later states within a person, subject to model assumptions. A contemporaneous network describes same-occasion residual associations. These networks are not interchangeable.
Psychometric networks should also be distinguished from social network analysis. In social networks, nodes often represent people or organisations and edges represent social ties. In psychological networks, nodes usually represent measured psychological variables. The open Learning Analytics Methods chapter on psychological networks provides a worked explanation of this distinction and of conditional associations.
Define nodes before looking at the network
Node selection is a theoretical and measurement decision. List each proposed node, its scale, response options, scoring direction, time frame, and reason for inclusion. Decide whether the network will use items or subscales. Item networks preserve symptom-level detail but can become large and may contain near-duplicate wording. Subscale networks are more compact but may conceal important variation within constructs.
Avoid selecting nodes because they produce an interesting picture. Data-driven inclusion changes the question and can exaggerate apparent structure. Record exclusions and reverse scoring in the protocol. If different instruments measure nearly identical content, examine redundancy before estimation. Two nodes that are alternative wordings of the same experience can create a strong edge and distort centrality without adding substantive information.
Keep the psychological meaning consistent
Nodes should refer to compatible levels and time frames. Mixing a stable personality trait, a symptom during the past week, and a behaviour observed during one laboratory task may be defensible, but only with a theory explaining the joint system. If scores have opposite directions, recode or label them clearly so that an edge sign cannot be misread.
Measurement quality remains essential. Reliability problems, restricted range, floor effects, and differential item functioning can alter correlations and therefore edges. If groups or languages are compared, consider whether measurement invariance is required before interpreting network differences.
Match the estimator to the data
The familiar partial-correlation network assumes variables are suitably represented by continuous, approximately Gaussian data or an appropriate correlation transformation. Psychological data often violate that simple description. Binary symptoms, ordered response categories, counts, and mixtures of variable types require compatible models and transparent assumptions.
| Data structure | Possible network model | Main planning issue |
|---|---|---|
| Continuous, approximately normal variables | Gaussian graphical model | Check linearity, outliers, and covariance estimation |
| Ordered questionnaire items | Ordinal model or GGM using justified polychoric correlations | Ensure categories and sample support stable thresholds |
| Binary symptoms | Ising model | Check sparse categories and separation |
| Mixed continuous and categorical variables | Mixed graphical model | Specify each variable type correctly |
| Repeated measurements | Multilevel or vector autoregressive network | Separate within-person, between-person, and temporal relations |
Do not treat this table as an automatic selection rule. Distribution, sample size, missingness, research design, and estimator implementation all matter. Document the software, package version, correlation method, tuning parameter, and any transformation. A reader should be able to reproduce the exact network rather than infer settings from a plot.
Regularised and unregularised estimation
A network with many nodes has many possible edges. Estimating all of them from a modest sample can produce unstable values. Regularisation methods such as the graphical lasso shrink estimates, with some weak edges becoming zero. This can yield a sparse, readable network, but shrinkage also introduces bias and may remove small genuine associations.
Extended Bayesian information criterion selection is frequently used to choose among regularised models. Its tuning setting affects sparsity. Report that setting and justify it. Do not describe every retained edge as a confirmed relationship or every missing edge as proof of no relationship. An estimated zero reflects data, estimator, tuning, and model assumptions.
Unregularised estimation with model selection or thresholding may be appropriate in some settings. The choice should follow current methodological guidance, expected network structure, and simulation evidence relevant to the study. A useful dissertation can include a prespecified sensitivity analysis that compares reasonable estimators without choosing whichever produces the most appealing result.
| Decision | What to report | Why it matters |
|---|---|---|
| Correlation or likelihood input | Pearson, polychoric, mixed model, or another justified method | Determines what associations mean |
| Regularisation | Penalty, selection criterion, and tuning value | Changes sparsity and edge bias |
| Missing data | Method, variables used, and diagnostics | Complete cases may not represent the target population |
| Layout | Fixed plotting layout for comparisons | Node position is visual, not a statistical result |
| Robustness | Bootstrap type, repetitions, and stability summaries | Shows whether apparent features are precise |
Plan sample size for the intended network
There is no universal participants-per-node rule. Precision depends on the number of nodes, edge magnitudes, sparsity, distributions, estimator, missingness, and the conclusions to be drawn. A sample that produces a network plot may still be inadequate for ranking edges or comparing groups.
Plan around the primary inferential target. If the aim is accurate edge estimation, evaluate expected edge precision. If centrality is central to the research question, assess its expected stability. If two networks will be compared, plan for the relevant difference and group allocation. Simulation under several plausible data-generating conditions is often more informative than a generic ratio. Explain the assumptions and test less favourable scenarios.
Place this reasoning in the dissertation’s power analysis and sample size plan. Avoid retroactively claiming that a large-looking sample guarantees stability. The effective sample may be lower after exclusions or missing-data handling.
Handle missing data and unusual observations transparently
List missingness by node and examine patterns before choosing a method. Complete-case analysis can discard substantial information and can bias estimates when the retained cases differ systematically. Pairwise correlations use different participant subsets for different entries and can create technical or interpretive complications. Multiple imputation may be possible, but pooling networks requires a principled method rather than averaging attractive plots.
Align missing-data handling with the estimator and the assumed missingness mechanism. Conduct a sensitivity analysis when reasonable alternatives could change the network. The site’s guide to missing data in psychology dissertations explains how to document assumptions and diagnostics.
Outliers can strongly affect covariance-based networks. Inspect univariate and multivariate unusual observations, check whether values are valid, and use prespecified rules. Do not delete a participant solely because the network becomes cleaner. Report analyses with and without influential valid observations when their impact is consequential. Follow a documented outlier analysis.
Assess edge accuracy and network stability
A network plot without uncertainty is incomplete. Edge thickness can make small sample differences look definitive. Use bootstrap procedures suited to the model to obtain intervals around edge weights and to examine whether apparently different edges can actually be distinguished.
Case-dropping bootstrap analyses can assess how well centrality estimates remain related to the original estimates when subsets of participants are removed. Report the stability measure and its interpretation, not only a successful-looking graph. The University of Amsterdam repository provides the accessible paper Estimating Psychological Networks and Their Accuracy, which describes edge-weight confidence intervals and centrality-stability assessment implemented in bootnet.
Bootstrap intervals do not convert an exploratory network into causal evidence. Nor does an interval that excludes zero establish practical importance. Interpret magnitude, direction, precision, measurement quality, and theoretical plausibility together.
Interpret edges without causal overreach
A positive edge means two nodes are conditionally associated in the model. A negative edge means higher values on one tend to accompany lower values on the other after conditioning on included nodes. The sign can depend on coding and the selected node set. Omitted common causes, selection, measurement error, and conditioning can all influence the result.
Cross-sectional direction is especially limited. If insomnia and worry have an edge, the data alone do not show whether worry increases insomnia, insomnia increases worry, both processes occur, or an unmeasured factor affects both. Use language such as “was conditionally associated with,” not “influenced,” “drove,” or “led to.”
An absent edge is also conditional on the model. It can reflect a truly small relation, low power, regularisation, noisy measures, restricted range, or an indirect path through other nodes. Discuss these alternatives rather than treating blank space as proof of independence in the population.
Use centrality cautiously
Strength centrality sums the absolute weights connected to a node. Expected influence retains edge signs when summing them. Betweenness and closeness come from shortest-path ideas that may not map naturally onto many psychological processes. A high value is a mathematical property of the estimated network, not proof that a node causes the system or is the best treatment target.
Centrality depends on which nodes were included, how they were measured, and how the network was estimated. Redundant nodes can gain apparent importance because they share similar content. If centrality is reported, justify the chosen index, show its stability, avoid unsupported rank ordering, and state what substantive claim the index can and cannot support.
Compare networks with a prespecified question
Group comparisons may ask whether overall connectivity, a particular edge, or a complete structure differs. These are different hypotheses. Select the test and correction procedure accordingly. Do not infer a group difference because an edge is significant in one group and not significant in another; test the difference directly.
Use the same nodes, coding, estimator, tuning settings, and preferably the same plotting layout across groups. First consider measurement comparability and unequal sample sizes. A more weakly estimated network in a smaller group can look sparser even without a population difference. Limit the number of exploratory comparisons or label them clearly.
Separate cross-sectional and longitudinal networks
Cross-sectional networks mix between-person differences with any within-person processes and do not establish temporal order. Intensive longitudinal designs can separate average differences between people from fluctuations within people. A temporal edge can represent whether one variable predicts another at the next measurement after conditioning on other lagged variables. A contemporaneous network can describe residual same-occasion associations.
These models require adequate observations per person, suitable measurement intervals, and checks for trends and stationarity. The interval determines the process being captured: hourly worry and sleepiness may reveal something different from weekly assessments. Consult the guide to a psychology dissertation longitudinal study before treating repeated observations as ordinary independent rows.
A practical dissertation workflow
- Specify the question. Define the population, nodes, time frame, and estimand.
- Audit measures. Check scoring, distributions, reliability, redundancy, and missingness.
- Choose the model. Match the estimator to variable types and study design.
- Plan precision. Use simulation or model-specific evidence for edge, centrality, or comparison goals.
- Estimate reproducibly. Record software versions, functions, options, seeds, and exclusions.
- Evaluate uncertainty. Examine edge accuracy, centrality stability, and prespecified sensitivity analyses.
- Interpret conservatively. Separate description, statistical inference, theory, and causal speculation.
- Report completely. Provide the node list, analytic code, numerical estimates, uncertainty, and limitations where ethics and permissions allow.
Worked psychology example
Consider a dissertation on student distress with eight nodes: worry, sleep difficulty, fatigue, low mood, concentration problems, avoidance, perceived workload, and social disconnection. The research question asks how these experiences are conditionally associated across postgraduate students during an assessment period. It does not ask which symptom causes another.
The student preregisters an ordinal network because the items have five ordered response categories. They justify each node, check sparse categories and redundancy, document missingness, and simulate data under several plausible edge structures. They estimate the network, obtain bootstrap intervals, and examine the stability of expected influence. The results show a relatively precise positive edge between sleep difficulty and fatigue, while several visually different edges have heavily overlapping intervals.
The discussion states that sleep difficulty and fatigue were conditionally associated within the measured set. It does not recommend a sleep intervention based on centrality. It notes that cross-sectional self-report data cannot establish direction, the assessment period may limit generalisability, and omitted financial or health factors could affect the network. A longitudinal follow-up is proposed as a new design, not implied by the present data.
| Weak wording | Defensible wording |
|---|---|
| Sleep difficulty caused fatigue. | Sleep difficulty and fatigue had a positive conditional association. |
| Worry was the most important symptom. | Worry had the highest estimated expected influence, but rank uncertainty limited comparison. |
| No edge means no relationship. | No edge was retained under this estimator and tuning specification. |
| The network revealed treatment targets. | The exploratory pattern identifies hypotheses for longitudinal or experimental evaluation. |
Report the analysis so it can be evaluated
Describe the sample, recruitment, exclusions, node construction, response scales, missing-data method, estimator, correlation input, regularisation, tuning, software, uncertainty procedures, and sensitivity analyses. Give numerical edge estimates or a machine-readable matrix in addition to the plot. Explain visual conventions such as colours, line width, and layout. Make clear whether node placement was fixed or data driven.
Use a reporting checklist to catch omissions. The EQUATOR Network listing for reporting standards in cross-sectional psychological network analyses identifies guidance covering data, results, and statistical methods. Treat a checklist as a reporting aid, not proof that the design itself is valid.
Share code and de-identified data when consent, governance, and participant protection permit. If data cannot be shared, provide synthetic data, detailed code, a data dictionary, and a clear access procedure where feasible. Never disclose sensitive symptom combinations or small-group information merely to maximise openness.
Common mistakes to avoid
- Choosing nodes after viewing several networks without disclosing the exploration.
- Combining incompatible variable types in a default Gaussian model.
- Calling an edge causal or treating centrality as an intervention effect.
- Reporting a network plot without edge estimates or uncertainty.
- Using node position in a force-directed layout as an inferential result.
- Ignoring redundant items, missingness, outliers, and measurement comparability.
- Comparing two pictures instead of formally testing a prespecified difference.
- Describing a cross-sectional between-person network as a within-person process.
Frequently asked questions
How many participants are needed for psychology dissertation network analysis?
No fixed ratio works for every network. Plan for the number and type of nodes, plausible edge sizes, estimator, sparsity, missingness, and intended inference. Simulation is particularly useful when stability or group comparison is central.
Should nodes be questionnaire items or scale scores?
Use items when symptom-level relations answer the question and the sample supports the larger network. Use scale scores when construct-level relations are intended and aggregation is psychometrically justified. State what information aggregation removes.
Does a network edge show causation?
No. In a cross-sectional psychometric network, an edge is normally a conditional association. Causal interpretation requires stronger assumptions and an appropriate longitudinal, experimental, or causal design.
Is the graphical lasso always the best estimator?
No. It is a common regularised estimator, not a universal default. Data type, sample size, expected structure, and current methodological evidence should guide the choice. Report tuning and sensitivity decisions.
Can I compare networks from two groups?
Yes, if the comparison is prespecified, measurement is comparable, and the method tests the stated hypothesis. Use consistent nodes and settings, account for unequal precision, and control multiplicity where required.
What should accompany the network figure?
Provide node definitions, numerical edges, estimator settings, uncertainty or stability results, sample details, and sensitivity analyses. Explain edge colour and width and remind readers that layout distance is usually not an inferential statistic.
Which limitations are most important?
Discuss design-based limits on direction and causality, measurement error, omitted nodes, sample representativeness, estimator dependence, edge uncertainty, centrality stability, and the effect of missing-data decisions.
Conclusion
A strong psychology dissertation network analysis begins with a precise psychological question and ends with conclusions that respect the design. Define meaningful nodes, match the model to the data, plan for precision, disclose analytic choices, and treat stability checks as part of the main analysis. Most importantly, describe conditional associations as conditional associations.
If you need dissertation support, seek ethical guidance that helps you understand and defend your own design, code, and interpretation. Retain responsibility for the research, protect participant data, follow institutional rules, and acknowledge any permitted assistance transparently.
