Research story · Psychology, mathematics & evidence synthesis

The distance to
a different conclusion.

A result can be estimated precisely and still sit close to a different conclusion. Before treating a positive or negative regression coefficient as a settled account of a relationship, it helps to ask a second question: how far would one input have to move to bring that coefficient to zero?

My research, The Fragility Margin: A Closed Form for Sign Reversal in Meta-Analytic Regression, turns that question into a quantity researchers can report beside the coefficient itself. The result connects a simple algebraic identity to a difficult problem in evidence synthesis: the published number may be clear while the evidence assembled beneath it is uneven.

A literature assembled into a matrix

Meta-analysis combines findings from multiple studies. Meta-analytic regression goes further, using a matrix of synthesised correlations to examine several predictors together. It can produce influential statements about which characteristic retains an association with an outcome once the other characteristics are taken into account.

Those conditional statements depend on the whole arrangement of correlations. A trait’s association with an outcome and its overlap with other traits both matter. Moreover, different cells of the matrix may have been estimated from different collections of studies, with different measurements and participant populations.

A standard error addresses an important question about estimation uncertainty under the analysis’s assumptions. The fragility margin supplies a different piece of information: the distance from the current input to the point at which a particular conditional coefficient reaches zero.

The distance has a simple form

The calculation holds the predictor correlation matrix fixed, along with every other predictor–outcome correlation. It then varies one predictor’s own correlation with the outcome. For the specified nonsingular standardised regression, the margin is the absolute value of that predictor’s coefficient divided by its variance inflation factor, or VIF.

In compact form: margin = |coefficient| ÷ VIF.

The VIF describes how much a predictor overlaps with the other predictors. Equivalently, the margin multiplies the coefficient’s magnitude by the share of that predictor’s variation left unexplained by the others. This is the predictor’s overlap with its peers, rather than the model’s overall success at explaining the outcome.

Take an illustrative coefficient of 0.12 and a VIF of 3. The distance is 0.04 correlation units in the relevant direction. Moving that input to the threshold brings the coefficient to zero; crossing the threshold reverses its sign. The arithmetic requires no simulation when the necessary coefficient and collinearity diagnostic are already available.

What the audit reveals

The manuscript applies the calculation to two published articles examining personality and workplace outcomes. Together, their four regressions supply 22 coefficients. The audit covers the coefficients in those selected articles, making the denominator explicit rather than choosing only conspicuous examples.

One example concerns openness and interpersonal deviance. Its coefficient is approximately 0.008, with a margin of about 0.0072 correlation units. That places its conditional sign close to the threshold under the specified one-input change.

For comparison, the psychopathy coefficient in the same interpersonal-deviance model is approximately 0.274. With a VIF near 1.65, its margin is about 0.166. The contrast gives readers a direct account of the different distances involved, rather than requiring them to infer sensitivity from coefficient size alone.

The paper also divides each margin by the reported between-study standard deviation for the relevant input. Fifteen of the 22 scaled margins are below one; the median is 0.69. In this audit, many thresholds are therefore closer than one reported between-study standard deviation.

Three questions about uncertainty

That scaling is descriptive. A between-study standard deviation measures dispersion across studies; it is not the sampling standard error of a pooled estimate. A margin of 0.69 standard deviations does not give a probability that the sign is wrong, nor does it demonstrate that an attainable alternative selection of studies would move the pooled input by that amount.

The distinction separates three useful questions. How precisely was this estimate obtained? How far is the specified model from a sign threshold? And do the measurements and study populations make the comparison substantively coherent?

The last question matters even when the algebraic margin is comparatively large. The manuscript examines a counterfactual substitution using a later synthesis of psychopathy and workplace behaviour. Replacing the earlier association with the later input changes the conditional estimate’s direction in the otherwise fixed model.

The two syntheses differ in included measures and study composition. The comparison demonstrates why substantive reconstruction can matter beyond the uncertainty represented by an older interval. It does not, by itself, identify a universal true trait effect or establish an arithmetic mistake in the earlier regression.

Look beneath the pooled comparison

The paper’s reconstruction of one review identifies 22 articles behind a psychopathy–counterproductive-behaviour estimate and 12 behind a Machiavellianism estimate. Across 33 distinct articles, only one is shared. A side-by-side table can therefore conceal substantial differences in the evidence supporting its columns.

This motivates a practical reporting change: show the studies behind each estimate and the overlap between them. The reconstruction uses article-level information and one author’s coding of short sample descriptions, so it supports scrutiny of composition rather than a causal explanation of the difference between constructs.

The margin itself is similarly focused. It changes one input while holding the others fixed. A joint analysis would need to represent dependence among inputs and determine which alternative correlation structures are admissible. A larger one-input distance cannot certify the robustness of every conclusion drawn from the model.

Make the next question easier to ask

The proposal is to report the coefficient, its fragility margin and, where appropriate, the descriptive scale together. Add transparent information about study overlap and measurement choices, and readers gain a much clearer starting point for evaluating a conditional claim.

This is where mathematics and psychological measurement reinforce one another. Algebra identifies exactly what changes under a declared perturbation. Substantive research determines whether that change corresponds to a meaningful comparison. Reporting both makes precision more useful: it shows not only where the estimate sits, but how the evidence supports the interpretation placed on it.