Research story · Psychology · Measurement · Evidence synthesis

Comparing traits,
comparing evidence.

Two research syntheses can disagree about personality because they have compared different evidence under the same labels. Before asking which conclusion wins, ask who answered the questions, whose behaviour was recorded, which items were used, and whether the people being compared were the same.

My reanalysis of seven Dark Triad databases brings those design choices into view. It examines a long-running question: how should researchers distinguish Machiavellianism and psychopathy when their measures often overlap, yet some findings suggest different relationships with harmful conduct?

The opportunity is to make evidence synthesis more informative. A pooled correlation carries more meaning when readers can see the measurement process and study population that produced it.

Carry the design with the effect size

The study, Machiavellianism and Psychopathy Separate on Conduct That Costs Other People, reuses published tables and reported inclusion rules. Seven databases contribute to the design comparison; six supply effect estimates for 29 outcome categories. This is a reanalysis of existing syntheses, rather than seven new experiments or a new search of the entire primary literature.

For each category, the analysis compares psychopathy’s correlation with the outcome against Machiavellianism’s. It then records three features that can disappear inside an overall average: the source of the outcome score, whether the outcome concerns conduct imposing a cost on another party, and the instrument family used to measure the traits.

All the trait measures in this collection are participant self-reports. Outcome measurement varies. Keeping those two sides separate is essential: a supervisor’s outcome rating does not turn the predictor into an independently observed personality assessment.

What the broad average captures

Across 25 archival outcome categories, the pooled difference is 0.009, with a reported 95 per cent interval from −0.017 to 0.036. The manuscript reports equivalence within a stated comparison bound of 0.05. That describes the pooled contrast in this collection, rather than proving that two constructs are identical in every respect.

The composition of the evidence clarifies its reach. Twenty-three of those 25 categories are participant-supplied, one uses another source and one mixes sources. Much of the comparison therefore concerns relationships among scores people supplied about themselves.

That is substantive evidence about how the measures relate. The next question is whether the same pattern appears when the outcome concerns harmful conduct and includes information beyond the participant’s own account.

Keep the complete comparison visible

Five categories meet the study’s combined criteria of harmful conduct and outcome scoring that is not purely participant-supplied. All five are mixed-source categories: the original syntheses combine participant reports with other reports or records. They should not be described as five wholly independent observations of conduct.

Across all five, the pooled difference is 0.029, with a reported interval from −0.063 to 0.122. The estimate is imprecise and the categories differ substantially. Restricting the pool to the four contemporary categories gives 0.065, with an interval from 0.015 to 0.116.

Both results belong in the account. The four-category estimate points towards a stronger association for psychopathy in that subset. The complete five-category comparison does not supply the same clear pooled separation. Their difference creates a research question about evidence design, rather than permission to discard the inconvenient category.

The intervals are also approximate: source reviews and outcome categories share parts of the primary literature, and their dependence is not fully recovered from the printed tables.

A comparison between literatures

The older counterproductive-work-behaviour category points in the opposite direction. Reconstructing its article sets reveals a striking asymmetry. Nineteen of the 22 articles supplying psychopathy effects were coded as involving police, military or prison staff. None of the 12 articles supplying Machiavellianism effects fell into that category.

The sets share one article among 33 distinct articles. These are counts reconstructed from printed descriptions, not participant percentages, and shared article identity does not guarantee identical participants.

The finding establishes a problem of comparability. A difference between two correlations may incorporate differences between the populations and instruments supporting them. The occupational imbalance does not, by itself, prove what caused the contrary estimate. It identifies something a clean construct comparison must address.

Shortening a scale can change the question

The instrument analysis compares full-length measures, the Short Dark Triad and the Dirty Dozen across common outcome categories. It focuses particularly on content concerning disinhibition, such as impulsive or risky conduct.

In the available disinhibition-related subsets, the reported pooled differences are 0.149 for full-length measures, 0.211 for the Short Dark Triad and −0.033 for the Dirty Dozen. Those subsets contain only two or three outcomes, depending on availability.

The cross-instrument inference needs its own design check. Comparing the original measures with the Dirty Dozen category by category yields a reported p-value of 0.067, whereas treating the pooled families as independent gives 0.019. The paper relies on the paired comparison. Instrument families were not randomly assigned, so the contrast is evidence to investigate rather than a causal verdict about shortening scales.

Separate the construct from the content

A further explanation remains open. If a personality scale asks about rule breaking or conduct closely related to the outcome, part of its association may come from that overlap. Moving the outcome assessment to another informant does not remove content already embedded in the predictor.

The current evidence cannot separate this criterion-contamination explanation from a genuine distinction between constructs. The next useful comparison would measure the traits in the same participants, obtain suitably independent outcomes, and examine scores with and without the relevant conduct items. Such a design would make the competing explanations more directly testable.

A more useful synthesis

The practical recommendation is to carry outcome source, study overlap and instrument content into the research record alongside each effect size. These details show what a comparison is capable of establishing and where a better measurement design would add most.

The findings concern group-level research associations; they do not validate screening or decisions about individual people. Their value is methodological: connecting psychological theory, measurement and statistical synthesis so that apparently conflicting conclusions can be traced to the evidence that supports them.