Research story · Accounting, supply chains & measurement

Who enters
the data?

A researcher opens a database of supplier–customer relationships and finds a customer responsible for eight per cent of a company’s sales. Should that observation stay? Deleting it may seem to make the sample more consistent. But the deletion changes both the companies being studied and the measure used to describe their customer dependence.

My research, The Disclosure Threshold as a Selection Boundary: Estimating the Voluntary Disclosure Rate for Major Customers, makes that choice measurable. It asks what a reporting threshold can reveal about relationships that enter the data voluntarily, and what happens when researchers remove them.

The missing denominator

The historical reporting setting centres on a ten-per-cent revenue boundary. The standards examined in the paper require disclosure of the existence and amount of a relationship at or above that boundary. Disclosure of the customer’s identity is a separate issue, governed by a different set of considerations.

Below the boundary, some relationships still appear. The difficulty is calculating how often. The database records disclosed relationships; those never disclosed are absent. Counting visible eight-per-cent customers supplies a numerator, while the total number of such customers remains unknown.

This problem reaches beyond accounting. A supply-chain map, a concentration measure or a study of commercial dependence can inherit the selection process that made its underlying records visible. Understanding the observed network requires understanding how its edges entered the archive.

A boundary that reveals selection

The study applies an established density-discontinuity identification argument to this reporting setting. Its contribution lies in the application, estimate and consequences for research design.

The intuition is straightforward. Suppose the true distribution of relationship sizes is smooth and positive around ten per cent. Relationships just below and just above the boundary then provide a local comparison. If the probability of being reported changes at the cutoff, the observed density can jump even though the underlying economic density does not.

Under those conditions, the ratio of observed densities immediately below and above identifies the ratio of reporting probabilities on the two sides. The unknown underlying density cancels. This recovers a local, relative reporting rate from a database whose undisclosed relationships cannot be counted directly.

The smoothness condition carries real weight. Trading relationships might themselves respond to a reporting rule, and recorded shares can contain measurement error. The interpretation therefore depends on distinguishing a change in observation from a change in the economic relationships being observed.

What the data show

The paper constructs 168,118 disclosed supplier–customer–year observations involving 10,674 suppliers over fiscal years 1976–2026. Customers appearing across several supplier segments are combined, and customer sales are expressed relative to a carefully constructed supplier-revenue denominator.

Reporting is heavily rounded. The baseline removes whole-percentage-point heaps, leaving 98,408 relationships for the local-density estimation. Estimated density is 2.861 just below the threshold and 5.198 just above it. Their ratio is 0.550, with a reported bootstrap 95 per cent interval of 0.530–0.573.

Read that result as follows: near the cutoff, the estimated reporting rate below the line is about 55 per cent of the rate above it, given the identification assumptions. It is not an estimate that 55 per cent of all customers are disclosed. Recovering an absolute probability requires information about compliance above the threshold.

The estimate also has a sensitivity profile. Several alternative denominators, industry restrictions and rounding treatments remain near the baseline. Results change with other choices: a specification excluding a narrow interval around the cutoff reports 0.395. Making those differences visible gives readers a more useful account than treating 0.550 as invariant.

Deleting observations changes the question

Removing every subthreshold relationship eliminates 16,191 of 92,108 supplier-years entirely: 17.6 per cent of the panel. These are supplier-years whose disclosed customers all fall below the boundary. Among those remaining, 19.9 per cent lose at least one relationship.

Yet total disclosed customer-share rankings among the survivors remain highly similar, with a Spearman correlation of 0.980. That favourable result matters. The deletion rule does not scramble every ranking or invalidate every measure.

It does leave two distinct changes to investigate. First, some supplier-years disappear. Second, concentration is recomputed for those that remain. Stable rankings within the surviving group cannot establish that losing part of the original population is inconsequential.

The paper separates these channels in conventional regressions of supplier outcomes. With concentration measured as total disclosed customer share, the operating-margin coefficient becomes 43 per cent larger in magnitude after deletion. The capital-intensity coefficient increases by 19 per cent. A positive association with revenue per employee weakens and crosses the conventional significance threshold.

A Herfindahl measure is substantially steadier in these comparisons, with coefficients changing by about seven per cent or less. The research therefore supplies a differentiated result: sensitivity depends on the measure and outcome. These regressions describe conditional associations, rather than causal effects of concentration, and the outcomes have different coverage in the panel.

The denominator can manufacture a historical break

A second finding concerns the construction of supplier revenue. The source files can contain several reporting vintages for the same segment and fiscal year. Summing them as though each represented additional revenue inflates the denominator and pushes measured customer shares down.

In the paper’s reconstruction, that mistake creates an apparent break around the accounting-standard transition in 1998. Collapsing to the latest reporting vintage removes the artificial break in this construction.

The lesson is especially important for robustness checks. Comparing the same companies over time or switching segment categories can preserve an error if every calculation inherits the same duplicated denominator. More variations of a calculation cannot repair a shared misunderstanding of what its rows represent.

Make selection part of the research design

The practical contribution is a transparent sequence: establish what the reporting rule observes, construct the relationship and revenue units correctly, examine selection near the boundary, and report what a sample restriction changes.

That approach links accounting institutions, econometric identification and supply-chain research. The reporting archive becomes something to explain as well as something to analyse. A defensible dataset is built by making its observation process explicit, so readers can see which commercial relationships support the result and which remain outside its view.