Research story · Mathematics, econometrics & measurement
The price
of repetition.
An expensive measurement returns a binary answer. Should you call the instrument again? Repeating everything wastes resources; stopping too early can discard the very information needed to understand how the instrument behaves across different items.
My research, Sharp Identification Cost and the Rare Path Limit of Stopped Measurement, examines this decision mathematically. Its central result is striking: within a specified binary-response model, a carefully chosen stopping rule can preserve all the moment information available through a measurement horizon at a worst-case expected cost of two calls per unit.
The practical importance lies in the next result. The histories carrying some of that information can be extremely rare. A design can be economical in average calls while requiring a very large number of units to estimate its hardest features accurately.
What remains stable when we repeat?
The model begins with independent units, each having a persistent, unobserved probability of producing a one. Conditional on that probability, repeated calls are independent and use the same response law. Across units, the probabilities can vary in an unrestricted way.
This allows persistent differences in item difficulty or response propensity. It also makes the complete sequence informative: several responses from the same unit share the same underlying propensity, even though that propensity is unknown.
The assumptions describe a stable repeated-measurement setting. Learning, fatigue, anchoring or a changing device require a different model. Repeating a classifier does not automatically reveal ground truth, and recovering response-propensity groups does not assign those groups substantive meanings without further evidence.
Design the stopping tree
A stopping policy specifies what to do after each observed sequence, up to a maximum number of calls. Its final records might include a single response, a short sequence or a much longer one. Both the outcomes and the point at which measurement stopped are part of the data.
The paper translates those terminal histories into polynomials of the latent response probability. Their probabilities depend on a finite set of moments: the mean propensity, the mean squared propensity, and higher powers through the chosen horizon.
This produces a constructive identification test. A further split in the stopping tree adds information only when it contributes a polynomial relationship not already spanned by the existing records. Counting the number of possible final histories is insufficient. What matters is whether they provide genuinely independent information about the moment vector.
The two-call result
For binary horizons of at least three, the paper characterises the deterministic policies that minimise the worst-case expected call count while preserving the full moment vector. Exactly two complementary staircase policies achieve the bound of two.
One stops immediately after an initial one. After an initial zero, it continues while subsequent responses are ones, stopping at the next zero or the maximum horizon. The other exchanges zeros and ones.
Most units can stop quickly while a few continue much further. That is how the average remains at or below two even as the permitted horizon grows. Two is an expected acquisition cost across the response process, not a rule limiting every unit to two observations.
The qualification to binary responses is substantive. The paper also studies larger response alphabets and shows that the same constant does not carry over unchanged. The identification logic generalises more broadly than this particular cost optimum.
Rare histories set the price of precision
The longest branch is reached only by a particular response sequence. Under the paper’s condition that latent propensities stay away from the favourable endpoint, the chance of reaching that branch falls geometrically with the horizon.
The information about the final conditional response falls with that reach probability. The relevant effective sample size is the number of units multiplied by their probability of reaching the final call. A huge panel with almost no units on the critical branch can still tell us little about that conditional feature.
This distinguishes logical identification from useful estimation. Knowing that different moment vectors imply different observable distributions is one achievement. Collecting enough of the distinguishing observations to separate them precisely is another.
Randomisation makes the distinction sharper. Assign a small positive fraction of units to an identifying staircase and stop everyone else after one call. Logical identification survives, and expected cost approaches one as that fraction shrinks. At the same time, information in directions requiring repetition collapses. Exactly one call is an unattained limit of this identifying construction, rather than a free route to the same statistical precision.
A good fit can hide richer variation
A further result concerns the interpretation of latent classes. With measurements through horizon 2K−1, every distribution of stable response propensities has a representation on at most K points that matches the observed moments.
For example, a three-call experiment can be compatible with a two-point representation even when the underlying heterogeneity is richer. A well-fitting two-class model therefore need not establish that the world contains two genuine types. The representation can summarise everything that experiment observes while leaving other features unresolved.
One additional degree of measurement makes a support restriction testable within the maintained model. The paper derives a moment-matrix condition that richer support violates. Its local-power analysis also distinguishes a new type with very little weight from one whose propensity lies close to an existing type.
Passing that test with finite data does not prove the proposed number of types. Rejecting the restriction calls for a richer account, without identifying an arbitrary latent distribution or diagnosing whether an unstable response process caused the failure.
Preserve the experiment you actually ran
The design implications are concrete: record the assigned policy, retain complete terminal histories, check the stability of the response process, and budget for the information required by the question. Saving only the final answer discards the duration and path that made the stopping design informative.
The research connects econometric identification, probability, sequential acquisition and measurement practice. It gives a disciplined way to ask what repetition buys. The strongest design is one whose cost, identifying content, precision and interpretation have each been established, so that economical data collection produces evidence suited to its intended use.