Research story · Mathematics, AI & engineering assurance
Equivalent optima.
Different policies.
Two mathematical formulations can have exactly the same best physical decision and still lead an AI system to behave differently. Research on distribution-preserving reformulations identifies what must survive the journey from a solver model to a stochastic decision layer.
An engineering team rewrites a scheduling model to make it easier for a solver to process. Every feasible production schedule remains available. Every physical cost remains the same. The best schedule is unchanged. If the model feeds an AI system that samples decisions or learns a probability distribution, another correctness question has just appeared.
My manuscript, Equivalent Optima, Different Policies: Certified Distribution-Preserving Reformulations for AI Decision Systems, studies that interface. It asks when a change of mathematical representation preserves the behaviour expected by the component that consumes it, including both its output probabilities and its learning objective.
More descriptions can change what gets selected
Consider a production quantity represented by one nonnegative integer. Rewrite it as the sum of two nonnegative integers. A quantity of two now has three descriptions: zero plus two, one plus one, and two plus zero. The physical output is identical in every case, but the encoded model contains several ways to describe it.
A deterministic optimiser concerned only with the physical cost may be indifferent to that multiplicity. A sampler that gives equal weight to complete descriptions can favour physical decisions with more descriptions. The transformation has preserved what is possible and what is cheapest while changing how often a decision is selected.
The paper makes the decoder explicit: it maps each encoded assignment back to the physical action it represents. It also treats the reference weights as part of the model. Uniform weighting over descriptions and uniform weighting over physical decisions are different choices, with different consequences.
Certify the distribution the consumer needs
The first correctness requirement concerns inference. At a fixed parameter setting, does decoding the transformed policy produce the intended probability distribution over physical decisions? A representation should be checked against that requirement, rather than receiving a single undifferentiated label of “equivalent.”
For the finite log-linear policy families studied in the manuscript, the research derives an exact characterization. Group encoded descriptions by their decoded physical decision, account for their reference weights and compare the complete distributions of the extra score features introduced by the encoding. Those weighted residual spectra must agree across the groups.
The result applies across the declared parameter domain under its stated support and reference assumptions. Testing a handful of parameter values or comparing only a mean and variance cannot supply that universal certificate. The checker retains the full finite information needed to determine whether the condition holds.
Learning can impose a second contract
Even a representation that passes the distribution test can change learning. The distinction appears when the learner penalises departure from a reference distribution over the complete encoded assignment, including auxiliary variables. That penalty can contain an extra term invisible in the physical output.
If physical and auxiliary behaviour share parameters, the learner may change its physical policy to reduce the auxiliary penalty. At the same parameter value, the original and decoded policies still agree. After training on different objectives, the selected parameter values and physical policies can differ.
The paper separates these two contracts. It characterizes when the complete-policy relative-entropy objective is preserved and identifies the remaining nonnegative penalty when it is not. It also provides a certified correction that removes the extra term and recovers the physical objective.
This distinction avoids overdiagnosis. A learner whose loss depends only on the preserved physical output has no additional discrepancy of this kind. Untied auxiliary parameters can also absorb an auxiliary penalty without altering the physical optimum. In a separate convex policy-class result, a constant penalty can shift objective values while leaving every optimal physical policy unchanged.
A controlled demonstration of the mechanism
The computational work uses designed scheduling problems with explicit production, demand and ramp constraints. It includes small exhaustive checks, structured dynamic programmes, analytic controls and matched learning experiments. The tests examine the proposed mechanism rather than estimating how common it is in deployed industrial systems.
One experiment adds auxiliary bits whose behaviour shares the physical policy’s parameters. Marginalising the bits preserves the physical distribution at every fixed parameter. Training with the complete-policy penalty nevertheless changes the resulting physical policy in the constructed cases.
Across twenty designed cost vectors at a regularization coefficient of one, four tied auxiliary bits per period produce a median total-variation distance of 0.24468 from the intended physical optimum. This measures a difference between probability distributions, not a percentage of failed decisions or a measured productivity loss.
The corrected and untied implementations agree with the physical implementation to numerical precision in the reported study. Separate analytically solved constructions demonstrate the mechanism at a global optimum; the multidimensional tied-bit numerical runs are not presented as globally certified solutions.
Certificates need a maintenance rule
A useful guarantee must remain attached to the model it actually checked. The implementation uses exact rational arithmetic and rejects malformed inputs, incompatible reference weights and changed source specifications. A bounded-integer compiler produces both transformed models and the information needed to check their consumer-specific contracts.
Adding a later restriction can invalidate an earlier certificate even when every physical decision remains feasible. The restriction may remove different proportions of the descriptions representing different decisions. Reference weights may need refreshing, after which the distribution and learning conditions must be checked again.
Composition is equally explicit. Compatible certified layers can be combined, while their auxiliary penalties add. A later representation change cannot simply be assumed to cancel an earlier penalty. The result depends on matching intermediate references and score features.
Correctness belongs at the interface
The implementation has a defined boundary: finite supports, declared decoders and references, specified policy families and a bounded rewrite fragment. Exact checking of supplied information does not establish that an arbitrary implicit model has been completely enumerated, or that every industrial solver uses the certified weights.
The broader engineering contribution is a clearer division of responsibility. Optimisation verifies feasible actions and costs. Probability semantics verifies how actions are selected. Learning analysis verifies the objective being trained. Software assurance binds those checks to the representation that reaches the consumer.
That makes reformulation a question about meaning as well as mathematics. Before declaring two models interchangeable, identify what the downstream system does with them. Preserving the optimum is one contract. Preserving the behaviour of an AI decision system can require more, and the research makes those additional obligations precise enough to test.