Imagine two instruments designed to distinguish between several competing explanations of a physical system. Both produce noisy observations. Which instrument gives us more information?
One way to compare them is to pick a particular task and see which performs better. But there is a more ambitious possibility: perhaps the observations from one instrument can reproduce everything the other would tell us. If so, it can stand in for the other across all decision problems.
This idea, formalised by Blackwell, is the starting point of our work on comparing statistical experiments. It connects questions about noisy observations to measures of distinguishability, convertibility of quantum states, catalysts, and even to betting games. Throughout these connections, the same question keeps returning: what can we do with the information we have?
A statistical experiment describes how the distribution of observations depends on an unknown hypothesis. With finitely many hypotheses, say d\geq2, we write P=\{p^{(k)}\}_{k=1}^d, where each p^{(k)} is the distribution predicted by hypothesis k.
We say that P is at least as informative as another experiment Q if a single processing rule can transform its observations into observations distributed exactly as in Q:
Tp^{(k)}=q^{(k)}\quad\text{for every }k=1,\ldots,d.
The crucial word is “single”. The rule cannot depend on which hypothesis is true—that is precisely what we do not know. For finite outcome spaces, the distributions form the columns of a matrix, and the processing rule is a stochastic matrix. The condition becomes Q=TP, which gives this comparison its name: matrix majorization.
Things become more surprising when we allow an experiment to be repeated a large number of times. In this way, a conversion that is impossible with one observation may become possible with many independent repetitions. In this large-sample setting, we ask whether, for every sufficiently large n, there is a processing rule satisfying
T_n\left(p^{(k)}\right)^{\otimes n}=\left(q^{(k)}\right)^{\otimes n}\quad\text{for every }k.
Another possibility is to borrow an auxiliary experiment—a catalyst. It helps the conversion succeed, while being returned unchanged and independent of the output, conditional on each hypothesis:
T\left(p^{(k)}\otimes r^{(k)}\right)=q^{(k)}\otimes r^{(k)}\quad\text{for every }k.
The catalyst can then be used again in a subsequent experiment. We restrict which catalysts are allowed; an auxiliary experiment that simply reveals the true hypothesis would make the problem trivial.
In our 2023 work with Muhammad Usman Farooq, Tobias Fritz, Erkka Haapasalo and Marco Tomamichel (see Matrix majorization in large samples), we identified quantities that govern these transformations. They are multivariate Rényi divergences: measures of how distinguishable a tuple of d distributions is. For distributions with a common finite support, their central expression is
D_{\underline{\alpha}}(P)=\frac{1}{\max_k\alpha_k-1}\log\sum_i\prod_{k=1}^d\left(p_i^{(k)}\right)^{\alpha_k}.
Here the parameters \underline{\alpha}=(\alpha_1,\ldots,\alpha_d) sum to one, i.e. \alpha_1+\ldots+\alpha_d=1. Either all satisfy 0\leq\alpha_k<1, or one of them exceeds 1 and the others are non-positive. Certain limits in these parameters also give valid quantities.
All these quantities cannot increase after performing some data processing on the distributions, also known as the data-processing inequality. Under the common-support assumptions, strict dominance in all these quantities guarantees conversion in large samples, and also catalytic conversion. Conversely, conversion requires the corresponding non-strict inequalities. Allowing arbitrarily small error in the target experiment gives a necessary-and-sufficient characterization of large-sample and catalytic conversion. The catalyst itself is still returned exactly.
The mathematical techniques behind these results use Tobias Fritz’s real-algebraic theory of preordered semirings and its so-called Vergleichsstellensätze, which are theorems that single out the quantities that characterize large-sample and catalytic transformations. See Abstract Vergleichsstellensätze for preordered semifields and semirings II.
The common-support setting leaves out a particularly revealing kind of observation: one that rules out a hypothesis altogether. An outcome with probability zero is qualitatively different from an outcome that is merely unlikely. In subsequent work in 2024 with Frits Verhagen, Marco Tomamichel and Erkka Haapasalo (see Matrix majorization in large samples with varying support restrictions), we allowed the hypotheses to have different supports—the sets of outcomes they consider possible. We studied both a broadly unrestricted case with at least some common overlap and a case where one distribution’s support contains all the others.
The result was more than a technical extension. Changing which outcomes are possible changes which divergences govern conversion, sometimes dramatically. The work also connects these conditions to catalytic transformations in quantum thermodynamics. In some sense, this work shows why the zero probabilities deserve as much attention as the positive ones.
The next extension moves beyond finite probability distributions. Instruments often produce continuously varying readings. The 2026 work Multivariate majorization of continuous statistical experiments by Erkka Haapasalo extends the large-sample and catalytic comparison theory of finite distributions to general standard Borel outcome spaces, while retaining finitely many hypotheses.
The results require the distributions within each experiment to have the same null sets and bounded likelihood ratios. Within this setting, multivariate Rényi divergences again give sufficient and almost necessary conversion conditions. For experiments whose hypotheses give pairwise distinct distributions, the optimal rate takes the form
r(P\to Q)=\inf_{\underline{\alpha}}\frac{D_{\underline{\alpha}}(P)}{D_{\underline{\alpha}}(Q)},
with the infimum over the admissible parameters. This rate measures how many target observations can be generated per source observation in the limit of many repetitions. Each divergence imposes an upper limit; the tightest limit determines the best possible rate.
There is also a more tangible way to understand these seemingly abstract quantities: through betting.
In recent joint work by Erkka Haapasalo with Andrés F. Ducuara and Ryo Takakura in the 2026 preprint Multivariate Rényi divergences characterise betting games with multiple lotteries, we consider an agent betting on several lotteries associated with the same random event. The agent’s preferences reflect different attitudes towards risk. Under the paper’s utility model and fair-odds assumptions, the optimised certainty equivalent—the guaranteed reward the agent regards as equally attractive as participating—is an exponential of a multivariate Rényi divergence.
The parameters thus acquire an economic meaning through attitudes towards risk. A conditional version also describes the value of side information available before placing bets. This connects the mathematics of distinguishability to an explicit decision-making task.
Quantum experiments add another dimension to the problem. Hypotheses now correspond to density operators, and a quantum channel replaces the classical processing rule. Quantum states need not commute, making their comparison substantially richer than that of classical distributions.
In a the 2025 paper Conditions for Large-Sample Majorization of Pairs of Flat States in Terms of α–z Relative Entropies by Frits Verhagen, Erkka Haapasalo and Marco Tomamichel, we study pairs of flat states: states that are uniform on their supporting subspaces, together with certain generalisations.
For these pairs of states, the relevant divergences are quantum \alpha–z relative entropies. These quantities give strict sufficient conditions for exact large-sample and catalytic conversions, and necessary-and-sufficient conditions in approximated settings. They also determine optimal conversion rates under the stated assumptions. The \alpha and z parameters are independent in this setting, giving this two-parameter family an operational interpretation.
This provides a concrete quantum extension of the classical theory, while pairs or tuples of general quantum states remain beyond the current scope.
These developments raise a deeper question: why do these particular measures of information appear in various settings?
The 2025 paper Barycentric decompositions for extensive monotone divergences by Erkka Haapasalo investigates measures that cannot increase under processing and that are additive when independent experiments are combined. In its algebraic framework, these measures can be expressed as positive mixtures of a special family of basic extremal divergences. The classical family can be fully identified.
Quantum theory, however, has a richer structure. Many established quantum divergences are extremal members of the convex set of all divergences: they cannot be obtained by averaging other divergences. This helps explain why several different quantum generalisations of classical Rényi divergences have enduring roles. This work’s results gives a structural explanation for that diversity.
The quantities of distinguishability between quantum states or probability distributions that we uncover in our studies find an operational meaning in settings related to the conversion of states, specifically in the large-sample or catalytic setting. Extending our findings to more general quantum states or to tuples of more than two states would give even more insight into this rich mathematical and physical framework.