The big idea¶
Neyrinck’s result — that Gaussianizing restores power-spectrum information — raises a sharper question: which non-linear transform is best, and how much information can any local transform recover? Carron & Szapudi answer it with Fisher information and cosmological perturbation theory. They show that at each perturbative order there is a polynomial that exhausts the information on a given parameter; this polynomial is the Taylor expansion of the maximally efficient “sufficient” observable. The corresponding optimal local transform “is essentially the simple power transform with an exponent related to the slope of the power spectrum; when this is -1, it is indistinguishable from the logarithmic transform.” The transform Gaussianizes the distribution and recovers the linear density contrast — a direct equivalence between undoing the non-linear dynamics and efficiently capturing Fisher information. Their transforms stay close to optimal even deep into the non-linear regime, .
Why this matters for the fat tail¶
The companion observation (building on Carron 2011) is the one that bites in the gravoturbulent problem: in the large-variance regime a large fraction of the information escapes the entire hierarchy of -point moments. When the density PDF has a heavy power-law tail — BM19 with , where formally diverges — the moments are dominated by the rarest cells and carry almost no information about the parameters. The fix is not “more moments” but a non-linear transform that Gaussianizes the distribution, after which the (transformed) two-point function is information-rich. This is precisely why progenax carries the log-density two-point and the peaks-over-threshold tail estimator rather than linear-density moments.
Use in progenax¶
Formal justification for the log-density variable used throughout Differentiable inference — natal cloud parameters from cluster substructure: the log transform is (near-)optimal for a spectrum, Gaussianizing the field and concentrating the information in .
Explains why the linear-density 2-point is avoided — its moments are information-poor and divergent for the collapsing slopes .
Notes¶
The discrete (Poisson-sampled) extension of these “sufficient observables” is Carron & Szapudi (2014); the empirical demonstration on the Millennium simulation is in Neyrinck, Szapudi & Szalay (2011).
“Optimal” here is for local (one-point) transforms; genuine phase information (filaments) still requires higher-order or morphological statistics, consistent with the 1pt+2pt scope of the progenax inference.
- Carron, J., & Szapudi, I. (2013). Optimal non-linear transformations for large-scale structure statistics. Monthly Notices of the Royal Astronomical Society, 434, 2961–2970. 10.1093/mnras/stt1215