Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Validation audit report — 2026-06

San Diego State University

1. What was verified

Every row is backed by a committed test and a regenerated figure whose measured value is shown against the tested tolerance (no self-consistency-only checks).

Module

Page

Tests

Figures

Key anchor(s)

Date

profiles/Plummer

Plummer equilibrium

20

5

a=0.7664rha=0.7664\,r_h exact; Beta(3/2,9/2)\mathrm{Beta}(3/2,9/2) speed law; Q=0.50Q=0.50

2026-06-08

profiles/King

King profile validation

32

5

c(W0)c(W_0) vs King 1966 Table II (Δ≤0.002); density vs DF oracle (1e-10)

2026-06-08

profiles/EFF

EFF profile validation

23

5

slope γ\to-\gamma; γ=5\gamma=5\equiv Plummer (exact); f(E)0f(\mathcal{E})\ge0

2026-06-08

kinematics/Michie

Michie-King anisotropy validation

12

5

β(r)\beta(r) vs DF moment oracle; isotropic King limit (2.7e-3)

2026-06-08

kinematics/rotation + OM anisotropy

Rotation & anisotropy validation

10

5

vϕ=ΩRv_\phi=\Omega R (exact); OM β=r2/(r2+ra2)\beta=r^2/(r^2+r_a^2) exact

2026-06-08

substructure/CW04 Q + azimuthal

Substructure (CW04 Q) validation

14

8

CW04 2004 Table 1 (3D0/1/2); σΣ\sigma_\Sigma–Q anti-correlation

2026-06-08

diagnostics/Λ_MSR + segregation

Mass segregation validation

8 + 13 + 27

9

Λ_MSR vs analytic ground truth; energy-sorted generator; differentiable observables (soft Λ_MSR / radial / Σ–m) vs exact soft→hard limit + Fisher identifiability

2026-06-09

cluster/two-component

Two-component populations (superseded)

1

per-component mass + dynamical-state recovery

2026-06-08

All four velocity DFs are true equilibria with no external virial rescale (unscaled Q=T/VQ=T/|V|: Plummer 0.50, King 0.505, EFF 0.502, Michie 0.501; 100% bound), and every released structural parameter (rh,rc,a,W0,γ,M,Ω,rar_h, r_c, a, W_0, \gamma, M, \Omega, r_a) passes an autodiff-vs-finite-difference gradient check (agreement 10-610-11).

2. Scientific trustworthiness

Verification rests on three independent legs:

  1. Physics-anchored tests — each asserts a quantitative match to an analytic result, an independent oracle, or a published table — never self-consistency.

  2. Independent oracles where no closed form exists (King density via direct velocity integral; Michie β\beta via DF 2nd moments) — these catch implementation bugs a self-referential test cannot.

  3. Paper-grounding against held PDFs (King 1966; Cartwright & Whitworth 2004; Merritt 1985; Elson, Fall & Freeman 1987), not memory.

A useful trust grading:

Tier

Meaning

Examples

A — exact/analytic

matches a closed form to ~machine precision

Plummer aa, EFF ρ(a)\rho(a) & γ=5\gamma=5\equivPlummer, rotation vϕ=ΩRv_\phi=\Omega R, OM β\beta, all gradients

B — published anchor

reproduces a published table within stated tolerance

King c(W0)c(W_0) vs Table II; CW04 Q vs Table 1

C — approximate / finite-N

statistically anchored or a documented surrogate

Monte-Carlo QQ/dispersions (±0.02–0.04); differentiable q_approx (kNN)

Limits stated honestly (not papered over)

Limitation

Status

EFF sharp truncation (γ=3\gamma=3) is ~5–8% sub-virial

intrinsic to truncating an empirical (non-DF) profile; use King for a strict lowered-DF equilibrium, or mild EFF truncation. Documented on the page.

Michie β\beta is suppressed below pure Osipkov-Merritt

real physics (lowering term breaks f(Q)f(Q)); validated vs the DF’s own moment oracle, not textbook OM.

q_approx over-reads ~0.1 for concentrated configs (Q>0.85Q>0.85)

faithful in the substructure regime it is used for; gradient is a kNN soft-surrogate (median AD-FD 0.9%, cell-boundary worst-case ~20%).

King scalar rtr_t not differentiable (argmax crossing)

deferred + documented; the profile shape in W0W_0 is differentiable.

No LIMEPY cross-validation

by design — validated against the original papers, not the external package.

Inference loop not validated end-to-end on data

gradients verified at the IC / forward-model level; structural-parameter recovery on real/mock observations is the research program, not a shipped result.

Bottom line: trustworthy for production ICs and a methods paper. Not a claim that the differentiable-inference pipeline is validated end-to-end against data.

3. Implementation completeness

Layer

State

Validated (publication tier: unit + integration + validation + figures)

Plummer, King, EFF, Michie; matched isotropic DFs + OM anisotropy + Michie anisotropy + rotation; CW04 Q + azimuthal variation; Λ_MSR + energy-sorted segregation; two-component clusters; imf-statistics (25 tests + 5 figures, 2026-06-08); binary-imf (24 orbital tests + 5 figures incl. the regenerated “confidently wrong” recovery, 2026-06-08); environment-imf (Marks+2012/Jeřábková IGIMF: new validation tier of 12 published-table tests + 5 figures, 2026-06-08); analytical-test-cases (two-body Kepler, Chenciner–Montgomery figure-eight, solar-system Kepler III, harmonic oscillator: 12 tests + 5 figures, 2026-06-09); tidal-truncation (Jacobi radius vs the restricted-3-body L1 point + analytic Plummer truncation: new validation tier of 9 tests + 5 figures, 2026-06-09).

Released, tested, not yet figure-validated ⚠️

(none — analytical-test-cases figure-validated 2026-06-09; see the validated row).

Released, unit-tested only ⚠️

(none — tidal-truncation hardened-and-validated 2026-06-09: new tests/validation/test_tidal_physics.py, 9 tests vs the L1 Lagrange point + analytic Plummer, + 5 figures; see the validated row).

Experimental (repo-only, not in the wheel) 🔬

gravoturb turbulent/fractal density-field subsystem (src/experimental/, AC1–AC17).

Deferred / not implemented ⏳ 🚧

differentiable King rtr_t (IFT, plan written); unified differentiable lowered-model family (Wilson/Woolley/multi-mass); multi-mass equipartition; self-consistent rotating equilibria.

4. What it lets us model

The architecture is compositional and differentiable end-to-end: any SpatialProfile pairs with any VelocityDF, plus IMF, binaries, anisotropy, rotation, and substructure diagnostics — flowing into gravax (dynamics) and fluxax (photometry). You can generate, as differentiable initial conditions:

5. Example science questions

Scoped to the validated, differentiable capability:

  1. Structural inference — jointly infer (W0 or γ,rc,M,ra,Ω)(W_0\text{ or }\gamma, r_c, M, r_a, \Omega) from projected density + kinematics by HMC, with calibrated uncertainties.

  2. Anisotropy-vs-rotation degeneracy — two distinct anisotropy routes (exact-OM stretch vs suppressed self-consistent Michie) + rotation transforms: can projected data distinguish radial anisotropy from rotation?

  3. Primordial vs dynamical substructure — generate clumpy/smooth ICs, quantify with the validated Q / azimuthal metrics, evolve in gravax: how fast does measurable substructure erase, and does mass segregation build dynamically vs primordially?

  4. Binary imprint on the IMF — with the Moe & Di Stefano engine + binary-aware likelihood: how biased is a single-star IMF fit with unresolved binaries, and at what NN is it “confidently wrong”?

  5. Tidal-field coupling (once rtr_t is differentiable — see roadmap) — can a population of cluster limiting radii constrain the Galactic potential?

6. Improvement recommendations

Recommendation

Why / payoff

R1

Differentiable King rtr_t (Approach B / implicit function theorem)

unlocks tidal-field inference (science Q5); plan already written, ~15–20 LOC + a 6th gradient panel.

R2

Figure-validate IMF / binary / analytical (61 tests already pass)

closes the three ⚠️ released gaps; scripts/validate_imfs.py already exists — needs pub-figure rewrite + measured tables + pages.

R3 ✅

Tidal-truncation validation tier — DONE 2026-06-09

tests/validation/test_tidal_physics.py validates jacobi_radius / apply_tidal_truncation against the restricted-3-body L1 point + analytic Plummer; 5 figures.

R4

q_approx: extend calibration above Q0.85Q\sim0.85 or hard-document scope; consider a soft-rank neighbour surrogate

removes the concentrated-regime over-read and smooths the gradient cell-boundary noise.

R5

EFF sharp-truncation: optional virial rescale or more prominent guidance

lets users get a virial γ=3\gamma=3 IC without silently inheriting the ~5–8% sub-virial offset.

R6

End-to-end inference validation (SBC-style recovery on mocks through gravax/fluxax)

this is the real outstanding gap: gradients are verified at the IC level, not against data.

R7

Retrofit a Measured column onto the older figured pages (two-component; mass-segregation done 2026-06-09)

consistency with the new pages’ tolerance/measured convention.

7. Remaining modules to validate + plot

Released, tested, missing only figures/pages (do these next):

8. Incomplete — harden first, then validate + plot

These are not merely unplotted; they need implementation/hardening before a validation tier is meaningful:

9. Changelog

Date

Milestone

2026-06-08

Plummer, King (retrofit), EFF, Michie, rotation+OM-anisotropy, substructure/CW04 Q, azimuthal-variation — all to publication standard (5 figures each; +measured tables). Shared scripts/_plotstyle.py; LIMEPY reframed across 13 pages + lowered-model-family roadmap; validate_profiles.py retired. Released-core 830 → 866 (+36 validation tests). This report authored.

(prior)

Λ_MSR + energy-sorted segregation, two-component clusters, IMF + Moe binary engine, tidal machinery — implemented and tested.