Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

OED for robustness — when binaries lie to your mass estimate (Stage 4)

The first OED demo whose headline is a BIAS, not a recovery — a binary-blind survey design over-weighs a cluster by 184% with 41× false confidence, and a binary-aware design removes it

San Diego State University

The three OED demos before this one each end with the optimiser rediscovering a piece of physics: anisotropy wants the outskirts, concentration wants the core, dynamical mass wants an interior depth. Each is a clean result — and each is one a referee could wave away as “Fisher-OED recovers the obvious.” This page is different. Its headline is not a recovery; it is a bias. We hand a careful, binary-blind observer the optimal survey design for weighing a cluster, let the real cluster harbour the binaries that every real cluster has, and watch the analysis over-estimate the mass by 184% while reporting a 4.5% error bar — a measurement that is confidently, expensively wrong. Then we show the design that protects against it.

Why this arc exists: a bias a referee cannot dismiss

The sibling OED demos are strong engineering — one differentiable Fisher backbone, re-pointed at three science targets — but each headline is either already-known physics or contingent on an arbitrary RV/PM error split that flips with the budget. None is a result a referee couldn’t read as “of course the optimiser put the proper motions where the signal is.”

This demo monetises a capability nobody else has wired to OED: progenax carries a faithful, differentiable Moe & Di Stefano (2017) binary-population engine, so we can ask the question that actually bites in real cluster mass measurements — what does an unmodelled population systematic do to an “optimal” design, and what design defends against it? The answer is a parameter bias larger than its own forecast uncertainty, which is the one kind of result a Fisher-OED demo cannot be accused of staging. It is the canonical binary M/LM/L systematic of dwarf-galaxy and young-cluster dynamics, turned into a design problem.

The physical question: why binaries masquerade as mass

You weigh a star cluster with the virial theorem: a hotter cluster — larger line-of-sight velocity dispersion σlos\sigma_{\rm los} — is a heavier one, Mσ2R/GM \propto \sigma^2 R / G. That logic has a silent assumption: that every velocity you measure is a single star orbiting the cluster. It is not. A large fraction of massive stars live in binaries, and in a single-epoch radial-velocity survey you cannot resolve the binary — you measure the star’s orbital velocity around its companion folded into its cluster velocity. Those orbital motions are fast (tens of km/s for short-period massive binaries), so they inflate the measured dispersion above the true cluster value.

The inflation is not noise — noise averages out. It is a systematic offset in the second moment. If the cluster’s true dispersion is σcluster\sigma_{\rm cluster} and a fraction fbinf_{\rm bin} of the stars carry a binary velocity drawn from a population with variance VbinV_{\rm bin}, the observed dispersion-squared is

σobs2(R)  =  σcluster2(R)falls with radius  +  fbinVbinflat pedestal  +  εRV2,\sigma_{\rm obs}^2(R) \;=\; \underbrace{\sigma_{\rm cluster}^2(R)}_{\text{falls with radius}} \;+\; \underbrace{f_{\rm bin}\,V_{\rm bin}}_{\text{flat pedestal}} \;+\; \varepsilon_{\rm RV}^2,

a flat pedestal fbinVbinf_{\rm bin}V_{\rm bin} added on top of the real cluster heat at every radius. Feed σobs\sigma_{\rm obs} to a virial analysis that does not know about the pedestal, and the analysis reads the extra dispersion as extra mass. For our young-massive-cluster fiducial the pedestal is fbinVbin=6.9\sqrt{f_{\rm bin}V_{\rm bin}}=6.9 km/s sitting under a central cluster dispersion of only 9.0 km/s — and the cluster term falls to 0.4 km/s in the outskirts, while the pedestal does not budge. That is the whole story: in the cold outskirts the binary pedestal dominates the signal.

The intuition the optimiser is about to fall for — and then defeat

Hold one fact in mind: the cluster mass MM scales the dispersion amplitude and shows up at every radius; the binary fraction fbinf_{\rm bin} adds a flat offset. The only thing that tells them apart is radial leverage — the cluster dispersion has a shape σ(R)\sigma(R) that falls from core to edge, while the binary pedestal is flat. A design that ignores binaries chases the loudest signal into the cold outskirts and parks its budget exactly where the pedestal is largest relative to the cluster — and so confuses pedestal for mass most badly. A design that models binaries instead spends its stars on the core↔outskirts contrast that breaks the MfbinM\leftrightarrow f_{\rm bin} degeneracy. Same instrument, same budget, opposite outcome.

Inputs and assumptions

The design optimises a fixed budget of NtotalN_{\rm total} radial-velocity measurements allocated across K=12K=12 log-spaced projected-radius bins — RV only, single-epoch. The science target is the dynamical mass MM; the EFF shape parameters (γ,a)(\gamma, a) are photometrically pinned nuisances, and the anisotropy radius rar_a and binary fraction fbinf_{\rm bin} are kinematic nuisances the radial allocation must disentangle. Everything is computed pre-data — the cross-model Monte Carlo afterwards is a calibration gate, not part of the design loop.

Model inputs

Input

Meaning and role

Status (fiducial)

MM

Total dynamical mass — the science target; scales the cluster dispersion amplitude. c-optimality minimises its marginal variance.

target (M=4×105MM=4\times10^5\,\Msun; index 0 of θ\boldsymbol\theta, no prior)

rar_a

Osipkov–Merritt anisotropy radius — a kinematic nuisance with a weak (50% fractional) prior; the RV-only M ⁣ ⁣raM\!\leftrightarrow\! r_a degeneracy is regularised by it.

nuisance (ra=3r_a=3 pc; index 1; σprior/ra=0.5\sigma_{\rm prior}/r_a=0.5)

γ\gamma

EFF outer slope — the concentration shape knob. Photometrically pinned (measured from the surface-brightness profile).

nuisance (γ=2.7\gamma=2.7; index 2; tight 10% prior)

aa

EFF scale radius — parameterised directly (EFF is natively (a,γ,rt)(a,\gamma,r_t) with no closed-form rhr_h inversion; see the modelling choice). Photometrically pinned.

nuisance (a=1a=1 pc; index 3; tight 10% prior)

fbinf_{\rm bin}

Binary fraction — the misspecification axis. In the binary-aware model it is a free nuisance with a weak (50%) prior, recovered from radial leverage; in the binary-blind model it is silently fixed at 0.

truth 0.5 (index 4; weak prior or absent)

VbinV_{\rm bin}

Variance of the single-epoch binary velocity blend — a build-once population scalar from the Moe & Di Stefano PPqqee engine over a Maschberger massive-primary IMF with ZAMS flux weights.

fixed (Vbin=94.7V_{\rm bin}=94.7 (km/s)2^2 σbin=9.73\to\sigma_{\rm bin}=9.73 km/s)

density model

EFF (Elson–Fall–Freeman), projected under Osipkov–Merritt. Analytic density \Rightarrow no ODE / no OOM (see the modelling choice).

held fixed (the misspecification CONTROL)

kinematic channel

RV only (σlos\sigma_{\rm los}) — no proper motions, no channel split (the binary M/LM/L systematic lives in the RV second moment; the dwarf/cluster M/LM/L regime).

fixed (the design allocates stars among radii)

εRV\varepsilon_{\rm RV}

Per-star RV error, entering the per-datum variance δσ2=(σ2+ε2)/(2n)\delta\sigma^2=(\sigma^2+\varepsilon^2)/(2n).

fixed (εRV=1.0\varepsilon_{\rm RV}=1.0 km/s)

KK bins, budget, optimiser

K=12K=12 log-spaced bin centres Rk[0.2a,0.95rt]=[0.2,17.1]R_k\in[0.2\,a,\,0.95\,r_t]=[0.2,17.1] pc; rt=18r_t=18 pc; multi-start Adam over the softmax simplex.

numerical choices

units, GG

STELLAR (M\Msun, pc, Myr); errors converted to km/s explicitly.

known / fixed

The Fisher is built in the dimensionless lnθ\ln\theta metric Important, so σ(lnM)\sigma(\ln M) is a fractional precision. The pinned scales are the load-bearing ones: σcluster,central=8.98\sigma_{\rm cluster,central}=8.98 km/s and σbin=9.73\sigma_{\rm bin}=9.73 km/s, a contamination ratio σbin/σcluster=1.08\sigma_{\rm bin}/\sigma_{\rm cluster}=1.08 at the centre — rising to 24\sim24 in the outskirts, because the cluster term collapses while the pedestal stays flat. That the binaries rival the cluster heat at this operating point is pinned before any inference, exactly so that H1 below has a chance to bite.

The modelling choice: EFF-OM, RV-only, and a held-fixed cluster

Two facts about the forward model set up the whole demo, and stating them plainly is the honest way to read the result.

  1. EFF density is analytic — no ODE, no OOM. The Elson–Fall–Freeman profile has a closed-form density ρ(1+r2/a2)γ/2\rho \propto (1+r^2/a^2)^{-\gamma/2}, so project_dispersion does a plain quadrature with no differential-equation solver in the tape. This is what makes the cross-model Monte Carlo runnable — the King/Michie ODE solvers in the concentration demo forced its Michie calibration to OOM at \sim28 GB. The EFF slope γ\gamma doubles as a concentration knob, so the bias vector spans (M,ra,γ,a,fbin)(M, r_a, \gamma, a, f_{\rm bin}) with no ODE anywhere. We parameterise by the EFF scale radius aa directly rather than a derived half-mass radius, because EFF has no closed-form rh(a,γ,rt)r_h(a,\gamma,r_t) and inventing one would inject a spurious unpinned inversion.

  2. The cluster forward model is the held-fixed CONTROL. This is a misspecification study, and a misspecification study is only clean if it changes one thing. So the EFF-OM cluster model — density shape, OM anisotropy law, RV projection — is identical between the design, the truth, and both fits. The only difference between the “disaster” and the “fix” is whether the analysis models the binaries. Every number on this page is the cost of that one omission, with the cluster physics deliberately frozen so it cannot be blamed.

The method, specialised: c-optimality with a misspecified vs marginalized Fisher

This demo pours the additive backbone Equation into a c-optimal-for-MM radial allocation, exactly as the sibling designs do. One jacrev of the cluster term at the truth gives the per-bin Jacobian J=σlos/lnθJ=\partial\sigma_{\rm los}/\partial\ln\boldsymbol\theta (reverse-mode by policy); the design enters only through the per-bin counts. The criterion minimises the MM variance with the nuisances profiled out:

c ⁣:minz  (F1)MM,F(z)=bnbMb+diag(priors),Mb=2JbJb ⁣σobs2(Rb)+εRV2.c\!:\quad \min_{\mathbf z}\;(F^{-1})_{MM}, \qquad F(\mathbf z)=\sum_{b} n_{b}\,M_{b}+\mathrm{diag}(\text{priors}), \qquad M_b=\frac{2\,J_b J_b^{\!\top}}{\sigma_{\rm obs}^2(R_b)+\varepsilon_{\rm RV}^2}.

The whole demo is the contrast between two Fishers built on this one equation:

Because the Fisher is additive and design-linear, the expensive projection is differentiated once; the fbinf_{\rm bin} grid for the maximin criterion below only updates the σobs(fbin)\sigma_{\rm obs}(f_{\rm bin}) denominator — no re-jacrev.

The headline (H1): a confident, expensive, wrong mass

The pre-registered hypothesis (H1, locked before the Monte Carlo ran) was: fitting the binary-blind model to binary-contaminated RV data biases M^\hat M high, by more than the design’s own forecast σ(M)\sigma(M) — false confidence. The rival H0 was that the binary-blind fit would absorb the flat pedestal into its nuisances and stay unbiased. H1 is confirmed, decisively.

The binary-blind c-optimal-for-MM design, fit with the binary-blind model on mock catalogues that do contain Moe binaries, recovers

M^M=2.84(bias +184%),forecast σ(M)M=4.5%,biasforecast=41×.\frac{\hat M}{M} = 2.84 \quad(\text{bias } +184\%), \qquad \text{forecast } \frac{\sigma(M)}{M}=4.5\%, \qquad \frac{\text{bias}}{\text{forecast}} = 41\times.

The estimate is nearly triple the truth, and the design’s own error bar — the precision it promised — is 41 times smaller than the error it actually made. This is the worst failure mode in inference: not a wide error bar (which warns you), but a tight one around the wrong answer.

The control is what makes it airtight. Run the same design and fit on mocks with fbin=0f_{\rm bin}=0 — where the fit model exactly matches the generative model — and the bias is 0.3%-0.3\%, well inside the forecast. The entire +184%+184\% is the binaries; nothing else moved.

False confidence: a binary-blind design weighs the cluster at 2.84\times its true mass
with a 4.5\% error bar. The recovered \hat M/M_{\rm true} for the binary-blind
c-optimal-for-M design fit on binary-contaminated mocks (vermilion, \approx2.84), drawn
with its own \pmforecast-\sigma error bar — visually a speck. The unbiased truth is the
dashed line at 1; the f_{\rm bin}=0 baseline (blue square, mock \equiv fit model)
recovers \hat M/M\approx1. The bias (+184\%) is annotated against the truth; the callout
states the headline — the claimed error bar is 41\times smaller than the bias. From the
env-gated cross-model Monte Carlo (48 draws).

False confidence: a binary-blind design weighs the cluster at 2.84×2.84\times its true mass with a 4.5%4.5\% error bar. The recovered M^/Mtrue\hat M/M_{\rm true} for the binary-blind c-optimal-for-MM design fit on binary-contaminated mocks (vermilion, 2.84\approx2.84), drawn with its own ±forecast-σ\sigma error bar — visually a speck. The unbiased truth is the dashed line at 1; the fbin=0f_{\rm bin}=0 baseline (blue square, mock \equiv fit model) recovers M^/M1\hat M/M\approx1. The bias (+184%+184\%) is annotated against the truth; the callout states the headline — the claimed error bar is 41×41\times smaller than the bias. From the env-gated cross-model Monte Carlo (48 draws).

Why it happens: the design parks its budget in the contaminated outskirts. The truth
cluster \sigma_{\rm los}(R) (vermilion) falls steeply from 9.0 km/s in the core to 0.40
km/s at the edge — that fall is the mass signal. The binary-inflated observable
\sqrt{\sigma_{\rm cluster}^2+f_{\rm bin}V_{\rm bin}} (green) sits on a near-flat pedestal at
the \sqrt{f_{\rm bin}V_{\rm bin}}=6.9 km/s floor (dotted). The binary-blind design’s per-bin
allocation (blue bars, right axis) piles into the outer bins — exactly where the pedestal
dominates the signal — so the binary-blind fit reads the pedestal as extra mass. The design
optimised for M under the wrong model finds the worst possible place to spend its stars.

Why it happens: the design parks its budget in the contaminated outskirts. The truth cluster σlos(R)\sigma_{\rm los}(R) (vermilion) falls steeply from 9.0 km/s in the core to 0.40 km/s at the edge — that fall is the mass signal. The binary-inflated observable σcluster2+fbinVbin\sqrt{\sigma_{\rm cluster}^2+f_{\rm bin}V_{\rm bin}} (green) sits on a near-flat pedestal at the fbinVbin=6.9\sqrt{f_{\rm bin}V_{\rm bin}}=6.9 km/s floor (dotted). The binary-blind design’s per-bin allocation (blue bars, right axis) piles into the outer bins — exactly where the pedestal dominates the signal — so the binary-blind fit reads the pedestal as extra mass. The design optimised for MM under the wrong model finds the worst possible place to spend its stars.

The fix (measure-and-marginalize): model the binaries, recover them, lose the bias

The remedy is not to throw away the outskirts; it is to model the pedestal. Add the fbinf_{\rm bin} parameter to the fit (and to the design Fisher), let the data determine it, and the bias collapses:

That recovery of fbinf_{\rm bin} comes with a caveat sharp enough that it gets its own section below — the mock is generated at the truth fbinf_{\rm bin} and the prior is centred there, so the unbiased recovery alone proves only self-consistency. What proves fbinf_{\rm bin} is genuinely identifiable is the prior-insensitivity test, not this number.

H2 — the binary-aware design is also tighter

The fix is not only a better fit; it is a better design. Comparing the two designs under the same binary-aware (marginalized) fit, the binary-aware allocation reaches a target σ(M)\sigma(M) with fewer stars:

precision gain  =  σ(M)binary-blind designσ(M)binary-aware design  =  1.33×(at the operating point),\text{precision gain} \;=\; \frac{\sigma(M)_{\text{binary-blind design}}}{\sigma(M)_{\text{binary-aware design}}} \;=\; 1.33\times \quad\text{(at the operating point)},

confirming the pre-registered H2 threshold (1.3×\geq1.3\times). This margin is thin at the fiducial ratio σbin/σcluster=1.08\sigma_{\rm bin}/\sigma_{\rm cluster}=1.08 on purpose — and the sweep shows it grows to 6×\sim6\times as binaries come to dominate. H3 is confirmed too: the binary-aware allocation is not a monotone rescaling of the blind one — it reshuffles, pulling budget toward the radii that constrain fbinf_{\rm bin} to break the MfbinM\leftrightarrow f_{\rm bin} degeneracy (the blind design spread over {1.5,7.6,17}\{1.5, 7.6, 17\} pc; the aware design drops 7.6 pc and concentrates 2.26 pc while keeping 17 pc).

The honest test of identifiability: prior-insensitivity, not truth-centred recovery

This is the caveat to read before quoting the f^bin=0.50±0.08\hat f_{\rm bin}=0.50\pm0.08 recovery as proof that binaries are measurable. The cross-model mock is generated at the truth fbin=0.5f_{\rm bin}=0.5, and the binary-aware fit’s prior is centred on that same value. So an unbiased recovery, by itself, demonstrates only self-consistency — the fit returns the value it was pointed at. That is necessary but not sufficient.

What actually proves fbinf_{\rm bin} is identifiable from the data is a prior-insensitivity test: loosen the fbinf_{\rm bin} prior by 100×100\times and the recovered f^bin\hat f_{\rm bin} moves by only 0.4%0.4\%. The posterior is set by the radial leverage in the data, not by the prior — which is the genuine evidence that the core↔outskirts contrast measures the binary fraction. We state this plainly because it is the difference between a real measurement and a tautology, and the truth-centred recovery number alone cannot tell them apart.

Robustness, honestly: maximin is a near-null

The measure-and-marginalize fix assumes you trust the Moe & Di Stefano fbinf_{\rm bin} enough to model it. What if you do not — what if you want a design that is good across a whole range of possible binary fractions you refuse to commit to? That is maximin (robust) design: minimise the worst-case σ(M)\sigma(M) over fbin[0,0.7]f_{\rm bin}\in[0, 0.7].

The honest answer here is a near-null, and reporting it as such is part of the integrity of the arc. Because σ(M)\sigma(M) is monotone in fbinf_{\rm bin} (more binaries \Rightarrow a larger pedestal \Rightarrow less mass information), the worst case always sits at the upper endpoint fbin=0.7f_{\rm bin}=0.7 — and the marginalize design, already built to handle a free fbinf_{\rm bin}, is already nearly maximin-optimal. The maximin design gives up only +0.16%+0.16\% of precision at the truth to gain +0.24%+0.24\% at the worst case: a 0.2%\sim0.2\% hedge. The lesson is not “maximin wins” — it is that measure-and-marginalize is the right tool here, and the expensive worst-case machinery buys almost nothing on top of it. A null reported faithfully is a finding.

The robustness hedge is nearly free — and nearly pointless. Forecast \sigma(M)/M vs the
assumed binary fraction f_{\rm bin} for the marginalize design (blue, optimal at the truth
f_{\rm bin}=0.5) and the maximin design (vermilion, optimal at the worst case
f_{\rm bin}=0.7). The two curves nearly coincide: \sigma(M) is monotone in f_{\rm bin},
so the marginalize design is already near-maximin-robust. The callout quantifies the \sim0.2\%
hedge; the inset shows the two per-bin allocations barely differ. An honest near-null.

The robustness hedge is nearly free — and nearly pointless. Forecast σ(M)/M\sigma(M)/M vs the assumed binary fraction fbinf_{\rm bin} for the marginalize design (blue, optimal at the truth fbin=0.5f_{\rm bin}=0.5) and the maximin design (vermilion, optimal at the worst case fbin=0.7f_{\rm bin}=0.7). The two curves nearly coincide: σ(M)\sigma(M) is monotone in fbinf_{\rm bin}, so the marginalize design is already near-maximin-robust. The callout quantifies the 0.2%\sim0.2\% hedge; the inset shows the two per-bin allocations barely differ. An honest near-null.

The sweep: the bias is catastrophic where binaries dominate

The fiducial operating point (σbin/σcluster=1.08\sigma_{\rm bin}/\sigma_{\rm cluster}=1.08) is one slice through a continuum. Sweeping the system mass — which sets σcluster\sigma_{\rm cluster} while σbin\sigma_{\rm bin} stays fixed — traces the bias across contamination regimes, and it is the regime-of-validity map that turns a single number into a story.

The binary-blind-fit bias runs from +26%+26\% (massive, hot clusters where binaries are a minor perturbation) to +850%+850\% (low-mass, cold clusters where the pedestal swamps the cluster heat). Across the entire sweep the binary-aware fix holds the residual bias at 0\approx0. And the H2 precision-gain grows from 1.3×\sim1.3\times at the fiducial point to 6×\sim6\times where binaries dominate — so the thin operating-point margin is the floor, not the typical case. The colder the cluster, the more the binary-blind design lies, and the more the binary-aware design is worth.

Bias, fix, and OED payoff across the contamination ratio. Left axis vs
\sigma_{\rm bin}/\sigma_{\rm cluster} (top axis: system mass): the binary-blind-fit realized
M-bias (vermilion) rises from +26\% to +850\% as binaries come to dominate the cold
outskirts of lower-mass clusters; the binary-aware-fit residual (green) stays flat at \approx0
— the fix holds everywhere. Right axis: the H2 precision-gain (sky) grows from \sim1.3\times
to \sim6\times. The fiducial operating point (the single-point +184\% headline) is marked.

Bias, fix, and OED payoff across the contamination ratio. Left axis vs σbin/σcluster\sigma_{\rm bin}/\sigma_{\rm cluster} (top axis: system mass): the binary-blind-fit realized MM-bias (vermilion) rises from +26%+26\% to +850%+850\% as binaries come to dominate the cold outskirts of lower-mass clusters; the binary-aware-fit residual (green) stays flat at 0\approx0 — the fix holds everywhere. Right axis: the H2 precision-gain (sky) grows from 1.3×\sim1.3\times to 6×\sim6\times. The fiducial operating point (the single-point +184%+184\% headline) is marked.

Validating a pre-data calculation: two gates

A Fisher forecast is a promise — and this demo’s promise is the unusual one that a design will be biased. Both halves are checked.

Gate 1 — the gradients are correct (AD-vs-FD)

The load-bearing new derivatives are the fbinf_{\rm bin} Fisher block and the binary σ2\sigma^2-inflation term. We verify them against finite differences to rel<103\mathrm{rel}<10^{-3} (Richardson where a fixed-step FD is truncation-limited), per gradient-validation. The analytic fbinf_{\rm bin} column fbinVbin/(2σobs)f_{\rm bin}V_{\rm bin}/(2\sigma_{\rm obs}) matches AD to 1016\sim10^{-16}.

Gate 2 — the bias is real (cross-model Monte Carlo)

The bias claim is gated by a cross-model Monte Carlo: generate WITH Moe binaries, fit WITHOUT. The mock is forward-model-consistent — the cluster mock is the fit model, so the fbin=0f_{\rm bin}=0 baseline is unbiased by construction and any bias is attributable to binaries alone. Each draw places the design’s per-bin star counts directly, draws each star’s cluster velocity from N ⁣(0,σlos2(R))\mathcal{N}\!\left(0,\sigma_{\rm los}^2(R)\right) with σlos\sigma_{\rm los} from project_dispersion at the truth, adds a flux-weighted Moe blend Δ\Delta (from a build-once KorbK_{\rm orb} pool rescaled to Var=Vbin\mathrm{Var}=V_{\rm bin}) to the per-star Bernoulli(fbin)\mathrm{Bernoulli}(f_{\rm bin}) fraction, adds per-star εRV\varepsilon_{\rm RV}, forms the per-bin σ^\hat\sigma, and fits M^\hat M by MAP in the lnθ\ln\theta Gauss–Newton metric (honest realized-σ^\hat\sigma weighting; unpopulated bins dropped, not floored). The H1 accept rule was fixed in advance: accept iff the mean bias(M^)/M>2σforecast/M\mathrm{bias}(\hat M)/M > 2\,\sigma_{\rm forecast}/M and the 2SEM2\,\text{SEM} band does not straddle that threshold. It is met by a wide margin (bias/forecast =41×=41\times; SEM 0.06). The binary-aware fit’s accept rule — bias<2σM,marg|\mathrm{bias}| < 2\,\sigma_{M,\rm marg} — is met by the +5%+5\% residual.

The pre-registration (locked before the Monte Carlo)

To keep the bias result honest, the hypotheses and accept/reject rules were locked before the Monte Carlo ran, via research-workflow:discriminating-experiment-design. A reject on any of them was defined in advance as a reportable finding, not a failure — and an H1 reject would have descoped the whole arc.

Pre-registered hypotheses (locked 2026-06-19)

Hypothesis

Accept rule (fixed in advance)

Outcome

H1 — the bias. Binary-blind design + binary-blind fit on contaminated data biases M^\hat M high beyond its own forecast.

bias(M^)/M>2σforecast/M\mathrm{bias}(\hat M)/M > 2\,\sigma_{\rm forecast}/M, positive, 22\,SEM not straddling.

ACCEPT (+184%+184\% vs 4.5%4.5\%; 41×41\times)

H0 — the rival. The binary-blind fit absorbs the pedestal into nuisances; M^\hat M stays unbiased.

biasσforecast/M|\mathrm{bias}|\le\sigma_{\rm forecast}/M.

rejected (the fbin=0f_{\rm bin}=0 control is the unbiased case; fbin=0.5f_{\rm bin}=0.5 is not)

H2 — OED payoff. The binary-aware design beats the blind one under the binary-aware fit.

precision-gain 1.3×\geq1.3\times.

ACCEPT (1.33×1.33\times at the operating point; 6×\sim6\times in the sweep)

H3 — non-obvious allocation. The binary-aware allocation is not a monotone rescaling of the blind one.

per-bin weight rank-order changes.

ACCEPT (drops 7.6 pc, concentrates 2.26 pc)

Quantitative results

All numbers are from the gated CLI’s run-record (EFF-OM YMC operating point, Ntotal=5000N_{\rm total}=5000, K=12K=12 bins, cross-model MC 48 draws):

OED binary-robustness results

Quantity

Binary-blind (naive)

Binary-aware (fix)

M^/M\hat M/M (cross-model MC, fbin=0.5f_{\rm bin}=0.5)

2.84\mathbf{2.84} (+184%+184\% bias)

1.05\approx1.05 (+5%+5\%, within forecast)

Claimed / honest σ(M)/M\sigma(M)/M

4.5%\mathbf{4.5\%} (claimed)

6.9%\mathbf{6.9\%} (honest, marginalized)

bias / forecast ratio

41×\mathbf{41\times} (false confidence)

<2<2 (calibrated)

fbin=0f_{\rm bin}=0 control bias

0.3%-0.3\% (unbiased — isolates the binaries)

Recovered f^bin\hat f_{\rm bin}

— (not modelled)

0.50±0.08\mathbf{0.50\pm0.08} (radial leverage)

Prior-insensitivity (100×100\times looser prior)

f^bin\hat f_{\rm bin} moves 0.4%0.4\% (identifiable)

H2 precision-gain (binary-aware design)

1.33×\mathbf{1.33\times} (fiducial); 6×\sim6\times (binary-dominated)

Maximin hedge over the marginalize design

0.2%\sim0.2\% (honest near-null)

Sweep: MM-bias range across system mass

+26%+850%+26\% \to +850\%

0\approx0 throughout

What the optimum means: science implications

1. “Optimal” is only optimal under a correct model. The whole demo is a warning about the silent assumption inside every OED (and every analysis): the design that minimises your forecast variance under the wrong model can maximise your real error. A tighter forecast is not a better measurement — it is only a better measurement if the model is right. The binary-blind design’s 4.5%4.5\% was the most dangerous kind of number.

2. Population systematics are a design problem, not just an analysis problem. You cannot fix binaries purely after the fact if your survey spent its budget in the contaminated outskirts. The binary-aware design spends differently because it knows the pedestal exists — it earns the fbinf_{\rm bin} constraint from radial leverage at the telescope, not from a prior at the keyboard.

3. progenax can do this because it is differentiable end-to-end. Wiring a faithful Moe & Di Stefano binary population into a Fisher design requires (information)/(observing strategy)\partial(\text{information})/\partial(\text{observing strategy}) through the binary kernel. That is the throughline of the whole OED section: the same additive backbone that targets anisotropy, mass, and concentration here defends a measurement against a model error — it extends OED from “what to measure” to “what to measure robustly.”

Current scope and planned extensions

How to run

# the cheap design + forecast + mechanism figure (no MC; CI-safe, ~1 min)
env -u VIRTUAL_ENV uv run --no-sync python scripts/demo_oed_binary.py

# the full result: cross-model MC (the H1 bias + headline figure) + maximin + sweep
PROGENAX_RUN_OED_BINARY=1 env -u VIRTUAL_ENV uv run --no-sync \
    python scripts/demo_oed_binary.py --run-mc --maximin --sweep \
    --outdir docs/website/60-science-demos/optimal-design/figures

The cheap default path computes the binary-blind design, its forecast σ(M)/M\sigma(M)/M, and the mechanism figure, and exits 0. The env-gated cross-model Monte Carlo (the H1 bias, the false-confidence figure) is @slow and out of CI — it is the gate test_H1_naive_design_is_biased_beyond_forecast (and the binary-aware test_fix_binary_aware_fit_is_unbiased) in tests/unit/test_demo_oed_binary.py, run with PROGENAX_RUN_OED_BINARY=1. The figures and run-record land alongside this page.

References

The shared Fisher / Cramér–Rao / projection theory and its references are on the OED formalism page. The line-of-sight projection into σlos\sigma_{\rm los} is Binney & Mamon (1982); the Osipkov–Merritt anisotropy law is Merritt (1985); the EFF density profile is documented on the EFF profile theory page. The binary population — the PPqqee engine behind VbinV_{\rm bin} — follows Moe & Di Stefano (2017), with ZAMS flux weights from Tout et al. (1996); the single-epoch binary-inflated dispersion machinery is the binary dynamical-mass demo (B12). The companion OED designs are anisotropy (where the proper motions belong), dynamical mass (how deep to survey), and concentration (where concentration information lives) — this page extends them to model robustness.

References
  1. Binney, J., & Mamon, G. A. (1982). M/L and velocity anisotropy from observations of spherical galaxies, or must M87 have a massive black hole? Monthly Notices of the Royal Astronomical Society, 200, 361–375. 10.1093/mnras/200.2.361
  2. Merritt, D. (1985). Spherical stellar systems with spheroidal velocity distributions. The Astronomical Journal, 90, 1027–1037. 10.1086/113810
  3. Moe, M., & Di Stefano, R. (2017). Mind your Ps and Qs: The interrelation between period (P) and mass-ratio (Q) distributions of binary stars. The Astrophysical Journal Supplement Series, 230, 15. 10.3847/1538-4365/aa6fb6
  4. Tout, C. A., Pols, O. R., Eggleton, P. P., & Han, Z. (1996). Zero-age main-sequence radii and luminosities as analytic functions of mass and metallicity. Monthly Notices of the Royal Astronomical Society, 281, 257. 10.1093/mnras/281.1.257