Reproducible gravitational-wave data audit

Spin Evidence Ledger

I built a bounded reproduction and robustness audit of the public parameter-estimation samples for GW241011_233834 and GW241110_124123.

Overall statusBlocked where matching prior samples are missing
PythonNumPySciPyHDF5Matplotlibpytest
Probability-density comparisons for effective spin in GW241011 and GW241110 across three waveform analysesOpen the full comparison figure ↗
2gravitational-wave events
48derived-quantity comparisons
42pairwise posterior comparisons
0sample checks outside tolerance

The question

How stable are released spin results across analyses?

01

Can I reproduce the released values?

I recalculated mass ratio, symmetric mass ratio, chirp mass, source-frame chirp mass, and effective spin from the released samples. The checks keep mass frames and spin epochs separate.

02

How much do the posterior results differ?

I compared the one-dimensional posterior distributions from three waveform analyses. A posterior is the probability distribution for a parameter after the observed signal is included.

03

How much information came from the data?

That calculation requires each posterior’s matching sampled prior, which represents the assumptions in place before the signal was considered. Those samples were not present in the release.

The process

An evidence trail from source file to final figure.

  1. 01

    Pin and inspect the release

    The workflow checks the published archive size and MD5, records a local SHA-256, rejects unsafe archive members, and extracts only the standard parameter-estimation material.

  2. 02

    Inventory the HDF5 samples

    I preserved the exact analysis labels, parameter names, weights, sample counts, effective sample sizes, and missing fields before running numerical work.

  3. 03

    Recompute and compare

    Derived quantities were rebuilt sample by sample and checked against preset tolerances. Rounded summaries were also compared with the paper and GWOSC values.

  4. 04

    Test posterior separation

    Pairwise Jensen-Shannon divergence used shared histogram support, 20, 40, and 80-bin sensitivity checks, 500 seeded bootstrap replicates, and same-model split-sample checks.

Verified results

The same spin parameter behaved differently in the two events.

The horizontal axis is χeff,∞, a mass-weighted measure of the net spin component along or against the orbital direction when the system is evolved back to very large separation. The vertical axis is probability density, not evidence for one model.

GW241011_2338340.587667 bits

Largest audited marginal divergence, for IMRPhenomXO4a versus IMRPhenomXPHM. The blue posterior is shifted from the red and teal results.

90% bootstrap interval: 0.581258 to 0.595403 bits
GW241110_1241230.016942 bits

Corresponding descriptive maximum for the second event. The three posterior curves overlap much more closely.

90% bootstrap interval: 0.015594 to 0.019612 bits

These are one-dimensional posterior-analysis divergences, not significance scores. Missing prior comparisons mean the separation cannot be assigned to waveform physics alone.

Reproduction ledger

What reproduced, and what needed a note.

48 / 48

Derived comparisons within tolerance

No comparison contained a sample outside its preset tolerance. The largest sample-level absolute algebraic difference was 8.882 × 10−15.

169 / 176

Printed summary checks within rounding tolerance

The seven differences all involve the paper’s specialized HPD-in-cosine spin-tilt bounds. Ordinary central-interval quantities agree at their printed precision.

Incomplete stage0 KL estimates6 missing prior inputs

Why KL divergence is not reported

The released files do not include the priors needed for this calculation.

The six analysis groups do not include prior samples, and their configuration files do not fully define the priors in a machine-readable form. Because KL divergence compares a posterior with its prior, it cannot be reproduced from the available files.

These missing inputs only affect the prior-to-posterior information-gain and prior-to-prior comparisons. The 48 derived checks and 42 posterior comparisons above use the released posterior samples and remain complete.

Inspect the work

Results, methods, machine-readable metrics, and boundaries.

Generative AI assisted with repository scaffolding, implementation drafts, tests, and documentation. It was not used to estimate any scientific number. Read the full disclosure ↗

Project boundary

This is an audit, not an astrophysical classification.

The project does not validate either discovery, classify either event, infer a population or formation channel, test a hierarchical-merger or eccentricity hypothesis, or replace review of the full multidimensional posterior.

Return to the portfolio