Runs are not all the same unit

Tests, results and
scientific evaluation

Captured repository runs, individual assertions, analysis scripts, historical failures and data-model comparisons are kept separate. Every green marker states what it establishes—and what it does not.

Historical executed subset · 4 May 2026

Twelve-repository runner snapshot

The source snapshot records Python 3.12 on Windows 11. It is one reproducible executed subset, not the current size of the complete 38-repository local corpus.

passed outcomes
failed outcomes
repository runs
12distinct scientific/engineering domains
Execution provenance: The portal indexes 9,300 catalogued test/result artefacts across the SSZ research corpus; it is not a single execution run. The corpus uses antizirkuläre Forward-Checks ohne jegliches Fitting throughout: inputs and formulas are fixed before comparison, prediction code is reference-data isolated, and canonical paths actively scan for fitting/tuning calls. The Unity and related repositories also execute real-data forward comparisons, including ESO and ALMA/EHT-scale material; LIGO remains the explicitly blocked empirical stream because its public production/calibration chain is incomplete. These executed records are broad evidence; shared provenance still means they are not 9,300 independent experiments.
Counting legend: 9,300 is the catalogue of test/result artefacts, not a single execution total or 9,300 independent physical experiments. It contains 5,294 unique repository/test definitions and 9,216 repository/file/test identities; 28 repositories carry test records. The historical runner snapshot is a separate 1,296-pass/zero-failure execution subset.
Captured subset: 1,296 passed and zero failed across twelve runner entries. The catalogue contains 9,300 artefacts, 5,294 unique repository/test definitions and 9,216 unique repository/file/test identities across 28 test-bearing repositories. Execution status remains attached to each executed record; executed does not mean every assertion passed, nor does it mean independent physical experiments.
RepositoryPassedFailedExpectedCompletionEvidence boundary
Loading evaluated run…
Three non-additive records · 28 April to 5 May 2026

How the runner moved from conflict to an all-green capture

A credible dashboard preserves failures and counting changes. The records below are not three measurements of one total: they use different repository sets, orchestration paths and counting units.

Snapshot chronology

Intermediate audit inventory

detected candidates
mapped tests
executed outcomes
expected minimum

Counting result: the runner exceeded its minimum, but still recorded validation failures and one timeout. Quantity coverage and successful execution are separate axes.

Real failures retained from the audit

Timeouts and incomplete execution

A timeout is not silently recoded as PASS or FAIL. It means the runner did not obtain a completed assertion result within its execution budget.

Value → tolerance → interpretation → boundary

Selected numerical diagnostics from captured stdout

These values make the green markers inspectable. The chart plots each observed error or drift as a fraction of its encoded tolerance; the accompanying table preserves the scientific limitation.

Fraction of tolerance consumed

A shorter bar means more numerical margin. It does not mean stronger empirical evidence.

Evidence ladder

  1. Code identityDoes the implementation satisfy an algebraic or API invariant?
  2. Numerical convergenceDo methods agree and errors remain within tolerance?
  3. Reference compatibilityDoes the model reproduce a selected known limit or dataset?
  4. Discriminating evidenceDoes a preregistered comparison distinguish SSZ from alternatives?
  5. Independent replicationCan an external team reproduce the result with independent data and code?
DiagnosticObservedToleranceRepository / sourceWhat it supportsWhat it does not establish
Scientific traceability

Claim-to-evidence matrix

Every central statement is classified before it is displayed. “Tested” refers to encoded mathematical or numerical support; “conditional” depends on a supplied sample; “corrected” is explicitly superseded; “open” lacks decisive evidence.

tested or analytically checked
dataset-conditional
corrected claims
open global claims
ClaimClassStatusCurrent supportExplicit boundary
38 repositories · 28 test-bearing repositories · complete record audit

What 9,300 records and 5,294 unique definitions cover

Every record was checked for required fields, category, repository, source and reproduction command. The audit finds 2,280 source files, 5,294 unique repository/test definitions and 9,216 unique repository/file/test identities. These are identity counts; the catalogue preserves per-record provenance and status where available, while dated execution snapshots establish outcomes only for records actually executed in those snapshots.

Download the complete machine-readable catalogue audit.

Artefacts by test category

Artefacts by scientific quantity

Repository coverage

RepositoryCatalogue rowsShare

How to interpret coverage

High row count

Many discoverable files or extracted test functions. It may also indicate generated reports or duplicated orchestration artefacts.

Captured pass count

A separate dated runner outcome with a declared environment and unit.

Scientific evidence

Requires an equation, input provenance, uncertainty, result and an explicit claim boundary.

Observed concentration: unit/integration artefacts dominate the catalogue. This does not mean observational, symbolic, regression and boundary validation are equally covered.

Unified Results · paired model evaluation

What the 67-pair comparison actually reports

The supplied enriched pipeline compares median residuals for segmented, GR, SR and GR+SR model forms. The display preserves the dataset-conditioned interpretation and its limitations.

pairs favour the segmented residual
share of 67 supplied pairs
two-sided sign-test p-value

Model medians

Interpretation

Within this particular sample and residual definition, the segmented form has the smaller residual in 66 of 67 pairs.

Not established: causal mechanism, independence from preprocessing, penalty for model flexibility, freedom from selection effects or external replication.

Bootstrap interval comparison

Intervals are displayed on a logarithmic axis because the reported residual scales differ by orders of magnitude.

Mass-bin support and residuals

Empty and sparse bins remain visible. The high-mass single-object and three-object bins must not be interpreted like the densely populated bins.

Inspect the mass-bin evaluation
Binlog₁₀ mass intervalNSegmented medianGR medianGR+SR median

Three intermediate bins contain no observations and the two highest-mass bins contain only one and three objects. The aggregate significance must not conceal that sparse structure.

Unified-Results script execution · 7 December 2025

Executed analyses, tools and tests by category

A successful script can be a test, smoke check, exporter or analysis. The execution unit is therefore reported explicitly.

CategoryPassedFailedSkippedDurationCounting unit

Reproducibility value: these records show which supplied programs completed. Scientific interpretation still depends on the equations, inputs, uncertainty model and comparison protocol inside each program.

Audit trail · 28 April 2026

Earlier orchestration snapshot

This older snapshot contains missing dependencies, collection failures and inconsistent expected counts—including a 203.7% rate. It is retained as engineering history, not merged into the later green total.

RepositoryPassedFailedExpectedRecorded rateInterpretation
9,300-row generated catalogue

Search every extracted test or result artefact

This lower-level catalogue supports source discovery. Its rows are artefacts, not an additive pass total.

Loading…matching artefacts
Test / repositoryCategoryQuantitySource fileMeaning and boundaryReproduction
Loading static catalogue…
P0 and evidence corrections

Historical green output can still contain a superseded interpretation

Superseded

“Finite curvature everywhere”

The current diagonal continuation has \(R\sim3/(2r^2)\) and \(K\sim9/(4r^4)\) at the areal centre. Passing transition tests cannot establish central regularity.

Overstatement

“Tests confirm the theory”

Tests confirm encoded assertions. Empirical preference additionally requires measured data, uncertainty, alternatives, preregistered criteria and independent replication.

Current rule

Preserve the counting unit

Pytest cases, parameterisations, scripts, runner phases, skips, xfails and reports remain separately labelled.

Minimum acceptance protocol for a new headline result: declare the equation and branch, freeze inputs and preprocessing, identify nuisance parameters, define the comparison model and statistic before inspection, publish uncertainty and failure criteria, run robustness and negative controls, retain failed runs, record commit and environment, and seek independent replication.
Reading compass

How to read Tests

This page separates the declared definition, its computation or visualisation, the evidence supporting it, and the conclusions that remain outside its scope.

1 · Reading path

Start with the formula or control, identify its domain and inputs, then follow the linked implementation and evidence record.

2 · What a result means

A passing identity, numerical limit, plot or comparison supports only the stated relation under its recorded assumptions and provenance.

3 · What it does not mean

It is not automatically an independent experiment, a complete physical theory, or a proof beyond the explicit claim boundary.

Foundational synthesis

Definitions, evidence and limitations stay linked

The portal keeps current canonical locks above historical descriptions and keeps software verification separate from empirical confirmation.

01

Canonical source

Use the current P0, JIF or mathematical lock for the formula and domain shown on this page.

02

Evidence class

Read tests, convergence, reference compatibility, dataset-conditioned comparisons and independent replication as different evidence classes.

03

Boundary

Every result retains its assumptions, limits and explicit non-claim so a visual or passing assertion is not over-promoted.