Runs are not all the same unit

Tests, results and
scientific evaluation

Captured repository runs, individual assertions, analysis scripts, historical failures and data-model comparisons are kept separate. Every green marker states what it establishes—and what it does not.

Historical executed subset · 4 May 2026

Twelve-repository runner snapshot

The source snapshot records Python 3.12 on Windows 11. It is one reproducible executed subset, not the current size of the complete 37-repository local corpus.

passed outcomes
failed outcomes
repository runs
12distinct scientific/engineering domains
Captured subset: 1,296 passed and zero failed across twelve runner entries. The complete inventory contains 9,300 records, 5,294 unique repository/test definitions and 9,216 unique repository/file/test identities across 28 test-bearing repositories. Catalogue records are not silently converted into executed PASS outcomes.
RepositoryPassedFailedExpectedCompletionEvidence boundary
Loading evaluated run…
Three non-additive records · 28 April to 5 May 2026

How the runner moved from conflict to an all-green capture

A credible dashboard preserves failures and counting changes. The records below are not three measurements of one total: they use different repository sets, orchestration paths and counting units.

Snapshot chronology

Intermediate audit inventory

detected candidates
mapped tests
executed outcomes
expected minimum

Counting result: the runner exceeded its minimum, but still recorded validation failures and one timeout. Quantity coverage and successful execution are separate axes.

Real failures retained from the audit

Timeouts and incomplete execution

A timeout is not silently recoded as PASS or FAIL. It means the runner did not obtain a completed assertion result within its execution budget.

Value → tolerance → interpretation → boundary

Selected numerical diagnostics from captured stdout

These values make the green markers inspectable. The chart plots each observed error or drift as a fraction of its encoded tolerance; the accompanying table preserves the scientific limitation.

Fraction of tolerance consumed

A shorter bar means more numerical margin. It does not mean stronger empirical evidence.

Evidence ladder

  1. Code identityDoes the implementation satisfy an algebraic or API invariant?
  2. Numerical convergenceDo methods agree and errors remain within tolerance?
  3. Reference compatibilityDoes the model reproduce a selected known limit or dataset?
  4. Discriminating evidenceDoes a preregistered comparison distinguish SSZ from alternatives?
  5. Independent replicationCan an external team reproduce the result with independent data and code?
DiagnosticObservedToleranceRepository / sourceWhat it supportsWhat it does not establish
Scientific traceability

Claim-to-evidence matrix

Every central statement is classified before it is displayed. “Tested” refers to encoded mathematical or numerical support; “conditional” depends on a supplied sample; “corrected” is explicitly superseded; “open” lacks decisive evidence.

tested or analytically checked
dataset-conditional
corrected claims
open global claims
ClaimClassStatusCurrent supportExplicit boundary
37 repositories · 28 test-bearing repositories · complete record audit

What 9,300 records and 5,294 unique definitions cover

Every record was checked for required fields, category, repository, source and reproduction command. The audit finds 2,280 source files, 5,294 unique repository/test definitions and 9,216 unique repository/file/test identities. These are coverage counts; only dated runner artifacts establish pass/fail outcomes.

Download the complete machine-readable catalogue audit.

Artefacts by test category

Artefacts by scientific quantity

Repository coverage

RepositoryCatalogue rowsShare

How to interpret coverage

High row count

Many discoverable files or extracted test functions. It may also indicate generated reports or duplicated orchestration artefacts.

Captured pass count

A separate dated runner outcome with a declared environment and unit.

Scientific evidence

Requires an equation, input provenance, uncertainty, result and an explicit claim boundary.

Observed concentration: unit/integration artefacts dominate the catalogue. This does not mean observational, symbolic, regression and boundary validation are equally covered.

Unified Results · paired model evaluation

What the 67-pair comparison actually reports

The supplied enriched pipeline compares median residuals for segmented, GR, SR and GR+SR model forms. The display preserves the dataset-conditioned interpretation and its limitations.

pairs favour the segmented residual
share of 67 supplied pairs
two-sided sign-test p-value

Model medians

Interpretation

Within this particular sample and residual definition, the segmented form has the smaller residual in 66 of 67 pairs.

Not established: causal mechanism, independence from preprocessing, penalty for model flexibility, freedom from selection effects or external replication.

Bootstrap interval comparison

Intervals are displayed on a logarithmic axis because the reported residual scales differ by orders of magnitude.

Mass-bin support and residuals

Empty and sparse bins remain visible. The high-mass single-object and three-object bins must not be interpreted like the densely populated bins.

Inspect the mass-bin evaluation
Binlog₁₀ mass intervalNSegmented medianGR medianGR+SR median

Three intermediate bins contain no observations and the two highest-mass bins contain only one and three objects. The aggregate significance must not conceal that sparse structure.

Unified-Results script execution · 7 December 2025

Executed analyses, tools and tests by category

A successful script can be a test, smoke check, exporter or analysis. The execution unit is therefore reported explicitly.

CategoryPassedFailedSkippedDurationCounting unit

Reproducibility value: these records show which supplied programs completed. Scientific interpretation still depends on the equations, inputs, uncertainty model and comparison protocol inside each program.

Audit trail · 28 April 2026

Earlier orchestration snapshot

This older snapshot contains missing dependencies, collection failures and inconsistent expected counts—including a 203.7% rate. It is retained as engineering history, not merged into the later green total.

RepositoryPassedFailedExpectedRecorded rateInterpretation
9,300-row generated catalogue

Search every extracted test or result artefact

This lower-level catalogue supports source discovery. Its rows are artefacts, not an additive pass total.

Loading…matching artefacts
Test / repositoryCategoryQuantitySource fileMeaning and boundaryReproduction
Loading static catalogue…
P0 and evidence corrections

Historical green output can still contain a superseded interpretation

Superseded

“Finite curvature everywhere”

The current diagonal continuation has \(R\sim3/(2r^2)\) and \(K\sim9/(4r^4)\) at the areal centre. Passing transition tests cannot establish central regularity.

Overstatement

“Tests confirm the theory”

Tests confirm encoded assertions. Empirical preference additionally requires measured data, uncertainty, alternatives, preregistered criteria and independent replication.

Current rule

Preserve the counting unit

Pytest cases, parameterisations, scripts, runner phases, skips, xfails and reports remain separately labelled.

Minimum acceptance protocol for a new headline result: declare the equation and branch, freeze inputs and preprocessing, identify nuisance parameters, define the comparison model and statistic before inspection, publish uncertainty and failure criteria, run robustness and negative controls, retain failed runs, record commit and environment, and seek independent replication.