Tests, results and
scientific evaluation
Captured repository runs, individual assertions, analysis scripts, historical failures and data-model comparisons are kept separate. Every green marker states what it establishes—and what it does not.
Twelve-repository runner snapshot
The source snapshot records Python 3.12 on Windows 11. It is one reproducible executed subset, not the current size of the complete 37-repository local corpus.
| Repository | Passed | Failed | Expected | Completion | Evidence boundary |
|---|---|---|---|---|---|
| Loading evaluated run… | |||||
How the runner moved from conflict to an all-green capture
A credible dashboard preserves failures and counting changes. The records below are not three measurements of one total: they use different repository sets, orchestration paths and counting units.
Snapshot chronology
Intermediate audit inventory
Counting result: the runner exceeded its minimum, but still recorded validation failures and one timeout. Quantity coverage and successful execution are separate axes.
Real failures retained from the audit
Timeouts and incomplete execution
A timeout is not silently recoded as PASS or FAIL. It means the runner did not obtain a completed assertion result within its execution budget.
Selected numerical diagnostics from captured stdout
These values make the green markers inspectable. The chart plots each observed error or drift as a fraction of its encoded tolerance; the accompanying table preserves the scientific limitation.
Fraction of tolerance consumed
A shorter bar means more numerical margin. It does not mean stronger empirical evidence.
Evidence ladder
- Code identityDoes the implementation satisfy an algebraic or API invariant?
- Numerical convergenceDo methods agree and errors remain within tolerance?
- Reference compatibilityDoes the model reproduce a selected known limit or dataset?
- Discriminating evidenceDoes a preregistered comparison distinguish SSZ from alternatives?
- Independent replicationCan an external team reproduce the result with independent data and code?
| Diagnostic | Observed | Tolerance | Repository / source | What it supports | What it does not establish |
|---|
Claim-to-evidence matrix
Every central statement is classified before it is displayed. “Tested” refers to encoded mathematical or numerical support; “conditional” depends on a supplied sample; “corrected” is explicitly superseded; “open” lacks decisive evidence.
| Claim | Class | Status | Current support | Explicit boundary |
|---|
What 9,300 records and 5,294 unique definitions cover
Every record was checked for required fields, category, repository, source and reproduction command. The audit finds 2,280 source files, 5,294 unique repository/test definitions and 9,216 unique repository/file/test identities. These are coverage counts; only dated runner artifacts establish pass/fail outcomes.
Artefacts by test category
Artefacts by scientific quantity
Repository coverage
| Repository | Catalogue rows | Share |
|---|
How to interpret coverage
Many discoverable files or extracted test functions. It may also indicate generated reports or duplicated orchestration artefacts.
A separate dated runner outcome with a declared environment and unit.
Requires an equation, input provenance, uncertainty, result and an explicit claim boundary.
Observed concentration: unit/integration artefacts dominate the catalogue. This does not mean observational, symbolic, regression and boundary validation are equally covered.
What the 67-pair comparison actually reports
The supplied enriched pipeline compares median residuals for segmented, GR, SR and GR+SR model forms. The display preserves the dataset-conditioned interpretation and its limitations.
Model medians
Interpretation
Within this particular sample and residual definition, the segmented form has the smaller residual in 66 of 67 pairs.
Bootstrap interval comparison
Intervals are displayed on a logarithmic axis because the reported residual scales differ by orders of magnitude.
Mass-bin support and residuals
Empty and sparse bins remain visible. The high-mass single-object and three-object bins must not be interpreted like the densely populated bins.
Inspect the mass-bin evaluation
| Bin | log₁₀ mass interval | N | Segmented median | GR median | GR+SR median |
|---|
Three intermediate bins contain no observations and the two highest-mass bins contain only one and three objects. The aggregate significance must not conceal that sparse structure.
Executed analyses, tools and tests by category
A successful script can be a test, smoke check, exporter or analysis. The execution unit is therefore reported explicitly.
| Category | Passed | Failed | Skipped | Duration | Counting unit |
|---|
Reproducibility value: these records show which supplied programs completed. Scientific interpretation still depends on the equations, inputs, uncertainty model and comparison protocol inside each program.
Earlier orchestration snapshot
This older snapshot contains missing dependencies, collection failures and inconsistent expected counts—including a 203.7% rate. It is retained as engineering history, not merged into the later green total.
| Repository | Passed | Failed | Expected | Recorded rate | Interpretation |
|---|
Search every extracted test or result artefact
This lower-level catalogue supports source discovery. Its rows are artefacts, not an additive pass total.
| Test / repository | Category | Quantity | Source file | Meaning and boundary | Reproduction |
|---|---|---|---|---|---|
| Loading static catalogue… | |||||
Historical green output can still contain a superseded interpretation
“Finite curvature everywhere”
The current diagonal continuation has \(R\sim3/(2r^2)\) and \(K\sim9/(4r^4)\) at the areal centre. Passing transition tests cannot establish central regularity.
“Tests confirm the theory”
Tests confirm encoded assertions. Empirical preference additionally requires measured data, uncertainty, alternatives, preregistered criteria and independent replication.
Preserve the counting unit
Pytest cases, parameterisations, scripts, runner phases, skips, xfails and reports remain separately labelled.