Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Jaxstro package assessment scorecard

This is the living assessment of Jaxstro as scientific infrastructure, a public package, and curriculum material. The companion Package and documentation SOTA assessment ranks investments; this page records current grades and what evidence is required to change them.

Assessment date: 2026-07-12; registry reconciliation: 2026-07-12. Grades describe the repository evidence reviewed on that date and must be re-audited when their supporting contracts change.

Grading rubric

Pluses and minuses identify position within a band. Grades describe current evidence, not effort or ambition.

Current grades

DimensionGradeEvidenceEvidence-backed rationaleDeficiency preventing the next gradePromotion evidence required
Internal differentiable-science foundationA−ValidationBroad shared primitives, explicit ownership, and extensive executable validationEvidence depth and downstream verification are not uniformGenerated contracts, uniform evidence classes, and pinned adoption records
Public scientific packageB+API referenceCohesive surface, minimal core dependencies, typed failures, and a public sitePublic reach and external adoption trail the implementationStable release, public compatibility evidence, and external use cases
Curriculum conceptB+Foundations: the ideas we will not assumeAn expert-reviewed first-principles spine connects models, dimensionality, linear algebra, derivatives, probability, inference, conditioning, and programs through predict → compute → auditLearner comprehension and transfer have not yet been studiedLearner evidence, misconception diagnostics, and concept-linked figures
Executable curriculumB+Executable research investigationsThree CI-executed investigations resolve public APIs to contracts, evidence, limitations, instructor notes, and a claim-calibrated rubricUnit breadth, automated feedback, optional exports, and learner trials remain limitedBroader investigation coverage and evidence from real course use
Architecture and ownershipAScientific contract registryThin-foundation ADRs, one-way ecosystem boundaries, and generated module contracts are explicitEvidence depth remains uneven at callable levelExpand callable classification without weakening the thin-foundation boundary
Numerical correctnessA−ValidationAnalytic cases, limits, round trips, convergence, and FD comparisons cover major kernelsValidation and performance artifacts remain uneven by modulePer-contract evidence coverage and missing-evidence ratchets
AD honestyARoot-findingSmooth, blocked, zero, value-first, validation-only, and certified implicit paths are distinguishedThe taxonomy is distributed across prose and testsGenerated callable-level AD contract matrix
JAX architectureA−Writing AD-safe scientific numericsJIT, VMAP, scan, PyTree, JVP/VJP, and gradient behavior are tested on substantial surfacesTransform and batching-cost behavior is hard to discover per callableGenerated transform matrix linked to executable evidence
Failure semanticsARoot-findingRoot, atmosphere, spatial, and spectral APIs expose typed or structured failure evidenceSome numerical helpers still return less-auditable scalar stateCallable-level boundary and failure records with coverage ratchets
Test architectureA−ValidationUnit, integration, and validation tiers contain analytic and independent checksNo public module-by-module branch-coverage ratchet is maintainedCoverage policy tied to semantic contracts rather than raw volume alone
Units and dimensional safetyB+Quantity system architectureUnitSystem is mature and quantity is substantially implementedTwo live dimensional surfaces remain while adoption is deferredDownstream parity, performance, serialization, and migration evidence
ProvenanceAScientific evidence indexSource cards, runtime manifests, full-card content digests, and a class-preserving evidence index are freshness checkedDownstream reproduction policies and adoption records are not yet uniformPinned downstream evidence manifests and compatibility records
API cohesionB+Scientific contract registryPublic modules and selected consequential callables now have generated, export-audited contractsMany public callables remain explicitly unclassifiedPrioritized callable-level coverage and downstream query evidence
Documentation correctnessA−ValidationExamples, routes, links, figures, and content contracts are testedActive guides and some command narratives can lag implementationCurrent CLAUDE.md plus generated claims and freshness gates
AccessibilityB+Root-findingAlt text and redundant visual encodings are enforced for new figuresStructural checks do not yet constitute learner-centered accessibility evidenceKeyboard, contrast, comprehension, and learner review gates
Performance evidenceB+Scientific evidence indexRootfinding and spectra use one units-explicit metric/comparison envelope with deterministic freshness checksCompile, graph-size, memory, and cost coverage remains uneven by moduleMethod-appropriate performance records for consequential public contracts
Downstream usefulnessB+Astro-first, science-general visionActive sibling packages motivate and consume selected shared foundationsAdoption claims are not generated from pinned downstream revisionsSymbol-to-project compatibility records and consumer validation
External reachBjaxstroThe package is science-general, permissively licensed, dependency-light, and publicly documentedStable distribution and independent adoption are limitedRelease evidence, external examples, and comparative positioning
Maintenance readinessB+Release checklistRelease gates, deterministic docs, Ruff, MyPy, and wheel smoke existType strictness and manually duplicated truth remain unevenContract-driven docs, stricter typing plan, and unified evidence infrastructure

Coverage by scientific area

Numerics — A−

Rootfinding, interpolation, integration, quadrature, distributions, splines, grids, optimization, ODEs, operators, and linear algebra have meaningful mathematical tests. Rootfinding is the current exemplar for separating a robust value, execution telemetry, finite-map sensitivity, and a certified implicit derivative. Other modules do not yet expose equally rich contracts and artifacts.

Coordinates and geometry — B+

Coordinate transformations, singularity documentation, and gradient checks are useful and science-facing. Frame conventions and geometric degeneracies need stronger visual and contract-level presentation.

Spatial methods — B+

Approximate candidate generation and exact pair acceptance are correctly separated, with explicit capacity and overflow semantics. Scaling, memory, and adversarial-configuration evidence remain limited.

Spectra and atmospheres — A−

Source semantics, structured outcomes, interpolation-policy evidence, and host versus JAX boundaries are unusually explicit. Some validation depends on optional local artifacts, so reproducibility boundaries must remain visible.

Units and quantities — B+

The canonical unit-system surface is mature and the quantity layer is deeply tested. Ecosystem adoption remains an evidence question, not an automatic migration.

Parameters and inference bridge — B

The selective bridge is appropriately narrow and avoids becoming an inference framework. More examples are needed around identifiability, constraint geometry, and cached derived leaves.

Testing and provenance — A

Gradient audits, numerical ratchets, source cards, and deterministic reports are connected by one scientific contract registry and a class-preserving evidence index. Evidence depth remains uneven across public callables, and source evidence still must not be mistaken for computational validation.

Grade-change policy

A grade does not improve merely because a feature or documentation claim lands. Promotion requires the relevant scientific contract, independent validation, limitation statement, reproducible artifact where metrics matter, and downstream adoption evidence where reuse is the justification.

Every change to a grade must record the date, supporting evidence, and the criterion newly satisfied. A regression in evidence or a newly discovered defect can lower a grade immediately.

Current hardening sequence

Scientific contract registry: implemented. The registry is deliberately not evidence-complete. Evidence depth remains uneven, and an unclassified symbol is not treated as supported.

Metric identitySymbolValueUnits
Registered public modulesN_module,contract16modules
Callable-level contractsN_callable,contract15callables
Explicitly unclassified public callablesN_callable,unclassified208callables
Module-inherited public record typesN_symbol,inherited125symbols

The completed foundation and next investment are:

Unified evidence infrastructure: implemented. Computational measurements, source provenance, and scientific policy remain separate evidence classes in one freshness-checked index. Method-specific scientific thresholds remain method-owned; the shared envelope validates identity, units, comparison truth, environment policy, and serialization without inventing cross-method acceptance rules.

Metric identitySymbolValueUnits
Indexed evidence artifactsN_artifact,evidence5artifacts
Distinct evidence classesN_class,evidence3evidence classes

Executable foundations curriculum: implemented. Optional readiness routing, first-principles foundations, executable research investigations, and instructor assessment now coexist with the unchanged module and API reference sections. This delivery does not automatically raise every pedagogy grade: learner testing, visual coverage, automated feedback, optional exports, and investigation breadth remain explicit promotion gates.

Metric identitySymbolValueUnits
Executable research investigationsN_unit,curriculum3investigations
Callable contract linksN_contract,curriculum7contract links
Indexed artifact linksN_evidence,curriculum2indexed evidence links
Units with instructor routesN_instructor,curriculum3instructor routes

The next curriculum investment is a contract-derived pedagogical figure suite plus learner-centered accessibility and comprehension evidence. Package-wide callable coverage and downstream adoption remain separate roadmap priorities.