Core performance baseline¶
For current-core regression monitoring, use run_core.py and compare.py.
See the consolidated workflow and CI policy for quick/stress
suites, versioned workload matching, advisory reporting and baseline refresh.
The commands below preserve the original Phase 3A–3D measurement workflows.
Run from a clean checkout with Python 3.12 and the existing six dependency:
python benchmarks/core_baseline.py --output /tmp/gem-baseline.json --trials 7 --target-seconds .02 --profiles
Add --historical to compare Matrix3/Matrix4 inversion with Phase 2F-1 in the
same interpreter. This option requires the historical commit object
2cbd899a47546b8c40f8244f57799969c86b0d87; shallow checkouts/source archives
can run every current-core benchmark without it. No runtime package changes,
optional benchmark dependency, RNG seed or downloaded asset is required.
Inputs are constants in prepare(). Ordinary vectors range from -3.75 to 3.75;
matrices are nonsymmetric diagonally dominant, with entries
3*identity + .13*(i+1)/(j+1), and a separate nonsymmetric multiplication operand.
Extreme norms use 1e-300/1e300, mixed norms span 600 decimal orders, and raw
inverse cases use uniform 1e-200/1e200 scales. Historical known answers cover
1e-300 through 1e300 using a well-conditioned triangular matrix. Quaternions
represent unit rotations around nontrivial axes; interpolation uses t=.37,
including a .001-degree near-angle case. No singular/unsupported operation is
timed as successful mathematics.
Bezier controls are 3D, bounded by eight units in X, two in Y and .5 in Z; a single-segment distance-tolerance sweep covers .1, .01, 1e-4, 1e-8 and 1e-15. The core stores this as squared tolerance. Eight segments use tolerance .001. The sweep records emitted point counts and reaches the depth-16 cap. SH directions use a deterministic midpoint-Z/golden-ratio-azimuth lattice; projection uses 64/256/1024 directions and 1/3/5 bands, precomputing basis values outside the timed calls. RGB radiance is finite and asymmetric, approximately [1,5], [.5,1.5], [.25,.25]. Basis-only cases cover degrees 0, 2, 8 and 12. L2 rotation/convolution/reconstruction uses fixed signed coefficients; negative coefficients are valid and do not imply negative radiance input.
Each case executes once before calibration. Calibration doubles its operation count until a trial reaches the requested wall time; seven trials then report all samples, median, median absolute deviation and extrema. GC is disabled only inside timeit and restored by it. Data preparation and import timing are separate. Returned object/ctypes allocation is included in public-operation timing; constructor and raw-kernel cases help locate overhead but are not a causal subtraction experiment. Ray transform cases include duplication to avoid accumulated mutable state, and report duplicate cost separately.
Fresh interpreter import timing excludes process launch and uses warm filesystem caches; it is not a disk-cold machine startup measurement. cProfile and single-call tracemalloc runs occur after timings and their instrumented durations are not benchmark timings. Profile cumulative times overlap; never sum them. Peak traced bytes measure Python allocation pressure, not RSS or total native allocation. Machine load, CPU frequency and timer overhead are uncontrolled. A no-op baseline is reported without subtracting it. Repeat runs on the same machine, compare medians and variability, and validate numerical behavior before performance claims.
Ray intersection is unavailable in the supported implementation and deliberately not invented for benchmarking. Core source and harness SHA-256 fingerprints identify the measured code; JSON contains environment and operation counts.
If a full repeat shows large outliers, recheck selected cases across interleaved rounds without a core change:
This executes nine cases in three rounds, seven trials each, with .05-second calibration. It isolates run-to-run variability from mathematical changes; it does not automatically declare performance gains or impose a noise threshold.
Matrix inverse matched comparison¶
python benchmarks/matrix_inverse.py --output /tmp/gem-inverse.json
python benchmarks/matrix_inverse.py --output /tmp/gem-inverse-repeat.json
python benchmarks/summarize_matrix_inverse.py /tmp/gem-inverse-summary.json /tmp/gem-inverse.json /tmp/gem-inverse-repeat.json
python benchmarks/verify_inverse_baseline.py /tmp/gem-inverse-equivalence.json
The scripts require the post-2G base git object
89cd4b97784d32625de65b8f465cfcbf6e102943. They reuse the calibrated timing
and profiling routines above. Run sequentially without a concurrent test suite.
Ordinary inputs use fixed seeds 310/311, plus uniform extreme scales and a mixed
binary-row-exponent dataset. Independent numerical tests remain separate from
bitwise baseline preservation and performance measurements. The bootstrap
summary is descriptive conditional evidence, not a universal confidence guarantee.
Vector and Quaternion comparison¶
python benchmarks/vector_quaternion.py --output audit/phase3c-benchmarks.json
python benchmarks/vector_quaternion.py --output audit/phase3c-repeat.json
python benchmarks/vector_quaternion.py --output audit/phase3c-followup.json --rounds 7 --target-seconds .04 --case vector2_subtract --case vector3_dot --case quaternion_squad4
python benchmarks/summarize_vector_quaternion.py audit/phase3c-performance-summary.json audit/phase3c-benchmarks.json audit/phase3c-repeat.json --follow-up audit/phase3c-followup.json
python benchmarks/verify_vector_quaternion.py audit/phase3c-equivalence.json
Requires the merged PR #35 git object; no additional dependencies. Both baseline modules are loaded independently, with the baseline Quaternion using the baseline Vector implementation. The unchanged Phase 3A Vector/Quaternion workload section supplies deterministic data. Raw list kernels, public function entry points and wrapper methods are labeled separately. Inputs/setup are excluded; result allocation is included. In-place cases include fresh receiver construction to avoid drift. Two executions of three alternating rounds, each with seven calibrated trials, provide six paired blocks per case. JSON stores individual timings and separate profiles. The seeded bootstrap interval describes these blocks, not universal speedups or a guaranteed population confidence interval. See the performance report.
Bezier and spherical harmonics comparison¶
python benchmarks/bezier_sh.py --output audit/phase3d-benchmarks.json
python benchmarks/bezier_sh.py --output audit/phase3d-repeat.json
python benchmarks/summarize_bezier_sh.py audit/phase3d-performance-summary.json audit/phase3d-benchmarks.json audit/phase3d-repeat.json
python benchmarks/verify_bezier_sh.py audit/phase3d-equivalence.json
python benchmarks/verify_hdr_reference.py --output audit/phase3d-hdr-reference.json --output-dir /tmp/phase3d-hdr-reference
Requires the merged PR #36 git object. Phase 3A workloads use independently loaded baseline/core Bezier and SH modules. Additional cases cover raw splitting, flatness/subdivision, full basis arrays and angular projection. Adaptive output counts are checked before timing. Deep cases retain the full depth-16 workload; slow single calls can exceed the .02-second calibrated target. Warm basis-layout cache costs are separate from cache-miss/resident-memory observations in the report. Input generation is excluded; result allocation and validated wrappers are included. Three alternating rounds of seven trials in each of two executions provide six paired blocks. cProfile/tracemalloc are separate from timing. The HDR verifier regenerates into a separate directory and compares coefficients, linear values, decoded RGB and all file hashes without replacing golden outputs. See the performance report.