Release v0.1 - #28
Open
ThrudPrimrose wants to merge 834 commits into
Open
Release v0.1#28ThrudPrimrose wants to merge 834 commits into
ThrudPrimrose wants to merge 834 commits into
Conversation
added 30 commits
September 29, 2026 20:31
…mpiler's, last in each Dockerfile, and a gate proves one is mapped. containers/lib/one_openmp.sh (called by numpy_on_openblas.sh, and again as the last install step of judge-agent-amd/cpu/cuda) links the system libgomp, spack's gcc-runtime copies and every wheel-bundled libgomp (torch/lib/libgomp.so.1 in the CPU image, xgboost.libs/libgomp-<hash>.so.1.0.0 in the CPU and AMD images) to the image compiler's file, refusing a copy that defines a GOMP version the compiler's lacks; LLVM's libgomp shim and 32-bit multilibs stay. one_openmp_gate.py then runs numba prange + BLAS, a gcc -fopenmp library and every optional wheel (torch, sklearn, xgboost, lightgbm, jax, cupy, tvm, pythran) in one process and asserts a single OpenMP runtime realpath is mapped (openmp_runtimes.py counts them; the package and containers/lib copies are pinned equal because the agent stage may not COPY hpcagent_bench). verify_image.py re-runs the gate in every finished image and tests/test_one_openmp_runtime.py in judge images, failing on any skip.
… submission (grading.single_openmp_runtime, off). native_call.openmp_runtime_gate counts the OpenMP runtime realpaths in /proc/self/maps at the end of every grading child. With grading.single_openmp_runtime on it raises OpenMPRuntimeConflict, which call_failure reports as NativeCallOpenMPConflict (a NativeCallHarnessFault: harness_fault, score_error); off, the child names the runtimes on its stderr and the grade stands. Off by default: clang, hipcc, Polly and OpenMP-offload builds link libomp, and -fopenmp=libgomp compiles the pragma to serial code, so enabling it now would turn every one of those grades into a judge fault.
…penMP probes (every schedule, tasks, taskloop, atomic, critical, simd, locks, threadprivate, barrier team size) that fail on a serial team. Not wired into any gate yet: the design (libomp everywhere) was withdrawn.
…pc contexts, a spawned grading child per family, per-context catalog libraries, per-context gates. Two runtimes in one process cannot see each other's parallel region, and no runtime serves every toolchain (clang cannot target libgomp, gcc emits libgomp calls, NVHPC has libnvomp). Every image carries a context per family under /opt/omp (containers/lib/omp_contexts.sh): gnu (the image gcc's libgomp and /opt/view), llvm (the libomp hipcc/amdclang resolve, libgomp.so.1 and libiomp5.so links to it inside the context only so numba's GOMP-ABI pool maps libomp, and /opt/omp/llvm/view, a second spack environment built %llvm with shared_linking runpath: openblas, scalapack, fftw, suite-sparse, superlu, superlu-dist, mumps, strumpack, hypre, arpack, magma, sundials, petsc, slepc) and, on the CUDA image, nvhpc (libnvomp, NVHPC's bundled BLAS and LAPACK behind a libopenblas.so.0). hpcagent_bench/omp_context.py maps a toolchain family (languages.submission_toolchain, offload legs included; numba imports; a prebuilt library's DT_NEEDED) to its context. _call_isolated(omp_context_name=) starts a child of another context as a spawned interpreter (run_forked(env=)) whose LD_LIBRARY_PATH leads with the context's lib/, for the candidate, the C/compiled references (context of the compiler that built them), the numba reference and the parallel oracle (llvm), the profiled child and the sanitizer child. Catalog tokens resolve in the context's view first; hpcagent_bench/omp_catalog.py measures which runtimes each catalog library's link closure maps per context (written to /opt/omp/catalog.json at image build) and sandbox.catalog_refusal refuses, up front, a library that would map another runtime. grading.single_openmp_runtime is on by default; libnvomp as the only extra runtime is logged (stderr, the grade's detail) and let through. containers/lib/omp_context_gate.py runs, per context in one process, the family's compilers (gcc/gfortran; clang, flang, hipcc, amdclang; nvc, nvfortran) on an OpenMP probe (every schedule, tasks, taskloop, atomic, critical, simd, locks, threadprivate), numpy/scipy BLAS, numba prange calling BLAS and torch, asserts every team exceeds one thread and exactly one runtime is mapped; omp_context_scan.py asserts no library of a context maps another runtime and numpy/scipy/numba carry no absolute RPATH. Dockerfiles, verify_image.py, tests and docs (anti_cheat.md, IMAGE_REQUIREMENTS.md, containers/README.md, library_requests.md) change together.
The judge's /submit grades a single-node submission with regrade.submit_grade: the code regrade finalize runs (final_grade over m=4 inputs x n=5 runs a side, Mann-Whitney per input) under regrade.final_settings, scoped to the request (config.scoped_environment), with the held-out cases riding with the first input and the sweep ended at the first failing input. /score keeps its keys. recording.record writes a credited submit together with its own final grade (a `final` row, of_grade_id = the submit, same numbers and grade_cells) in one transaction, so nothing times it again. The in-job re-timing goes: final_grade.py and its worker queue, wait_final_grades in run_cluster.sh, the chained grade-pending job (jobs.py action, sbatch, submit_common.sh). regrade finalize stays for submissions an older /submit protocol graded, stale and owed ones. MPI/ML-scaling tasks keep their grade. Docs, config comments and schema comments follow; tests pin the protocol, the single timing, the final row and finalize on an old-protocol row.
…oneapi toolchain family, no libiomp5 in any image. The oneapi family leaves languages.COMPILER_FAMILIES and REPORT_REFS, the icx/icpx/ifx blocks leave compilers.yaml and toolset.yaml, CPU_BASELINE_ICPX and ICX_OPT_REPORT leave flags.py, the cc_oneapi column leaves FRAMEWORK_META, the prompt's per-family table and the C++ TBB sentence name gcc and llvm only, the opt-reports skill drops the oneapi line, the OpenMP context map drops oneapi, and the CI install script, the setup action and the workflow install NVHPC only. containers/lib/parallelizer-gate.sh no longer probes icx or icpx. The AMD image no longer installs Intel MKL (spack intel-oneapi-mkl, /opt/intel-mkl, MKLROOT): its threading layer is libiomp5, a second OpenMP runtime no context runs on. libiomp5 stays in the runtime counter, so a wheel that bundles one is still a detected foreign runtime. envs/registry.yaml keeps the cc_oneapi key as "oneAPI (retired)": the list is append-only, removing it would repaint every entry after it in every published figure (tests/test_palette.py), and a recorded cc_oneapi row still resolves to a name in analysis and plots. Nothing indexes FRAMEWORK_META by a recorded column outside running one. Tests drop the oneAPI cases and keep every other compiler's.
…, Mann-Whitney credit
…y (AMD 655840); verify_image sources cache_env.sh so selfcontained_check sees HPCAGENT_BENCH_DATA_ROOTS (judge-cpu 655841)
…asserted on PATH, so a missing context is a bug); omp_contexts.sh fails on OMP_REQUIRE_NVHPC=1 without nvc.
…ing and loop_level_reasoning grade against best-of(numba, c), machine_learning against the compiled torch-autotune reference Oracle: resolve_oracle names the track's reference (compiled / numba / c / torch); oracle_kinds orders numba and C by the measured race leader (baseline_leaders.yaml, XL, taken at every preset; loop_level_reasoning starts from C, its verdicts' reference; KERNEL_COMPILED_HEAD sends bdf_newton_krylov to numba because its C leader is 4e-9 off NumPy at M). A reference that cannot answer raises ReferenceUnavailable and the grade moves to the other one; when none answers it is a harness_fault, never a numpy grade. Replaces COMPILED_ORACLE_KERNELS / PARALLEL_ORACLE_KERNELS (njit of the numpy source) with the kernels' numba references run in the sealed judge child (grading.numba_reference_outputs). Every grading road was moved off _numpy_reference: graded_score (public, held-out, re-verified checks, write probe), independent_verify (the dual leg is now the OTHER compiled reference: C beside numba, numba beside C, C beside torch), the distributed grade (oracle and one-node denominator), scaling anchors, score_cells, uncovered_grade. The write probe re-runs the reference that produced the expected outputs. numpy is only a timed denominator on an explicit machine_learning request; _numpy_reference / reference_function stay for tests and CI, where the compiled references are proven equal to it at S (tests/test_grading_never_numpy.py pins that no production call reaches them and drives real grades of each track with every road to numpy raising). Docs, config, api.Oracle, schema.sql and results_db note that scicomp/LLR never record 'numpy' as a denominator (the CHECK keeps it for old rows).
…per-kernel XL ceiling: sizing.KERNEL_XL_CEILING holds it at 10 GiB, the track's 4 GiB stays for every other kernel The oracle is now the numba reference, not numpy, so the numpy cost that capped XL at 2^25 is gone. Measured on one mi300 node in the judge image: the numba reference takes 0.72 s at 2^27 (0.8 s child wall, 19 GiB), 1.4 s at 2^28, 2.9 s at 2^29 (72 GiB); a real /score grade at 2^27 (numba python candidate, numba oracle, race numba 0.71 s, C cut at 12 s) took 366 s and peaked at 46 GiB in the judge process and 56 GiB in its largest child; /submit (public + 5 held-out, XL among them) 379 s and 58 GiB; the re-verify 404 s (the C dual leg times out at 300 s and is recorded not-applied). 2^28 cannot be graded: each timed rep redraws its inputs (5 s there) inside the 5 s guillotine floor. grading_cuts.yaml and baseline_leaders.yaml follow; xl_ceiling(track, kernel) is the one place scripts and tests ask, and tests/test_xl_ceiling.py pins the override to the kernel that needs it.
# Conflicts: # hpcagent_bench/harness/grading.py # hpcagent_bench/harness/scoring.py
…-Whitney alpha 0.1, the final grade's code under measurement.score; prompt, docs and schema comments follow
…=llvm; a bare %llvm bound Fortran only, 655986); STRUMPACK ~butterflypack on AMD and CUDA (flang cannot parse ButterflyPACK's fixed-form sources, 655985)
…e that fits the matrix is the submission's job; sizing bounds padded buffers by stored entries
…ery requirement (sundials was built for zen4, 656016); patchelf for scipy's un-isolated meson build (656017); format the OpenMP probes
…virtual's provider requirement takes a named spec (656527)
…lain Enum; typed cache; stale test paths and lambdas; docs toctree; the runner links wheel libgomp copies to the compiler's
…hwloc linked against libcudart.so.12 is not reused (daint 4947813)
…track-default tests follow the current build and denominator; the tsvc porter carries the averaged wavefront bodies; edge-shape test checks the constraint, not a label list
…ler-family spy records denominator builds only (the oracle's C build names no block); a pattern-only kernel is bounded by its index arrays
…host without them names the runtimes on stderr. pkg-config is asked with the context's PKG_CONFIG_PATH only (pkgconf drops a -L that LIBRARY_PATH names). The context gate test runs without inherited thread pins
… OpenMP context's LD_LIBRARY_PATH still leads (AMD 656542)
added 30 commits
October 3, 2026 17:37
… gate, and the prompt no longer points agents at man
…ree, native CPU builds, MAGMA in the llvm context only
…rict pyright project
… package, unit constants, HTTPStatus names, strict-typing fixes
… the image's /opt/rocm: PyTorch's ROCm wheel bundled a second libamdhip64 and librccl and crashed every ML grade; the image build asserts it and verify_image gates one HIP runtime and RCCL per process
…ndices, Owned, PromptFold, ResolvedOutput; task_token_totals returns the EpisodeTotals it reads
… string: sizing, scoring, curves, recording, the baseline curve and Harbor take the member; mpi.mode and the DB keep the spelling, parsed or written at that boundary, so a misspelled law fails at the config
… sandbox BuildResult, run_compiled_reference and the C reference a ReferenceRun (outputs, best_ns, hidden, samples_ns)
…l (status, payload), containers.Backends (spellings, passthrough)
…(one failing test: regrade replay StopIteration)
# Conflicts: # hpcagent_bench/stats/figures/scaling.py
…regrade worklist stages each setup through submit.sh
…t; env_spec's Model enum (members from the layer files) carries a reasoned ignore
…uncalled static inline in C++, and a kernel with no heap allocation never calls it (18 warnings broke the -Wall -Wextra ratchet)
…on errors carry their SDFG, fp16 mixed operands, framecode allocation placement)
…extended now refuses a gap in tuple return names (__return_1 with no __return_0)
…omic and a running job keeps the inode it mounted, so only the in-place writes (build, export, pull) refuse a mounted path
…alar version (two arrays sized by one runtime scalar combine again)
…ror, detail) NamedTuple; callers read .ok. measurement_statistics documents the ML final grade's 4 aligned inputs in one launch
…ross loops until its scalar is written
… node the login node's /run/user/<uid> is missing and the first role's probe read 0 GB and no digest
…, scaling inputs are drawn around XL, and the scaling tables key on the input (schema 4) - grade_under.Item gains `scaling: Scaling(laws, rank_counts)`; `grade-under worklist` asks it of every item whose task scales (service.scales, read under the setup's env; ml.grade_rank_counts = 1..16) and reports a scaling item without a recorded distribution. Worklist lines round-trip the enum laws (Item.as_json / of_json). - `grade-under run` / `job grade-under`: the per-task shape grades the plain items and leaves scaling ones owed; the gang shape (docs/jobs/grade-under.sbatch GANG_NODES, `--gang/--gangs`) grades the scaling items whose max(P) the gang places (scaling_grade.placeable_ranks). The plain path no longer misgrades an ML item as single-node. - scaling_grade is a library (grade, grade_into, run_shard, baseline curve); its run-dir worklist, auto/claims mode (scaling_claims.py) and hpcagent_bench/cluster/mlscale-grade.sbatch are deleted. ml_protocol_grade -> scaling_protocol_grade, ml_scaling_grade -> scales. - The scaling final grade draws its 4 inputs in [0.5, 1] x XL (metric.size_presets(anchored=True): a scaling kernel's `fuzzed` preset is its small correctness range, which had collapsed the inputs to one 64x512 shape), and each input is the P=1 base of its own strong and weak sweep (every sized problem of one P in one launch), anchored at its own torch time. - Results schema 4: `input` joins scaling_grades (grade, mode, input) and scaling_points (grade, mode, input, ranks); `python -m hpcagent_bench.harness.results_db upgrade` brings a v3 file up in place. The extractor emits scaling_input; the scaling figure folds a law's inputs by geomean per P. - Tests: fakes follow the multi-draw plan and the Draw API, BuildResult replaces SimpleNamespace builds, laws compare as ScalingLaw, worklist staging names a system. Docs: jobs/README, results_db, measurement_statistics, plotting, experiments LAUNCH/README, config and setups comments.
…y slice bounds, across nested-call regions, and refreshed by a loop that writes the size (vexx_k, cegterg, cp2k_grid_integrate, warpx parse again)
…ailed batched ML launch relaunches its draws alone, /score answers its timed cells again, and the in-process agent baselines are gone Production fixes found while fixing the tests: - scoring.ml_launch: a multi-draw launch that fails before any verdict relaunches each draw alone, so a weak law's grown size failing no longer erases the strong point at that P. - service /score: the payload built cells from dataclasses.asdict, whose tuple as_list read as empty, so every /score answered cells: []. - scoring.sanitized_run: a host whose AMD GPU arch cannot be detected leaves the HIP leg not applied instead of failing the recording. - sanitizers.HOST_CANNOT_MAP: "Shadow memory range interleaves" without the AddressSanitizer prefix is the host's refusal, not the submission's memory error. - observations_extract.clocks_agree_on_delta: no clock reading keeps the stored flag (the None branch was unreachable behind the type check). - grade_under.staged_env: a failed staging is retried instead of dying on mkdir. Deleted harness/baselines.py (bare/tools registry, ModelSpec/MODELS, the context ladder whose minimal rung was larger than no_hints) and agent --agent-baseline; BACKENDS moves to harness/agent.py and the agent verb calls runner.solve_task. Tests: one shared real rank-driver launch helper, hermetic setup envs, fakes matching the real API, typed mpi_shard_driver/cli/test_mpi_shard, underscore helpers renamed, annotation baseline rewritten.
…ach run's outputs against its own input, and the ML ranks check every repeat The protocol's name and its numbers are unchanged; what it verifies is not. Before, a grade re-ran 2 secretly chosen timed inputs after the loop and graded those; the timed calls' own outputs were never looked at, so a latent race that fired in one of 20 runs passed. - rep_variation: each (kernel, preset, datatype) cell has a fixed pool of 4 seeds (pool_seeds, from the route's unsalted secret seed); a call cycles it from a nonce-chosen offset (timed_seeds), so consecutive calls never share an input, and the public base seed is one untimed canonical call on both routes. final_seeds, derived_seeds, verify_indices, check_pool, pick_checks and the vary_inputs_untimed_base / repverify_* settings are gone. - native_call: sampled_calls keeps every timed rep's outputs (the copy off the device was already the call's own, outside the clock), spilled as each returns; _call_isolated returns an IsolatedCall with them. - scoring.score grades every timed run against the oracle on its pool input, one input's expected outputs held at a time, keyed by pool seed and kept structure so later grades of the cell (and the disk store for cached levels) reuse them; reason names the run. - mpi_shard_driver: every repeat's output shards are copied off the device after its clock stops and graded against the one reference, one repeat on the device at a time. - full_oracle_checks is runs_write_probe: the re-verified checks it also gated are gone, and every track's runs are checked.
…the shared timing, graded at repeat=5 so every pool seed is stored The two store tests overlapped almost entirely; at repeat=3 a grade touched only 3 of the cell's 4 pool seeds, so a later /score with the reference forbidden could need the unseen one.
…(torch_dist_curve passes the world)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.