Skip to content

Latest commit

 

History

496 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Optimizers, weight trajectories, and dense retrieval

Delivery complete: final status and agent handoff. The full scientific-source CI passed; all 24 runs, 120 checkpoints, analyses and the reviewed paper are delivered. The owner-approved old HF files have also been removed from the current trees, with valid backups and history preserved. Dated pending notes farther below are historical snapshots.

How do AdamW, Muon and NorMuon change a pretrained retriever's weights, and which changes matter for retrieval? This DenseOn study connects weight trajectories, functional representation utility and held-out retrieval. The target is a NAACL paper with reproducible model and analysis artifacts.

Complete DenseOn results. All three comparisons averaged over the declared four-rate grids are inconclusive. Validation-selected NorMuon exceeds AdamW by +0.437 nDCG@10 points (simultaneous 95% interval [+0.146, +0.729]); selected Muon is +0.317, with its interval crossing zero. These are one-seed, fixed-grid results, not a universal optimizer ranking. The complete weight, representation and crossed-continuation findings are in the reviewed paper and evidence summary.

The complete source payload is publicly available on GitHub main, with verified publication and anonymous readback. For completion status and remaining limits, read CURRENT_EXPERIMENT.md; agents must also follow AGENTS.md. Dated evidence is retained in PROJECT_STATUS.md.

Current progress

Snapshot: 2026-09-13, not a perpetual heartbeat.

Current status and agent instructions now contain only the active handoff; the full chronological records remain in the preserved history.

Read the complete-result paper PDF and revision notes. This reviewed local candidate contains the complete results, revised story and closest references: eight-page main text, 158-word abstract, all 13 pages visually inspected. The default build now selects this paper, and the complete Make release-artifact gate has passed. The conclusion above is generated from its authenticated numerical outputs. All 3,800 source-version tests pass locally; the source payload has been published and anonymously byte-verified. Independent hosted full-workflow acceptance remains pending. All3800 tests passed in a hosted run; the subsequent exact CPU-replay repair now passes complete numerical-to-paper reconstruction locally. See CPU replay conditions and CI runs. The final repair's complete local regression and reusable numerical-to-paper entry now pass; actual evidence is retained before a single publication push. Public source availability is not a claim that the full hosted workflow has passed.

The complete manuscript now has a stable source/build entry:

make -C paper current CURRENT_OUTPUT=/absolute/path/to/new-paper-build

The actual Make target and an extracted-wheel rebuild both pass; the latter loads all study modules from the wheel and reproduces the reviewed PDF text. See original package verification and completed release transition.

Component Verified state
Primary experiment 12/12 runs, five stages each; 60 checkpoints backed up and HF hash-verified
Primary retrieval 840/840 checkpoint/task evaluations; 14/14 pretrained-baseline tasks; validation-only selection complete
Weight and functional analysis Complete 60-state geometry/prediction panel and 61-state functional panel including the pretrained reference
Crossed state-by-operator continuations 12/12 runs and 60 additional checkpoints complete and backed up; 60/60 five-stage probes complete
Continuation full-corpus retrieval 168/168 complete, all workers and both coordinators exit zero
Complete continuation inference and paper All six tables/three contrasts verified; full numerical/PDF replay passed; new complete-result prose/reference revision built and all 13 pages visually reviewed
Current-paper source/build Stable paper/current and package entry implemented; 86 focused tests pass; actual Make and extracted-wheel builds pass
Complete numerical-to-reviewed-paper command Version-isolated installed-wheel execution passes, joining all eight shared inputs; 194 integration tests pass; run instructions
Distribution portability Original full audit passes; real relocated input read is unchanged; historical failures and all negative controls retained
Primary training source All 33 modules and 56 bindings match the actual training snapshot; all 12 run sources load in this checkout; inspection instructions
Local full source-version tests 3,800 passes: current 2,873, original analysis 733, original factorial 194; no failures/errors/skips; one-command test entry
Independent hosted CI All 3800 tests passed remotely; the latest full numerical child failed. Complete local repair regression and reusable emulated replay now pass; hosted full-workflow acceptance remains pending
Recovery and source delivery Both genuine GPU restores pass exact endpoints; scoped recovery add-on and current default/release entries work; complete source payload published and anonymously verified

All scientific training/evaluation and the statistical, document, numerical/PDF reconstruction and result-backup pipelines have finished locally. Local full source-version regression and distribution checks pass, and source publication is verified. No scientific training/evaluation is running. Both bounded warm-reducer restoration checks now pass bitwise endpoint equality and are terminal. Do not restart completed work. The earlier diagnosis retains both original failed comparisons; neither attempt replaces scientific states. Operational observations retain actual process exits and current-task timing evidence.

What the completed primary results support

Averaged over all four declared learning rates, the primary optimizer comparisons are inconclusive. Validation-selected Muon and NorMuon have higher retrieval point estimates than selected AdamW at all five retained stages. At the final stage, selected NorMuon–AdamW is +0.437 nDCG@10 points (simultaneous 95% interval [+0.146, +0.729]); selected Muon–AdamW is +0.317, with its interval crossing zero. These are paired-task intervals within each estimand family, not independent-training-seed intervals. Primary and validation-selected comparisons answer different questions. See all six endpoint contrasts and all rates and stages.

All twelve measured retrieval trajectories, retaining every declared learning rate

Stars mark validation selection; lines connect measured checkpoints. Crossing AdamW's final score earlier does not establish deployable wall-time savings: selection used full-horizon validation.

The weight/functional analyses test explanations, not just differences in spectra. Displacement and degrading-attribution mass predict retrieval under the original baseline, but no tested weight or functional predictor passes all four exploratory recipe comparators. The Muon–AdamW helpful-participation contrast is not stable across the sampled rotations. These findings do not establish “Muon uses more useful dimensions” or a causal gain mechanism. Read the weight controls, functional inference and functional controls. The complete crossed continuation now separates reached state from subsequent update rule under its fixed 50K-query design:

Contrast nDCG@10 points Marginal 95% interval
Muon-source minus AdamW-source state, averaged over continuation rules +0.322 [+0.047, +0.678]
Reset Muon minus reset AdamW, averaged over source states -0.526 [-0.897, -0.186]
State-by-operator interaction +0.208 [-0.123, +0.555]

These are three marginal intervals from the declared two-way seed/task bootstrap, not simultaneous intervals or independent primary-source replications. The result supports a distinction between the two reached weight states and the locally preferred continuation rule; it does not establish mediation or a universal optimizer ranking. Complete native evidence and scope.

Experiment design

Choice Primary experiment
Initialization DenseOn-unsupervised, revision 0edbd55684eb782bce55ee74c95b25c97cbe7f43
Data and objective Same deterministic 500,000 queries; one positive and seven hard negatives; no in-batch negatives; temperature 0.02
Schedule One epoch, global batch 128, maximum context 8192
AdamW learning rates 1e-6, 3e-6, 1e-5, 3e-5
Muon / NorMuon learning rates 1e-4, 3e-4, 1e-3, 3e-3
Retained steps 782, 1563, 2345, 3126, 3907
Evaluation and selection Fourteen pinned decontaminated BEIR tasks; separate frozen 4,096-query validation loss
Replication scope One primary training seed (42); learning-rate cells are not independent seeds
Actual training hardware Eight NVIDIA L20Z GPUs, split into two disjoint four-GPU pools; not literal H100 hardware

The v3 primary protocol binds data, model, runtime and checkpoint identities. This compares complete optimizer recipes: the primary AdamW rate varies across all its parameters, whereas Muon/NorMuon use auxiliary AdamW at 3e-6. AdamW's validation-selected rate is the upper tested grid boundary, not an established optimum.

The functional probe uses 224 fixed queries across fourteen tasks, eight candidates per query, 768-coordinate deletion with cosine re-normalization, shared random masks and three sampled rotations. It is not full-corpus compressed retrieval, and sampled rotations do not establish arbitrary-basis invariance.

The crossed continuation design uses two genuine source states × two reset optimizer rules × three data-order seeds: 12 continuations on the same fixed 50k queries. Initial hidden-update calibration is not trajectory-wide norm matching, and the three order seeds are not three independently trained source states. Only DenseOn is active; the scope amendment preserves historical LateOn artifacts without new LateOn experiments or pooled architecture claims.

Recover artifacts and reproduce analyses

Need Start here
Primary checkpoint weights and optimizer state Immutable checkpoint restoration and all sixty primary locations
Continuation checkpoints and five-stage probes Continuation restoration guide and checkpoint index
Inspect an already restored continuation save CPU-only authenticated checkpoint reader; original identities retained, not GPU-resume validation
Resume either of the two verified continuation saves Scoped restoration add-on; requires the original admitted four-rank worker, not a general new-host launcher
Complete continuation retrieval and statistical tables Immutable 593-file outcome snapshot and restoration
Full primary retrieval trajectories Retrieval tables, raw-score snapshots and figures
Weight / functional measurements Weight-analysis restoration, functional/calibration restoration
Complete primary + continuation numerical graph and original full PDF Paper-results reproduction; unchanged local closed bundle, separate reviewed-prose build
Complete numerical graph joined to the reviewed manuscript in one command Version-isolated paper reproduction; original numerical source and current document package both actually executed
Manuscript and acceptance requirements Paper directory, completion gates

Public data/model repositories are checkpoints and analysis artifacts. Use the guides' immutable revisions and digests, not mixed historical namespaces. The complete original numerical closure and versioned source/test entries are now included in the verified public GitHub payload. Numerical reconstruction is not a fresh encoding, GPU-training replay or physical second-host proof.

Development and release

Development only, in a separate environment:

uv sync --extra dev --extra eval --extra analysis
export PYTHONPATH="$PWD/src:$PWD"
# Install the hash-locked formal runtime first; see CONTRIBUTING.md.
CUDA_VISIBLE_DEVICES='' uv run --no-sync python -B scripts/test_source_roles.py --output /tmp/dense-source-tests-new

Do not install or upgrade packages in the live experiment environment. Formal execution uses formal_runtime.json, requirements-formal.lock and requirements-formal-flash.txt. Current execution authority and exact source bindings are in the agent handoff; archived controller commands are not current launch instructions.

The reviewed current paper contains both primary and continuation findings; its prose revision, complete visual review and document reconstruction have passed. The original full distribution audit now passes. The current default and complete release-artifact Make targets now pass using the original strict numerical/document components. The original mixed-source test attempt remains a recorded failure; all 3,800 cases now pass locally in their explicit source roles. Independent hosted CI remains pending and uses an authenticated runtime binary to accelerate installation without changing the scientific runtime or test gates. No engineering-error narrative belongs in the manuscript.

The former homepage preserves the detailed chronology and earlier reproduction/test records. AGENTS.md retains all source, process, external-access and publication boundaries. Tracking is available through W&B and the project issue; neither substitutes for native completion evidence. Never put credentials in source or released logs.

License and attribution

Apache-2.0. See LICENSE, THIRD_PARTY_NOTICES.md, CITATION.cff, CONTRIBUTING.md, CODE_OF_CONDUCT.md and SECURITY.md.

About

Reproducible DenseOn optimizer study: AdamW, Muon, NorMuon, decontaminated BEIR, dynamics, and packed-validation audit

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages