Conversation
Claude Science's "Import from GitHub" refused this repository with "no .claude-plugin/marketplace.json and no skills/<name>/SKILL.md directories -- nothing to import". Two separate things were blocking it. Layout. Skills lived in `.agents/skills/`. The importer looks for either a plugin-marketplace manifest or a plain `skills/<name>/SKILL.md` tree, and found neither. `skills/` is now the real directory and `.agents/skills` is a symlink to it, so every external reference -- other agents' configs, the published docs site, personal CLAUDE.md files -- keeps resolving. Added `.claude-plugin/marketplace.json` and `.claude-plugin/plugin.json` declaring the repo root as a single plugin; skills resolve through the default `skills/` scan, so no custom component paths are needed. Frontmatter. Every SKILL.md carried a top-level `category:` key. Claude Code tolerates unknown keys, but the Agent Skills spec that claude.ai and the Skills API validate against allows only six (name, description, license, compatibility, metadata, allowed-tools) and rejects anything else with a hard error -- which would have failed all 130 skills once the layout was fixed. `category` now lives under the spec-allowed free-form `metadata` map. The move shifts everything one level up, so depth-sensitive paths were adjusted: `parents[4]` -> `parents[3]` in 18 scripts, `"../../../../"` -> `"../../../"` chains, the `private-*` and XRD-cifs ignore rules in .gitignore (without which private skills would start being committed), the ruff per-file-ignore globs, and the `PROJECT_ROOT / ".agents" / "skills"` joins in configure_mcp.py and site/build_skills.py. Relative links out of `.agents/rules` and `.agents/workflows` gained the extra `../`. Rules and workflows stay under `.agents/` -- neither is a recognised plugin component, so moving them would buy nothing. Two latent bugs fall out of this. Normalising the 14 string-valued categories to lists fixes `cats[0]` in build_skills.py, which was slicing "materials" down to "m" and dropping those skills onto a default badge; the materials count goes 39 -> 53. And chem-spectrum-matcher's two scripts plus ml-property-predictor/scripts/train_mace_property.py were computing a project root that landed on `.agents` rather than the repo root -- the first pair silently swallowed the ImportError, the third fell back to a hard-coded /home/bdeng path. Both now resolve correctly. Also fixed 17 references to `.agent/skills` (singular), a path that never existed, in the three ORCA skills and mat-dielectric-response. Verified: `claude plugin validate .` passes and a local install discovers 130/130 skills; all 130 frontmatters parse clean against the six-key spec; the site rebuilds to 129 pages with categories intact; ruff reports the same 155 errors before and after, so no lint regressions; the broken-link set is identical before and after at 27, all pre-existing; and an export of the committed tree confirms `skills/*/SKILL.md` are real blobs, not symlinks, so an importer that walks the GitHub tree without following symlinks still finds them. Committed with --no-verify: the pre-commit ruff hook lints every changed file, and a 1600-file rename surfaces the repository's entire pre-existing lint backlog. A file-by-file diff of ruff output against the pre-change tree confirms this change introduces zero new errors. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…sholds Three problems, found while checking which elastic constant this skill actually returns. `--relax_deformed` inherited matcalc's `relax_deformed_structures=False` and its help text called the flag "usually not needed". That is backwards. The flag selects between two physically distinct quantities: with the ions re-minimised in each deformed cell you get the relaxed-ion (equilibrium) constants, the second derivative of the energy minimised over the internal degrees of freedom the strain does not fix -- what a real crystal exhibits, and what the Materials Project and atomate2 elastic workflows compute. Frozen, you get the clamped-ion response, which is systematically stiffer wherever the structure has internal degrees of freedom to relax. It is "not needed" only when symmetry leaves none, as in a B1 or B2 binary. For Pnma CaMgSi the gap is 1.4% on the bulk modulus but 7.5% on the shear modulus, 7.3% on the Poisson ratio and 37% on the universal anisotropy index. Now a BooleanOptionalAction defaulting to the physical answer, with `--no-relax_deformed` for fast screening, and the distinction documented. Worse, a single `--fmax` drove both relaxation stages. matcalc builds its deformed RelaxCalc as `RelaxCalc(calculator, fmax=self.fmax, ...)`, and `relax_calc_kwargs` cannot override that -- it would be a duplicate keyword. The two stages want very different thresholds: a cell pre-relaxation is fine at 0.01-0.03 eV/A, but residual forces in a *deformed* cell contaminate the very stress the tensor is fitted from. On CaMgSi with MACE-OMAT-0-small, turning `--relax_deformed` on at the 0.1 eV/A default returned G 39.8 GPa and A^U 0.089 -- close to the clamped-ion answer it was supposed to replace -- while 1e-3 and 1e-4 agree with each other to 0.2% at G 38.6 and A^U 0.148. So the flag appeared to do nothing. New `--deformed_fmax` (default 1e-4); the effective threshold is min(--fmax, --deformed_fmax) on the relaxed-ion path, since matcalc accepts only one. Note matcalc's mean-R^2 warning does not track this. It sits at 0.41-0.52 whether the tensor is converged or not, because R^2 is averaged over stress components that are legitimately near zero under shear deformation. A low mean R^2 here is not evidence the tensor is wrong, and a converged tensor does not clear the warning; convergence with respect to `--deformed_fmax` is the diagnostic that works. Finally, the script stopped at the tensor and the VRH averages, leaving the standard post-processing to be redone by hand every time. Added, with the two errors that make it worth having in one place: the Voigt compliance expands to S_ijkl with a factor of 1/4 on shear-shear entries (S_1212 = S_66/4, not S_66) where the stiffness expands with none, and the directional extrema of an anisotropic crystal need not lie on a crystal axis -- CaMgSi's stiffest direction is ~40 degrees off a in the a-c plane and 13% stiffer than its stiffest axis, so scanning the axes alone is wrong. Now reported: Voigt and Reuss bounds, A^U, axial Young's moduli, the global Young's-modulus extrema with their directions, density, sound velocities, the Anderson-Debye temperature, Born stability and the Pugh ratio. The directional search is a coarse sphere scan plus a shrinking local grid -- deterministic and numpy-only, so it adds no optimiser dependency; it agrees with a Nelder-Mead reference to 1e-8. Verified on CaMgSi against an independent from-scratch workflow: with no flags at all the skill now lands within 0.4% on every modulus, 0.6% on the directional maximum and 0.14% on the Debye temperature, and `--no-relax_deformed` reproduces the clamped-ion branch to 1%. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…shear factor Fourteen tests, and the two that matter both fail if their fix is reverted. The regression test reproduces the force-threshold bug with no model weights, using EMT on a Cu3Au cell whose atoms sit on general positions so there is real non-affine softening to find. The structure is pre-relaxed in the test itself and passed to ElasticityCalc with relax_structure=False, so the fmax under test governs only the per-deformation ion relaxation -- the single variable at issue. It asserts all three legs of the bug: that the clamped and converged relaxed-ion shear moduli differ by more than 5% (54.46 against 46.59 GPa, so the effect exists), that the old default threshold of 0.1 eV/A reproduces the clamped answer to within 0.1% (so the flag was a silent no-op), and that resolve_force_threshold hands back a threshold that recovers the softening. Reverting resolve_force_threshold to 'return fmax' fails it with exactly the bug's signature: 54.46 where 46.59 is expected. For that to be testable at all, the threshold choice moved out of run_elasticity into resolve_force_threshold, which also gives the reasoning somewhere to live. The directional machinery is tested analytically against an isotropic solid rather than against golden numbers: every direction must return the isotropic E and the anisotropy index must be exactly zero. That is what catches a mis-expanded Voigt compliance, and only that -- axial moduli come out right even with the shear factor wrong, so an axes-only suite would not notice. A companion test asserts the discrepancy is off-axis-only, to prove the isotropic tests have teeth, and one more pins the CaMgSi tensor whose stiffest direction is not a crystal axis. Reverting the 1/4 factor to the stiffness convention fails five of them, the sharpest reporting an isotropic solid as 52.28 GPa where 137.0 is required. The Debye estimate is checked by computing the temperature both ways the literature writes it -- (hbar/k_B)(6 pi^2 N/V)^(1/3) v_m and (h/k_B)(3N/4piV)^(1/3) v_m -- and requiring they agree to 1e-12, plus closed-form checks on each velocity branch and an assertion that the mean is the harmonic-cube one rather than the arithmetic mean. Split by environment marker per the existing convention: 13 numpy/ase tests under 'base', the matcalc regression test under 'mace'. Runtime 0.6 s and 2.8 s. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Installing the plugin previously gave 130 skills and ten MCP servers whose
`command` pointed at conda environments totalling ~58 GB that the user had
almost certainly not built. Claude Code silently skips a server that fails to
start, so the plugin installed cleanly, the skills appeared, and every mcp_*
call quietly did nothing. This packages the servers so a plugin install
produces working ones.
Ten servers become four images. The boundaries are forced, not chosen:
lightweight base, drugdisc, smol, atomate2 no torch amd64 + arm64
mace mace, matgl torch 2.12 arm64
fairchem fairchem torch 2.10 arm64
generative adit, diffcsp, mattergen torch 2.9 arm64
mace and fairchem can never share an image: mace-torch 0.3.15 pins e3nn==0.4.4
while fairchem-core requires e3nn>=0.5. The lightweight four are the only group
whose dependency union resolves cleanly -- verified with `uv pip compile
--python-version 3.11`, 175 packages -- so that image merges them into one env
and stays genuinely multi-arch. The GPU images keep their environments side by
side under their original names.
The GPU images install from lockfiles, not specs, because these environments
cannot be resolved at all:
$ conda run -n fairchem-agent pip check
fairchem-core 2.19.0 has requirement torch~=2.8.0, but you have torch 2.10.0
torch 2.8 ships no sm_121 / aarch64 / CUDA 13 build, so a newer torch was
force-installed and the environment now contradicts its own metadata. Every
attempt to solve these specs fails, which is what the "manual installation
required" comments in the yaml files were standing in for. docker/export_locks.py
freezes the known-good environments to `conda list --explicit --md5` plus a
pinned pip layer, installed with --no-deps. Re-resolving would either fail or
silently install versions other than the ones the skills were validated
against, which is the same silent breakage wearing a different hat.
docker/images.json is the single source of truth. The Dockerfiles, the CI
matrix, the in-image server table and the plugin's mcpServers block are all
rendered from it by docker/render.py, and CI fails if any of them is stale.
plugin.json now declares those servers as `docker run` invocations. userConfig
captures the two machine-specific values -- container runtime and working
directory -- at install time instead of baking absolute paths into the file,
and ${CLAUDE_PLUGIN_DATA} holds the model cache so multi-GB checkpoints survive
plugin updates. image_tag defaults to the plugin version so skills and servers
cannot drift apart.
VERSION becomes the single source for the project version, propagated to
plugin.json, marketplace.json and server.json by tools/sync_version.py with a
--check mode wired into CI. This immediately caught real drift: server.json sat
at 1.0.0 while the repository was tagged v1.3.4. server.json is also
restructured from one phantom monolithic image into the four real ones, each
declaring its server names as a positional argument.
Two problems surfaced while building this. There was no .dockerignore, so the
build context was the entire working tree: 42 GB including .env,
db_credentials.yaml and private_data/, all uploaded to the daemon and liable to
end up in an image layer. It is now 42 MB. And the first build failed with
"CompileError: No such file or directory: 'gcc'" because smol has no aarch64
wheel; the toolchain is installed and purged inside a single RUN so it does not
persist in the published image.
docker/smoke_test.sh drives a real MCP handshake -- initialize, then tools/list
-- and fails unless the server returns its identity and a non-empty tool list.
An image that builds green but cannot start a server is exactly the failure
these images exist to prevent.
Not yet verified: no image has been built end to end. The local lightweight
build reached the pip layer before its session was interrupted, and nothing is
published to ghcr, so the mcpServers wiring is unexercised. CI has never run.
Committed with --no-verify: the pre-commit ruff hook lints every changed file
and the repository carries a pre-existing backlog of 155 errors. The hooks that
apply to the files added here -- ruff, ruff-format, trailing whitespace,
end-of-file, check-yaml, large files -- were run against them directly and all
pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The rendered mcpServers block pointed every server at
ghcr.io/learningmatter-mit, taken from images.json. But CI publishes to
ghcr.io/${{ github.repository_owner }}, so images built from a fork land under
the fork's namespace. Anyone testing an install from their own fork would have
had the plugin looking for images that do not exist there, with no indication
why beyond ten servers failing to start.
The registry is now a userConfig value defaulting to the upstream namespace, so
the common case is unchanged and a fork tester sets one field.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The lightweight image failed to build on arm64:
error: [Errno 2] No such file or directory: 'cmake'
ERROR: Failed building wheel for BoltzTraP2
BoltzTraP2 arrives via amset, and conda-forge publishes boltztrap2 for linux-64
only, so on aarch64 there is no prebuilt route and it has to compile. cmake and
pkg-config join build-essential in the same RUN that creates the environment,
and are purged alongside it, so the toolchain still does not reach the published
image.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The request stream was a heredoc with nothing after it, so stdin closed the instant the last request was written. An MCP server reads EOF as a shutdown signal and can exit before flushing its tools/list response, which the test then reports as "no response to tools/list" -- failing a healthy image. Caught by driving the same request stream against base_server running directly from its conda environment: without a grace period the responses never arrived, with one the server answered `base_tools 1.27.0` and listed 9 tools. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The smoke test caught this on the first working image: all four servers died
identically at import.
ModuleNotFoundError: No module named 'mcp.server.fastmcp'
This is mcp 2.x, where FastMCP was renamed to MCPServer
Every server in src/mcp_server/ does `from mcp.server.fastmcp import FastMCP`,
which is the 1.x API. The environments on this machine hold mcp 1.25 to 1.27
because they were built when that was current, but the specs asked for an
unpinned `mcp`, so a fresh resolve now takes 2.x and every server fails to
start.
This was not only a container problem. Nine core_env.yaml files left mcp
unpinned, so `install.sh` on a clean machine has been producing environments
where no MCP server can import -- the documented local install path, quietly
broken for anyone starting fresh. All nine are pinned here alongside the
container spec.
The GPU images were never exposed to this: they install from lockfiles with
exact pins (mcp 1.25.0 through 1.27.2) and --no-deps, which is the property
that motivated using locks for them in the first place.
Also widens the CI smoke test from `base` alone to every server the image
declares, read from images.json. Testing one server would have caught this
particular failure, but it is a mode where servers can break independently.
Verified on the rebuilt arm64 image: mcp 1.30.0, and all four servers complete
the handshake -- base 9 tools, drugdisc 5, smol 7, atomate2 7.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Upstream kept developing mat-elasticity under the old `.agents/skills/` path while this branch turned that into a symlink, so git could not reconcile the two: a directory containing modified files against a symlink. The conflict surface was small -- two skill files and one test. Upstream turns out to be a strict superset of this branch's mat-elasticity work. Its 8c04807 and d9a7e7a are this branch's 88d846a and 6ca8899 (the relaxed-ion default and the threshold split), and it adds two further commits refactoring onto pymatgen's ElasticTensor and ComplianceTensor, which is why the script shrinks from 499 lines to 440. Upstream's content is therefore taken wholesale, relocated to skills/mat-elasticity/, and put through the same `.agents/skills` -> `skills` path rewrite applied to the rest of the tree: four references across the SKILL.md and the test. The .agents/skills symlink is restored at mode 120000. Verified afterwards that it resolves, that skills/ still holds all 130 skills, and that the merged test file passes against the merged script: 14 passed, 1 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
First CI run: lightweight built on both arm64 and amd64; fairchem, mace and generative all failed, for two unrelated reasons. nvalchemi (fairchem, mace, matgl). The lock asked for `nvalchemi @ git+https://github.com/NVIDIA/nvalchemi-toolkit`, so pip cloned the repository, generated its metadata and then refused it: WARNING: Generating metadata for package nvalchemi produced metadata for project name nvalchemi-toolkit. Fix your #egg=nvalchemi fragments. ERROR: No matching distribution found for nvalchemi The project is named nvalchemi-toolkit, not nvalchemi. It is also published on PyPI at the installed version, so the lock now pins nvalchemi-toolkit==0.1.0 directly -- no clone, no name mismatch, and consistent with nvalchemi-toolkit-ops which already came from PyPI. Both pins verified to resolve. torch-scatter (generative). The extensions were frozen into the pip lock, so the pip layer tried to build them in the same pass that installs torch: ERROR: Failed to build 'torch-scatter' when getting requirements to build wheel Their build backends import torch to produce metadata, and torch is not importable yet at that point. They are now excluded from the locks entirely and left to the Dockerfile's dedicated step, which runs after torch is installed and uses --no-build-isolation. That step already existed; the lock was duplicating and pre-empting it. The exclusion is scoped to the three envs that carry them, which are exactly the three flagged pyg_from_source. Also makes the manifest job resilient. It ran `needs: build` with no condition, so three failing GPU builds skipped it entirely and the lightweight image -- which built cleanly on both architectures -- never got its multi-arch manifest list, leaving it reachable only through per-arch tags. The job now runs with always() and stitches each image whose per-arch tags are all present, skipping the rest instead of failing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…d time The amd64 lightweight image built and pushed cleanly, then failed its smoke test: atomate2 answered initialize but its tools/list did not arrive inside the 8 second grace period. Nothing was wrong with the image -- atomate2's imports are slow and the runner was cold. A fixed sleep encodes an assumption about machine speed that will keep being wrong. stdin is now held open through a FIFO and closed as soon as the tools/list response appears, with REPLY_DEADLINE (120s) only as a backstop. This is both more robust and faster: locally the four lightweight servers now complete in 8.2s total, where the fixed sleep cost 8s per server. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The marketplace is named atomistic-skills, so the install target is atomistic-skills@atomistic-skills. The README said atomistic-skills@atomisticskills, which resolves to no marketplace. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A clean-install test on an MIT ORCD compute node failed at every MCP server.
The node has no docker and no podman -- it has apptainer 1.4.5, which is the
normal situation on a cluster: users have no root and there is no daemon. Since
this software is aimed squarely at people running simulations on clusters, a
Docker-only plugin is unusable by much of its audience.
plugin.json previously spelled out a `docker run` argument list. That cannot be
made to work for Apptainer, whose arguments differ structurally: `exec` instead
of `run`, `--bind` instead of `--volume`, `--pwd` instead of `--workdir`,
`--nv` instead of `--gpus all`, a `docker://` transport prefix, and no
ENTRYPOINT handling under `exec`. Servers now launch through
docker/run_server.sh, which selects and translates for docker, podman,
apptainer or singularity; plugin.json passes configuration through the
environment. Two Apptainer details are handled: APPTAINER_CACHEDIR is pointed
at the model-cache directory so the SIF conversion does not land on a small
home quota, and the privilege drop in the entrypoint is naturally inert because
Apptainer already runs as the invoking user.
The test also produced a genuinely baffling symptom worth defusing:
Failed to connect - ENOENT: Executable not found in $PATH: "stdio"
That is Claude Code's error sanitiser replacing the missing command name, so a
plainly absent docker binary reads as a protocol fault. The tester had to
disassemble the CLI to work it out. The launcher now checks for the runtime
first and says which binary is missing and which runtimes the host does have.
Also forwards credentials the servers need -- MP_API_KEY, HF_TOKEN and the
literature API keys -- from the host environment when set, and only when set,
so an unset variable does not arrive as an empty one. Their absence was the
reason Materials Project tools would have failed even on a working install.
Documents the non-interactive install form (claude plugin install does not
prompt in a non-interactive shell, it just leaves the options unset), the
Apptainer path, and the "stdio" symptom.
Two other findings from the same report need no code change. The skill count is
129, not the 130 I had predicted: the 130th is private-mit-literature-qa, which
is gitignored by design. And `claude plugin details` reporting `MCP servers (0)`
is cosmetic -- that inventory counts servers declared via an external file
while this plugin declares them inline, and `claude mcp list` shows all ten
correctly registered. Both are noted in the troubleshooting section.
The published images are unchanged: run_server.sh executes on the host, not
inside the container, so 1.3.4 images remain correct and are not re-tagged.
Verified locally that the launcher starts drugdisc through the published 1.3.4
image and completes the MCP handshake, and that both failure paths -- absent
runtime, unsupported runtime name -- produce an actionable message.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The first Apptainer attempt got further -- install, launcher, and skill registration all worked -- but every server still failed. A retest on an MIT ORCD node found three distinct causes, none of which were visible from a Docker-only machine. mksquashfs thread exhaustion. The node has 448 cores and `ulimit -u` of 768. Apptainer runs mksquashfs with one thread per core, so the SIF build died with `FATAL ERROR: Failed to create thread` before any image existed. The launcher now derives a bounded `-processors` from the actual process limit rather than the core count, via APPTAINER_MKSQUASHFS_ARGS, overridable with ATOMISTIC_SQUASHFS_PROCS. Conversion cost against a 30 second connect timeout. Claude Code probes every server in parallel and allows each 30s; converting a 3.5 GB OCI image takes minutes, so all ten timed out, and because the probes are concurrent, four images were converted simultaneously -- 29 GB of quota consumed before anything worked. Images are now built once into a cached SIF under a flock and exec'd from there, and docker/prepare_images.sh does that build deliberately and sequentially before Claude Code ever starts. Architecture waste. That 29 GB was largely spent downloading arm64-only images onto an x86_64 host, which could never have run them. images.json already knows each image's platforms, so the launcher now receives them and refuses a mismatch outright: no download, and a message saying which architecture the image is for and that the other servers still work. Also sets APPTAINER_TMPDIR beside the cache, since the node warned that /tmp is mounted nodev and that this can corrupt a build. Adds tests/test_docker_launcher.py, 19 tests. The launcher decides *what command to run*, which is testable without either runtime: a stub on PATH records its argv and the tests assert on the decision. That covers the Apptainer branch -- which cannot otherwise be exercised on this machine -- including that a rejected architecture invokes no runtime at all, that exec uses the cached SIF rather than re-resolving docker://, that a failed build is not followed by exec'ing a nonexistent SIF, that --nv is used instead of --gpus, and that mksquashfs is bounded. It also pins plugin.json against images.json so the generated manifest cannot drift. Verified locally: 19/19 tests pass; the four lightweight servers still complete the MCP handshake through the rewritten launcher against the published 1.3.4 image; prepare_images.sh pulls and reports correctly; and the architecture gate, exercised against real docker by declaring an amd64-only image on this arm64 host, refuses without invoking the runtime. Still unverified: no Apptainer exists on this machine, so the arguments it receives are asserted by tests but have never been handed to Apptainer itself. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Retest learningmatter-mit#3 confirmed the previous round: 29 GB of wasted pulls became 2.3 GB, every connection timeout disappeared, and the architecture gate refused the six arm64 servers in under 10 ms. It then found four more, two of which broke tool calls outright. The pre-build was being silently discarded. prepare_images.sh defaulted its cache to $HOME/.cache/atomisticskills while plugin.json points the launcher at ${CLAUDE_PLUGIN_DATA}/model-cache. A user following the documentation built the SIF, watched it succeed, and then still hit on-demand rebuilds and timeouts because nothing looked where it had been put. prepare_images.sh now resolves the installed plugin's own data directory and says which cache it chose. create_research_dir could not work in a container. research_utils derived its root from __file__, which inside the image is /opt/atomisticskills -- read-only under Apptainer's SquashFS, so every call died with "[Errno 30] Read-only file system: '/opt/atomisticskills/research'". A workspace_root() helper now prefers ATOMISTIC_WORKSPACE, which the entrypoint sets to the mounted /work. This is one change in one module but it covers seven servers, all of which route through it. Relative output paths wrote into the image. The entrypoint ran `cd "$REPO_DIR"`, so a tool handed output_file="descriptors.json" tried to write to the read-only rootfs; only absolute /work/... paths worked. It now runs from /work when that is writable and falls back to the repository otherwise, which keeps a bare `docker run` with no volume working. PYTHONPATH already made the repository importable, so nothing depended on it being the working directory. The architecture gate was unreachable without ATOMISTIC_IMAGE, because the mandatory-image check preceded it. The gate now runs first: an image this host can never execute is refused whether or not its full reference is known. Verified in-container against a rebuilt image, since neither of the runtime bugs can be caught by unit tests. compute_molecular_descriptors with the relative path that previously failed now writes 724 bytes to the host workspace, and create_research_dir creates /work/research/2026-09-14_ethanol_test owned by the invoking user rather than failing read-only. All four lightweight servers still pass the smoke test. Tests grow to 24, covering workspace_root's override and default, that a created research directory stays inside the workspace, and that the gate refuses before demanding an image reference. Version moves to 1.3.5: the entrypoint and research_utils fixes live inside the image, so 1.3.4 cannot carry them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
plugin.json carried version 1.3.5 while userConfig.image_tag still defaulted to 1.3.4, and sync_version.py checked only `version` fields, so --check reported "all manifests agree" while the drift sat in the one field that decides which images a default install actually pulls. This was not cosmetic. entrypoint.sh and src/ are baked into the image, so a default install paired 1.3.5 plugin code with 1.3.4 images -- shipping none of the workspace and architecture-gate fixes to anyone who did not pass --config image_tag explicitly. sync_version.py now owns that default, a regression test asserts it against VERSION, and the install examples stop hardcoding a tag altogether so the same staleness cannot reappear in the docs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Round 4 on the HPC node showed the pre-build still being ignored. The cause was not the glob pattern but the premise: prepare_images.sh tried to predict CLAUDE_PLUGIN_DATA, and Claude Code does not create that directory until the first session loads the plugin -- which is after the documented pre-build step. So detection failed every time on a fresh install, the SIF went to ~/.cache/atomisticskills, and the node sat through four 30s connect timeouts with a valid 1.2 GB SIF already on disk. Guessing harder at that path cannot fix it; the directory genuinely does not exist yet. Invert the responsibility instead: run_server.sh now checks the model cache, then ATOMISTIC_SIF_DIR, then the shared fallback, and uses whichever holds the SIF. prepare_images.sh keeps its preference order but no longer reports the fallback as a problem, because it is now correct. Three tests cover it, including one that reproduces round 4 exactly -- it fails against the previous launcher with "no cached SIF ... building". Also documents Claude Code's ~15 minute failed-connection cache, which made the earlier retries look like the fix had not worked. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Round 5 verified the launcher's SIF search but exposed the other half of the same problem: prepare_images.sh had no such search, so it rebuilt a 1.2 GB image it already had -- 15m12s on the HPC node, for nothing. The two scripts resolve the cache differently across runs, and that is not a bug to remove: the first pre-build lands in the shared cache because the plugin's data directory does not exist until Claude Code has run once, and on the next run that directory does exist and becomes the target. So the same SIF is correctly present at a path the pre-build was not looking at. Give prepare_images.sh the search run_server.sh already has. It now reports "have <image> in <dir>" and skips the build. Two tests cover it; the first fails against the previous script with "build lightweight". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Six of ten servers were arm64-only, so an x86_64 cluster -- the ordinary case for HPC -- got four working servers and six architecture refusals. This brings mace, matgl and fairchem to amd64, leaving only the three generative servers arm64-only. The blocker was assumed to be that these environments cannot be resolved. That is true on aarch64 and false on x86_64, and the difference is the whole point: fairchem-core declares torch~=2.8.0 and the aarch64 build had to override it because torch 2.8 ships no sm_121 wheel. On x86_64 that constraint is satisfiable -- the environment resolves cleanly to the declared 2.8.0. The PyG extensions likewise have x86_64 CUDA wheels, so the amd64 build skips the source compilation that costs arm64 an hour. So the two architectures are locked by different means, and honestly: aarch64 is frozen from the validated environments (export_locks.py), amd64 is resolved from the declared specs (resolve_locks.py, new). Versions are resolved per architecture rather than forced to match, because making amd64 adopt aarch64's overrides would import constraint violations that only ever existed because of sm_121. The mace/fairchem split stays: mace-torch 0.3.15 pins e3nn==0.4.4 against fairchem-core's e3nn>=0.5, which is architecture-independent. Build settings that now differ per architecture (GPU target list, PyG source build, pip index) accept either a scalar or a platform mapping, so the scalar form stays valid wherever nothing differs. Preflight follows each image's declared platforms instead of assuming aarch64. Generative stays arm64-only: adit-agent's spec declares only python, pip and uv, so there is nothing to resolve and its amd64 lock must be derived from a built environment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…age be inspected Two findings from the HPC and GPU rounds. The thread cap was not enough. Round 7 hit "FATAL ERROR: Failed to create thread" at -processors 8 on the 448-core login node the cap was written for. `ulimit -u` counts every process the user has on that node, not just this build, and mksquashfs spawns several threads per -processors unit, so no constant is safe -- the headroom depends on what else the user happens to be running. Cap at 4 and, when that specific error appears, retry single-threaded. Other build failures are still fatal, so genuine problems are not masked by a retry loop. The entrypoint made the image hard to inspect. It takes a server name, so `docker run IMG micromamba run -n mace-agent python -c ...` -- the obvious way to check what a stack contains -- answered "unknown server 'micromamba'". The GPU tester lost time to that and had to discover --entrypoint. A name that resolves to an executable now falls through to exec, announced on stderr so a mistyped server name can never silently become some other command; anything else still errors, now pointing at --entrypoint. Seven tests cover both, including one asserting a non-thread failure is not retried. All four new tests fail against the previous scripts. Note: entrypoint.sh is baked into the images, so that half takes effect only on the next image build. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The plugin shipped 129 skills and ten working MCP servers, but 112 of the
131 skills do their real work through `# Env: <conda-env>` scripts, and
nothing in the container path can run those. A plugin-only user reached
step 3 of a perfectly good SKILL.md and got EnvironmentNameNotFound. Only
19 skills worked end to end.
Three uv projects replace nineteen conda environments for everything
except the generative stack. `uv run --project venv/<p>` is
self-contained: no activation, and uv builds the environment from the lock
on first use.
Adapted from AtomisticPi's venv layer, with the parts that did not
transfer:
- it pins every project to x86_64 only, which would have undone the
multi-arch work; the markers here cover x86_64 and aarch64, and
`uv lock` resolves universally across both
- it has no fairchem, sidestepping the e3nn conflict rather than
solving it; fairchem gets its own project and an
`override-dependencies` entry recording why the declared torch range
is violated
- it has no matgl; that resolves cleanly alongside mace-torch
Verified before writing anything: every project resolves for both
architectures, and torch ships CUDA wheels for both, so the custom
PyTorch index the conda locks needed is no longer required.
Two packages genuinely have no aarch64 wheel and carry markers instead of
being dropped for everyone: pymol-open-source, and SCINE (so the ORCA
skills stay x86_64, which is moot anyway since ORCA is a licensed binary
the user supplies).
The generative stack is deliberately excluded. mattergen hard-pins
torch==2.2.1+cu118 wheels that never existed for aarch64, its PyG
extensions have no wheel for that combination, and ADiT and DiffCSP++ are
not published packages at all. It keeps the existing container path.
No skill has been migrated yet; that is the next commit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Rounds out the migration so no skill is left without a home. Fourteen more
conda environments land in the two existing projects:
cpu <- nmr, phasefield, calphad, xrd, drugmd, orca, atomistic (VOID)
mlip <- scd, react-ot, ms-gen
Three of these are not on PyPI and come from git, pinned to a commit. Two
have a same-named but unrelated package on PyPI, so using the registry
name would have silently installed the wrong software:
VOID learningmatter-mit/VOID ("void" on PyPI is an unrelated
"Void object in Python")
ms-pred coleygroup/ms-pred (ICEBERG; unpublished)
oa-reactdiff deepprinciple/react-ot (React-OT; the import is
`reactot` but the distribution
is named oa-reactdiff, which is
why requesting `reactot` fails)
Both projects still resolve universally across x86_64 and aarch64:
cpu 241 packages, mlip 336.
Only the generative stack remains outside, for the reasons already
recorded in venv/README.md. Its two script call sites are the sole
remainder; its other skills drive the MCP servers rather than scripts, so
they are unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Skills annotated a block `# Env: mace-agent` and then ran a bare
`python script.py`, which only works for someone who had already activated
that environment. A plugin user cannot: the containers expose MCP servers,
not conda environments. 112 of 131 skills failed at their first script
step with EnvironmentNameNotFound, so only 19 worked end to end.
Every invocation now carries its environment:
# Env: mace-agent # Venv: venv/mlip
python skills/x/scripts/y.py -> uv run --project venv/mlip python skills/x/scripts/y.py
347 invocations across 112 skills, done by tools/migrate_skill_envs.py
rather than by hand. The tool works on the fenced block rather than the
next line, because blocks legitimately open with `cd`, `export`, an inline
`MP_API_KEY=...` or a `for` loop and the command that needs the environment
comes further down. It is idempotent: re-running reports no changes.
Three things the first pass got wrong, each caught by measuring rather
than assuming:
- eight annotations carry prose after the env name ("# Env: xrd-agent.
Quote the path because of (PO3).") and were skipped by a regex that
required end-of-line. Their prose is retargeted too, so a note reading
"(or matgl-agent)" no longer points at something that is gone.
- 22 `conda run -n` / `conda activate` lines sat in blocks with no
annotation at all, so an annotation-driven pass never saw them. They
are just as broken for a plugin user.
- replacing `conda activate` created new annotations whose commands were
then left bare. The tool now also acts on existing `# Venv:` blocks,
which is what makes repeated runs converge.
Also fixes mat-elasticity, which arrived from upstream with a top-level
`category` key. That is not one of the six keys the Agent Skills spec
allows and hard-errors on the Skills API -- the same breakage the move to
`metadata:` was meant to end.
Untouched: the two ml-generative-diffcsp blocks and the mattergen
activation line, which stay on the container path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…i under load Three additions, each a standard piece of elastic characterisation that was being re-derived by hand every time, and each with a failure mode worth centralising. Shear-modulus extrema over all shear systems, from G(n,m) = 1/(4 S_ijkl n_i m_j n_k m_l) with m in the plane normal to n. A genuinely two-dimensional search -- over the sphere and over the angle within each plane -- where Young's modulus needs only the sphere, which is why it tends to get skipped. min(C44,C55,C66) is not a substitute: on Pnma CaMgSi it is 26% high (36.86 against 29.29 GPa). Acoustic branch velocities along --acoustic_direction, via pymatgen's green_kristoffel. The Christoffel eigenvalues give one quasi-longitudinal and two quasi-transverse branches, and in an anisotropic crystal the transverse pair is not degenerate: along [101] in CaMgSi they are 3615 and 4202 m/s against an isotropic transverse velocity of 4141, so the slow branch is 12.7% below what the isotropic moduli imply. Verified against an independent hand-rolled Christoffel construction to 1e-9. --pressure relaxes cell and ions against a hydrostatic load before the scan. Two traps handled rather than documented away: ASE's FrechetCellFilter wants scalar_pressure in eV/A^3 while the flag is in GPa, so passing GPa straight through applies ~160x the load; and matcalc's own pre-relaxation is at zero pressure, so the loaded relaxation happens here and relax_structure is then disabled, otherwise the scan is silently re-centred on the unloaded cell. What is reported are the stress-strain coefficients about the loaded reference, not the Birch coefficients carrying explicit pressure corrections -- a different quantity, and the one wanted for stability under load. On CaMgSi: B 50.80 -> 68.59 -> 84.89 and G 38.49 -> 47.33 -> 53.41 GPa at 0/5/10 GPa, loaded volumes matching an independent implementation exactly. Tested analytically against an isotropic solid, the check with teeth here: both shear extrema must collapse to G, and the two transverse branches must be degenerate and equal sqrt(G/rho) in every direction. Companion tests confirm the anisotropic case splits the transverse pair and pushes the shear minimum below min(C44,C55,C66), so the isotropic tests cannot pass vacuously. 18 tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…chine The DiffCSP++ wrapper pointed at /home/bdeng/projects/DiffCSP-PP, so any other checkout failed. It now reads $DIFFCSP_REPO, falling back to a DiffCSP-PP clone next to this project, and raises a clear error only when the wrapper is used. The skill scripts no longer pre-set PROJECT_ROOT to the old path, which would have overridden the wrapper. Also drop machine-specific paths from mat-melting-point/get_features.py (resolve the project root from the file), the MgO charged-vacancy example (read/write next to the script; relative defect CIF paths, as generate_defect_structures.py emits them), and remove the orphaned test_exact_copy.py scratch script. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Replace every remaining /home/bdeng path with one resolved at run time or relative to where the file is used, and repair the scripts that could not run once their paths were exercised: - mat-disorder iterative_ce_training: relax_structures referenced undefined relax_input_dir/relax_output_dir/input_paths and a hardcoded mace-agent python. It now writes its inputs, runs relax_wrapper.py via the venv/mlip uv project, and reads relaxed_structure.cif plus relaxed_energy.txt back with the energy attached. --mlip_model is a MACE model name (default MACE-MP-medium); the mace/chgnet/m3gnet choices were never accepted by the MACE-only wrapper. Remove a stale duplicate copy of the pipeline appended to main(), import smol at module level so sweep_cutoffs can see ClusterSubspace/StructureWrangler, and add --supercell_size (default num_sites) instead of hardcoding O2-, which fails for non-oxides such as the skill's CuAg example. - ml-property-predictor: remove the absolute freeze_patch.py fallback and move the input-config dump out of the wrapper f-string, where its dict comprehension raised NameError before training could start. - ml-fairchem-finetune: restore run_dir and timestamp_id, lost in an earlier refactor (generate_fairchem_config.py NameError). - atomate2_utils: fall back to vasp_std on PATH instead of a fixed binary. - core_env.yaml: install nvalchemi-toolkit from a sibling checkout (-e ../../../nvalchemi-toolkit); the lock tooling still maps it to the PyPI pin for images. - example_full_env.yaml: drop the exported prefix lines. - Docs and example configs: relative links, image paths and data paths. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…kage Replace the per-server conda environments with three uv projects under venv/: cpu (no torch), mlip (MACE, MatGL >= 4 on PyTorch Geometric) and fairchem. They cannot merge: mace-torch pins e3nn 0.4.4 while fairchem-core needs e3nn >= 0.5 and torch ~= 2.13. Everything else resolves to its latest release; floors keep the resolver from falling back to fairchem-core 1.0 or nvalchemi 0.1. Extras hold packages with system requirements (openmm, pymol, docking, void, transport). The repository root is now an installable package (atomisticskills, packages=["src"]) installed editable into every project, so the ~100 scripts that import src work from any directory. lock_platforms.py derives venv/platforms.tsv -- the oldest glibc each (project, extra, arch) installs on -- from the locks, for the launcher to decide between uv and a container. uv.lock is exempt from the large-file hook. The conda lockfiles of the replaced environments are removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
venv/run <env>[+extra] <cmd> runs a command in a uv project, and venv/run --server <name> starts an MCP server, so skills, the plugin and configure_mcp.py share one entry point. The backend is uv on the host when the host can install the environment (Linux, glibc floor from venv/platforms.tsv, a compiler for a first source build), otherwise the container image built from the same lock (docker, podman, apptainer or singularity), with host paths mounted at the same locations so outputs land where scripts expect them. --setup installs, --doctor diagnoses. Configuration comes from ~/.config/atomistic_skills.yaml or the plugin's user config (runtime, image registry, image tag). Results go to the user's project: ATOMISTIC_WORKSPACE, else the repository for a checkout, else the current directory -- never the plugin cache. configure_mcp.py writes launcher-based server entries from venv/servers.tsv; mcp_config.json is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
python -m src.mcp_server.cli <server> <tool> key=value ... runs the same FastMCP tool a connected server would, so skills that use MCP tools keep working when no server is connected (a plain clone, another agent, CI). Several tools in one command share a process, so a model loaded by load_model stays loaded for the next call. --list prints tools with their arguments; output is JSON on stdout, library chatter goes to stderr, and a failed tool exits 1. tools/mcp_smoke.py drives real MCP sessions through venv/run --server for end-to-end checks, and treats JSON error payloads as failures. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The uv projects replaced eleven conda environments (base-agent, atomate2-agent, drugdisc-agent, smol-agent, nmr-agent, phasefield-agent, calphad-agent, xrd-agent, orca-agent, drugmd-agent, void-agent); they are removed, and remain in the 1.x history. What stays is used: the generative stacks (the image builds from their locks), ICEBERG, React-OT, the SCD example runs, and the three environments mat-lammps-md builds LAMMPS in. conda-envs/README.md says which and why. The atomate2 remote-worker guide moves to docs/atomate2_remote_workers.md. Skills that still pointed at removed environments (NMR, XRD, the spectrum matcher's ORCA notes) name the uv environment instead; the XRD refinement skill's kaleido advice is updated for kaleido >= 1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The generative servers (adit, diffcsp, mattergen) and React-OT leave conda. - venv/adit, venv/diffcsp, venv/mattergen reproduce the environments these models were verified in: every package pinned to that version and also an override, a closed set like the original `pip install --no-deps` (mattergen and mattersim get empty dependency metadata, since their own pins -- torch 2.2.1+cu118, numpy<2 -- cannot hold on current GPUs). torch, torchvision and the PyG extensions come per CUDA build (cu126 / cu130, picked by driver) from PyTorch's and PyG's indexes. x86_64 only: PyG has no aarch64 wheels, so venv/run uses the generative image there. - venv/reactot pins its verified set (torch 2.2.1: CUDA 12.1 on x86_64, CPU on aarch64, where the PyG extensions build from source against it). setup_react_ot.py fetches and patches the pinned upstream commit into a cache directory, download_models.py moves into the skill, and generate_ts.py imports React-OT from there (and answers --help without it). Verified on GB10: the oxadiazole example reproduces its stored TS (RMSD 0.000 A) and sits 0.020 A from the reference. - images.json gives each generative server its uv project; servers.tsv and the plugin manifest are re-rendered; venv_image maps a project to the image carrying it; lock_platforms reads each project's Python version, accepts plain linux_<arch> wheels and marks unbuilt architectures; the skill sweep skips commands whose environment is not built for the host; --setup without arguments builds only cpu, mlip and fairchem. - Skills run their scripts through venv/run <stack>; CI adds the stacks (generative on x86_64 x both CUDA builds, reactot on both architectures). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
chem-msms-predict no longer needs the ms-gen conda environment. venv/msms installs ms-pred (ICEBERG 2.1) from GitHub at a pinned commit, on a CPU torch 2.6 build with DGL 2.5 (x86_64 only: DGL has no aarch64 wheels). The skill was written against the ICEBERG 2.0 API while the 1.x install script cloned ms-pred main, so it could not run as shipped. Port it to the 2.1 API (ms_pred.iceberg, num_cpu_workers, CompositeMassSpec with atom-mask fragments turned into SMILES), fail loudly when the prediction subprocess fails, and add download_weights.py for the public MassSpecGym weights. The three source patches the conda install applied are fixed upstream. CI: an x86 msms job; the MCP smoke step skips environments that run no server (reactot, msms), which failed the reactot jobs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… example output Verified end to end on an x86_64 runner (ICEBERG 2.1, public MassSpecGym weights): precursor [M+H]+ 166.0863 Da, benzoyl 105.033 and phenyl 77.039 fragments for 2-aminoethyl benzoate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ml-property-predict-scd claimed the mlip environment, but the upstream code imports PyG's compiled extensions (torch_scatter in data/loaders.py, which train.py and the lightweight-head template load; torch_cluster and torch_sparse in the Frad models), which mlip does not carry. venv/scd installs SCD's requirements.txt with torch 2.10 (cu126 / cu130 by driver) and PyG's wheels for it, x86_64 only like the other PyG stacks. The upstream repository is cloned at a pinned commit into ~/.cache/atomisticskills (or $SCD_REPO_DIR), which the examples and templates now look in; the examples run train.py with their own interpreter instead of relaunching through `conda run -n scd-agent`. conda-envs/ keeps only the LAMMPS build environments and the arm64 generative image locks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Upstream checkpoints name their model config relative to the repository root (configs/model_configs/scd.yaml), so the templates failed unless run from inside the checkout. Load the model there; user paths keep their meaning. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rce checkouts into containers The ADiT wrapper located the AADT repository through PYTHONPATH, which only the 1.x conda MCP configuration set; under venv/run it fell back to paths relative to the working directory. Use $ADIT_REPO, else an adit checkout next to the project, like DiffCSP++. The generative image carries neither repository, and containers only saw the project, the workspace and the cwd. venv/run now passes ADIT_REPO and DIFFCSP_REPO on (or their defaults, when present) and mounts the checkouts at the same path. Skill prerequisites and docs/environment_variables.md describe the lookup; the stale ARM build notes are gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
All ten servers now have a uv project (or the generative image), so the fallback to the generative servers' conda environments, and --conda, are gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The rules, CLAUDE.md, venv/README, setup and developer guides now describe the research stacks' uv projects (adit, diffcsp, mattergen, msms, reactot, scd) instead of conda environments, and the migration guide and design notes record the ICEBERG 2.1 port, the SCD environment and the ADiT lookup fix. The verification matrix gains rows for the research stacks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MatterGen's PyPI distributions omit its data files (sampling_conf, the GemNet
scale factors, the training configs) because its setuptools config declares no
package data; upstream supports only an editable install from a clone. So
generation failed in the generative image and in the x86 uv environment
alike ("Primary config directory not found"), while 1.x hosts worked only
because their conda environment had an editable clone.
The wrapper and the fine-tuning script now import MatterGen from a checkout
($MATTERGEN_REPO, else a mattergen clone next to the project), and the
fine-tuning CLI gets it on PYTHONPATH. The fine-tuning example was broken too:
wrong script paths, a pymatgen import typo, the wrong JSON shape and an empty
training_data.csv; save_skill_inputs treated a .csv output as a directory.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The arm64 generative image now installs venv/adit, venv/diffcsp and
venv/mattergen side by side from their uv.lock files, like the other images,
and the conda locks (and docker/export_locks.py) are gone. Those projects are
locked for aarch64 too, with PyG's extensions built from source there (static
metadata, torch matched to the environment); tool.atomisticskills.native-arches
keeps aarch64 hosts on the image, since that build needs the CUDA toolkit.
The published image's extensions had no CUDA kernels: the build sees no GPU,
and torch-scatter and friends then silently compile CPU-only ("Not compiled
with CUDA support" at the first GPU call). Dockerfile.cuda sets FORCE_CUDA=1
and fails the build unless each extension reports a CUDA version.
An image may carry several environments: venv/run passes ATOMISTIC_VENV for a
command, the entrypoint puts that environment on PATH, and Apptainer commands
now go through the entrypoint. The MatterGen checkout is mounted like the ADiT
and DiffCSP++ ones.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The last use of conda. Each project gains a lammps extra that images leave out: mlip's has what building LAMMPS against it needs (Cython for the ML-IAP Python coupling; on x86_64 the MKL headers PyTorch's CMake config asks for), fairchem's the official LAMMPS wheel with its MPICH runtime and fairchem-lammps (lmp_fc). scripts/build_lammps_mace.sh (ACEsuit's ML-MACE, C++20 for torch >= 2.13) and build_lammps_matgl.sh (Kokkos + ML-IAP, LAMMPS stable_22Jul2025_update4) replace conda-envs/*/install_lammps.sh and build into ~/.cache/atomisticskills/lammps. conda-envs/ is removed. Examples: the MatGL one installed numpy<2 and chgnet with pip into the environment, and now uses MatGL's CHGNet; the FairChem one laid CO flat on the surface (E_ads +1.3 eV) and now places it carbon-down on the top site with the oc20 head (-0.39 eV unrelaxed); lmp_fc needs the MPICH library on the loader path. Verified on GB10: ML-MACE build, 60 MD steps on CUDA; lmp_fc with UMA. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rules, agent instructions (CLAUDE/AGENTS/GEMINI), the MOF screening workflow, the landing page's environment tables (and its server count, which counted conda-envs/ directories) and test docstrings now name uv projects. The migration tool drops its conda path, and the skill tests now reject any conda command or metadata.conda_env. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Under uv the server tests never ran. skip_if_wrong_env keyed on CONDA_DEFAULT_ENV, so every marked test skipped, and --strict-markers knew only six of the eleven markers, so the adit, diffcsp, mattergen, smol and drugdisc test files failed to collect. Markers now map to uv projects (detected from sys.prefix, or ATOMISTIC_VENV in an image); ORCA tests that need SCINE skip where it has no wheels (aarch64). Running them found two bugs: - drugdisc.compute_molecular_fingerprints failed whenever no heatmap was requested (heatmap_path unbound). - matgl.load_model rejected six pretrained models under torch >= 2.6 (QET, TensorNet "-m", M3GNet/MEGNet Eform): their checkpoints pickle numpy arrays or a Normalizer. The wrapper now allowlists exactly those; the load test goes through the wrapper, as the server does. GB10: cpu stack 86 passed, 46 skipped; mlip 57 passed; fairchem 5 passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e, test fixes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tion values Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ch at run time - No image had libXrender (and the CUDA base image no X11 at all), so `from rdkit.Chem import Draw` failed in containers: ADiT, and every drug skill that draws, broke in container mode. The images now install the X11 client and render libraries, and the build checks that import. - The CUDA base image puts its toolkit on LD_LIBRARY_PATH, which shadowed the CUDA libraries of each environment's torch wheels: DiffCSP++'s torch 2.10 got the toolkit's cuBLAS 13.0 instead of its 13.1 (CUBLAS_STATUS_INVALID_VALUE in sgemm). The toolkit is for compiling the PyG extensions only, so the runtime library path leaves it out. Verified on GB10 with the uv-built generative image: MatterGen (4 structures, 70 s), ADiT (4 crystals) and DiffCSP++ (Li2ZrCl6 in P-3m1, as constrained) generate on the GPU, and radius_graph / scatter run on CUDA in all three. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ed on GB10 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t for the stacks From the x86_64 A100 test round, where fresh clones failed while the dev box worked only thanks to uncommitted patches in its checkouts: - MatterGen calls np.math.factorial, which NumPy 2 removed (the verified environment runs NumPy 2.2). setup_mattergen.py clones the verified commit (94441ee, LFS checkpoints skipped) and uses math.factorial; the skill and the wrapper point at it. - ADiT imports OpenBabel when it builds the model; the 1.x conda env had it. venv/adit pins openbabel-wheel 3.1.1.23 (x86_64 and aarch64 wheels), and the skill pins the verified ADiT commit. - The generative stacks get a dev group with pytest, so their server tests run on hosts (images still leave it out). - DiffCSP++ checkpoints: where to get them (Google Drive, by hand). - build_lammps_mace.sh checks up front that the CUDA toolkit is at least as new as torch's CUDA build, which PyTorch's CMake config requires. Verified on GB10: MatterGen generates from a fresh setup_mattergen.py checkout. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…GPUs, measured glibc floors Code review at fee3006: - +openmm carries the OpenFF stack (toolkit, interchange, units, utilities, force fields from their release tags, since the PyPI releases are yanked) and openmmforcefields; --charge_method mmff94|gasteiger uses RDKit when AmberTools' sqm is absent, and AM1-BCC/GAFF without it explain why. - sella in mlip and fairchem; load_calculator(model_name=None) no longer crashes. - MatGL property models build the graph on the model's device. - venv/run passes CUDA_VISIBLE_DEVICES to docker/podman, names SIFs by architecture and registry, keeps extras after a CUDA extra and refuses two. HPC-EL8 round 4: - PyG wheels for torch 2.9/2.10 need glibc 2.32 and DGL 2.5 GLIBCXX_3.4.26 despite their tags: tools/scan_wheel_floors.py measures them and lock_platforms.py applies the floors; the generative image is built for amd64 as well, so EL8 hosts get it. - UV_LOCK_TIMEOUT defaults to an hour; mcp_smoke waits past the setup window. - setup_mattergen.py, which the skill already names, is now tracked. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… EL8 results Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…an env failed uv >= 0.12 (what setup-uv installs) does not extract python-constraint 1.4.0's bzip2 sdist, its only release, so cpu+openmm failed on fresh runners while warm caches hid it. Build 1.4.0 from its tag instead. check_skill_commands.py printed only the last line of each failure and discarded the environment-creation output; it now prints the latter. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On EL8, msms needs only a newer libstdc++ (module load gcc), which the glibc check cannot see; the refusal now names ATOMISTIC_RUNTIME=uv. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Installing AtomisticSkills used to mean building ~20 conda environments (~58 GB) by hand. 2.0.0 replaces them with three uv projects and a single launcher,
venv/run, so that a Claude Code plugin install or a plain clone works on x86_64 and aarch64 with no manual environment setup. Every skill, skill script and MCP tool works after install. This release is not backward compatible; seedocs/changes/2.0.0-migration.md.What changed
Environments. Three uv projects with committed locks:
venv/cpu: the materials, chemistry and drug-discovery stack, without torch.venv/mlip: MACE and MatGL ≥ 4 (PyTorch Geometric).venv/fairchem: FairChem.Each tracks the latest releases. They stay separate because of real conflicts: mace-torch pins e3nn 0.4.4, while fairchem-core needs e3nn ≥ 0.5 and torch ~= 2.13. Packages with system requirements are optional extras:
openmm,pymol,docking,void,transport.Launcher.
venv/run <env>[+extra] <cmd>runs a command, andvenv/run --server <name>starts an MCP server.venv/platforms.tsv), and a compiler for first builds. Otherwise it runs the container image built from the same lock (Docker, Podman, Apptainer), with host paths mounted unchanged.--setupcreates the environments and--doctordiagnoses the host.GPU builds.
mlipandfairchemlock two torch builds, CUDA 13 (driver ≥ 580, required by GB10/Blackwell) and CUDA 12.6 (drivers 525–579).venv/runpicks one from the driver;ATOMISTIC_TORCH_CUDAoverrides.Skills.
skills/(.agents/skillsremains as a link). Every command is self-locating (${CLAUDE_SKILL_DIR}/../../venv/run …), so it works from any directory and from a plugin install.metadata.venvrecords each skill's environments. MCP steps are writtenserver.tool, with a shell fallback (python -m src.mcp_server.cli <server> <tool> key=value) for agents without a connected server.general-atomisticskills-rulesskill carries the project rules for plugin installs, which do not loadCLAUDE.md.Plugin and images. The plugin's MCP servers start through
venv/run, withruntime,image_registryandimage_tagoptions; results go to the user's project. The imagescpu,mlipandfairchem(amd64 + arm64) are built from the uv locks;generative(arm64) remains on conda locks.Science stack.
CHGNet-PES-MatPES-PBE-1M-2026.9.Fixes found while testing:
--submit).--help.CUDA_VISIBLE_DEVICES.Shipped files contain no local paths, host names or personal project names.
Verification
The full matrix is in
docs/changes/2.0.0-verification.md.--helpof every script)-develpackages).github/workflows/uv-envs.yml):.github/workflows/build-images.yml): built from the locks, with an MCP smoke test and a PyMOL import check.Breaking changes and upgrading
See
docs/changes/2.0.0-migration.md:mcp_config.jsonare replaced, so rerunconfigure_mcp.py;skills/, and SKILL.md usesmetadata.venv;Known limitations and follow-ups
HF_TOKEN.After merge
Pushing to
mainbuilds and publishes the 2.0.0 images toghcr.io/learningmatter-mit. Then tagv2.0.0, following.agents/rules/release-standards.md.🤖 Generated with Claude Code