A reference for writing numerical, simulation and large-data code without freezing the machine. Includes a machine-probe script, per-domain memory formulas, and guard patterns for BLAS threading, chunking and long runs.
This started as a skill that was supposed to change agent behaviour — "make the agent do the arithmetic first". It was measured twice against its own baseline and failed to show a gain both times:
Round Result 2 (2026-09-16) delta −0.133 — the skill lost 3 (2026-09-18) delta +0.000 — no measurable effect (14/15 vs 14/15) In round 3, both arms recovered the hidden dimensions, stated the correct footprints, and wrote streaming code — without the skill. Five pre-registered predictions about which items would discriminate all failed. On well-specified tasks with a capable executor, this advice is redundant: the executor already does it unaided.
Use this as a lookup reference, not as a behaviour guard. The technical content is correct and checked; the claim that loading it makes an agent behave differently is not. See
evals/hard/RESULTS.mdfor the full result and what remains untested.The one component with a demonstrated, standalone use is
scripts/probe_env.py— it reports what the machine actually has available, which no amount of model capability substitutes for.
Most agent hosts apply no memory or CPU quota to the processes they spawn — no cgroup,
no job object, no ulimit, no nice. One wrong size estimate is enough to take the machine
down: a np.zeros(10**10), a dense pairwise distance matrix over 50 000 points, a
read_csv with no chunking on a 40 GB file. The kernel starts paging, the desktop becomes
unresponsive, and the run is usually killed anyway.
The host's job timeout protects the session from hanging; it does nothing to stop a child process from consuming all the RAM before that timeout fires. There is no backstop, so any guard has to live in the generated code.
- A machine probe (
scripts/probe_env.py) reporting available RAM, physical and logical cores, swap or commit pressure, and a concrete budget. - Peak-memory formulas as
elements × itemsize × simultaneous_copies, with per-domain tables for MD, FEM, CFD, Monte Carlo, dense linear algebra, FFT, imaging, ML, dataframes and graphs. - A strategy table keyed on estimated peak versus available RAM — run, chunk, or refuse.
- Guard patterns: BLAS/OpenMP thread caps that actually take effect, a best-effort address-space limit, chunked accumulation, bounded process pools.
- Cost reporting: estimated versus measured peak, thread count, and any precision or algorithm fallback applied.
Clone, then copy the directory into whatever skills folder your host reads. SKILL.md is
the entry point; the host discovers scripts/ and references/ from there.
git clone https://github.com/Diraw/resource-aware-compute.gitProfer — skills are loaded from a workspace's skills/ directory, and a plain directory
copy is enough:
# replace <workspace-slug> with your workspace
mkdir -p ~/.profer/agent-workspaces/<workspace-slug>/skills
cp -r resource-aware-compute ~/.profer/agent-workspaces/<workspace-slug>/skills/Development builds use ~/.profer-dev/ instead of ~/.profer/. There is no install step to
repeat: edits take effect on the next agent run.
Copying the directory is the entire install — do not stop at leaving it anywhere else
(for example an E:\agent-skills\ working copy). Profer only reads the workspace's
skills/ directory, so a repo sitting elsewhere is invisible: it never triggers, and the
failure is silent. After copying, confirm the skill appears in the host's skill list before
concluding it works.
The user-global library at ~/.profer/global-skills/user/ is not enough on its own.
Every entry there also needs a skill.manifest.json next to SKILL.md, and without it the
directory is skipped silently — a plain cp into user/ looks installed and does nothing.
Use the host's own "promote to global" action, or copy an entry that already has a manifest,
rather than dropping a directory in by hand.
Claude Code — personal skills under ~/.claude/skills/, project skills under
.claude/skills/:
cp -r resource-aware-compute ~/.claude/skills/Any other host — copy the directory into the folder your host documents for skills.
Two warnings that come from testing this rather than assuming it:
- If your host ships a marketplace CLI (for example
npx skills add), check where it actually installs and whether the host reads that directory. Some hosts read one specific folder and ignore everything else, so a successful install can still be invisible. - A skill that is not discovered fails silently — no error, no log, it simply never
triggers. Confirm it appears in your host's skill list before concluding it works. Running
python scripts/probe_env.pyon its own is a quick sanity check that the files arrived intact.
Two ways to use this, and only one of them is validated.
As a reference (recommended). Read SKILL.md and the references/ files when you hit a
memory question. This is a lookup, and it is the use the evidence supports.
As a skill (not validated). Installing it and relying on the description to make an
agent pick it up is the configuration that was measured twice and showed no gain. It may
still trigger on simulation and large-array work, but do not expect it to change what the
agent does — in round 3, the arm without it produced the same estimates and the same
streaming code. Asking directly still works if you want the workflow applied:
Use the resource-aware-compute reference to check whether this fits in memory before running it.
The probe script is the one part that pays for itself, and it is standalone:
python scripts/probe_env.py # human-readable report
python scripts/probe_env.py --json # machine-readable
python scripts/probe_env.py --quiet # one line, e.g. for pasting into a reply
python scripts/probe_env.py --self # also report this process's peak RSSOS : Windows 11 (win32, Python 3.13.14)
CPU : 12th Gen Intel(R) Core(TM) i7-12700H
14 physical / 20 logical cores (source: platform-api, hybrid)
RAM : 28.04 GiB available of 47.63 GiB total
Commit : 32.69 GiB used of 74.63 GiB (43.8%) [commit charge, not disk swap]
Budget for this machine (all figures GiB, 1024^3 bytes)
run directly below : 7.01 GiB peak
chunk / memmap above : 16.82 GiB peak
do not exceed : 25.24 GiB peak
recommended threads : 13 (OMP/MKL/OPENBLAS_NUM_THREADS)
recommended processes : 13
Compare your own estimate in GiB (bytes / 1024^3) against the bands above.
WARNING: Hybrid CPU (mixed performance/efficiency cores). A thread count near the
logical count will overload the efficiency cores.
All memory figures are GiB (1024³ bytes). Convert your own byte estimate the same way before comparing, or you will be off by ~7% in the direction that makes an over-budget run look safe.
Standard library only, no network access, no writes. Linux reads /proc/meminfo; macOS
reads sysctl and vm_stat; Windows reads GlobalMemoryStatusEx and
GetLogicalProcessorInformationEx through ctypes. psutil is used when present but is
never required. --self reads the current process's peak RSS via getrusage (POSIX) or
GetProcessMemoryInfo (Windows), so a run can be checked against the estimate that
justified it.
| Path | Purpose |
|---|---|
SKILL.md |
The workflow: probe, estimate, choose, guard, run, report. Correct as reference; the behavioural claim is unvalidated. |
scripts/probe_env.py |
Machine budget probe. stdlib only, cross-platform. --self for peak RSS. |
references/estimation-cookbook.md |
Peak-memory formulas per domain. |
references/python-numeric.md |
numpy / pandas / scipy patterns that keep the peak down. |
references/parallel-and-gpu.md |
Thread versus process, oversubscription, VRAM, the OOM ladder. |
references/long-running.md |
Checkpointing, resumable loops, progress, background execution. |
references/host-notes.md |
Host-specific shell timeouts and background support. Optional. |
evals/evals.json |
The round-1/2 test cases. Has a 15/15 ceiling — kept for the record. |
evals/hard/ |
A test set designed to have no ceiling. Contains its own README, fixture generator, validator, and RESULTS.md (round 3: delta +0.000 — it did not discriminate). |
Worth being explicit, because a skill is a prompt, not a sandbox:
- It cannot enforce anything. No memory limit, no CPU cap, no sandbox. If the model ignores the instructions, nothing stops it. This raises the odds of good behaviour; it does not guarantee it.
- It cannot watch a running process. It has no way to notice that a job is thrashing and intervene.
- It is not a substitute for the right algorithm. It will help you notice that a dense N² matrix is the wrong structure — it will not invent the neighbour-list algorithm for you.
- It is not a job scheduler. It knows how to shape work to survive a kill; it does not queue or retry anything.
If you need real isolation, that belongs in the host (containers, cgroups, Windows job
objects) or in the code itself via a bounded runner. The address-space limit in SKILL.md
is the closest thing to enforcement that runs entirely in user space, and it does not exist
on Windows.
evals/evals.json contains tasks with objectively checkable expectations. Run them in
pairs (with and without the skill) and the effect can be measured rather than asserted.
None of the prompts uses the words "memory" or "performance".
Results as of 2026-09-18. Round 3 has now been run (see below): delta +0.000. Every number in this section except round 3 was measured against pre-1.0.2 revisions, so those rows neither describe the current revision nor establish that it is better.
SKILL.md's Step 2 was rewritten, the sparse/N² errors were corrected, units were disambiguated, and thedescriptionwas shortened after those measurements were taken. Treat rounds 1–2 as a historical record of what earlier versions scored.What is still true without re-running: the benchmark has a ceiling (the unskilled arm scored 15/15, so the skill could only lose points), and n = 1 per configuration, so a single expectation flipping moves the total by 0.067. See What this benchmark can and cannot tell you.
A test set designed to have no ceiling (evals/hard/): unlabelled binaries, dimensions
recoverable only via a sidecar, small stand-ins so that loading everything looks harmless.
Both arms scored 14/15; the single missed item was the same for both and turned out to be
a rubric defect, not an answer defect.
The premise — that a headerless binary would stop a capable executor from recovering the
size unaided — is false for this executor. Both arms cat-ed the sidecar, derived the
real dimensions, stated the correct footprints, and wrote streaming code. The skill added
nothing measurable on these three tasks, and two of the five pre-registered predictions
about which items would discriminate failed outright (0 of 5 held).
Full result, prediction scorecard and analysis: evals/hard/RESULTS.md.
This is the second measurement in a row where the skill did not beat its baseline
(round 2: −0.133). The consistent reading is that on well-specified tasks with a capable
executor, the advice is redundant — see the RESULTS file for what that does and does not
establish.
Three tasks × two configurations, one run each, graded blind: the two outputs of each task were relabelled A/B with a per-task mapping the grader never saw. Host: Windows 11, 14 physical / 20 logical cores, 48 GB RAM, 27–30 GB available.
| Eval | with skill | without skill |
|---|---|---|
| 1 — trajectory radius of gyration | 4/5 | 5/5 |
| 2 — FEM von Mises post-process | 4/5 | 5/5 |
| 3 — Monte Carlo speed-up | 5/5 | 5/5 |
| total | 13/15 (0.867) | 15/15 (1.000) |
Removing the stated data size from the prompts did not rescue the result. The unskilled arm read the format note by itself, found the real dimensions, said "your ~2 GB estimate is off, this is 38.4 GB", and wrote memory-mapped chunked code. So did the FEM run: it read the HDF5 attributes, discovered 300 steps × 2.4 M nodes, and streamed one step at a time.
Both failures in the skill arm are worth reading individually, because they are different kinds of event:
Eval 2 was a real defect, and the skill caused it. The arm with the skill estimated its
own runtime peak — "about 1.3 GB" — and never the size of the stored dataset (17 GB, 5.8 GB
per component). The arm without the skill stated both. The cause was SKILL.md itself:
Step 2 said "Estimate the peak, not the final result", and the model dutifully reported the
algorithmic peak. SKILL.md has been rewritten so Step 2 demands two labelled numbers —
the data footprint and the implementation peak — and warns explicitly that "peak 1.3 GB"
does not answer "how big is my data". This is the one concrete, actionable thing this
benchmark produced, and it came from the skill losing.
Eval 1 is a grading artefact. The item asked for results written incrementally "rather
than accumulating all frames in a Python list". The losing arm preallocated an
np.empty(n_frames) array — 3.2 MB at 400 000 float64 — and wrote the CSV after the loop.
The grader flagged it as borderline and failed it on the strict reading. Under a literal
reading of the item text that arm passes, making the totals 14/15 vs 15/15, delta −0.067.
The item has since been reworded to say what it meant (an O(n_frames) result array is fine;
per-frame data is not). Both numbers are reported here; neither is presented as the answer.
Measured against pre-1.0.2. See the version note above.
The same three evals against an earlier revision whose prompts stated the data scale ("400000 frames", "2.4 million nodes, 300 time steps"). Result: 15/15 vs 14/15, delta +0.067, with the only discriminating item being the thread/worker bound on eval 3. A capable executor converts a stated scale into a streaming implementation for free, so those two evals were measuring almost nothing. The prompts were rewritten; round 2 is the result.
That version of the table is kept rather than deleted, because a benchmark repaired after seeing the answer is worth exactly as much as you admit it is.
Applies to rounds 1 and 2 (the evals/ set). Round 3 in evals/hard/ was written
specifically to remove the two limits below; its own caveats are in evals/hard/README.md.
- It has a ceiling. The unskilled baseline scores 15/15. A skill can only lose points
here, never win them. If you want to measure the upside, the input has to be less
self-describing — an unlabelled binary, no format note, no attrs recording the production
size — or the executor has to be weaker. This is what
evals/hard/addresses. - One run per configuration. With n = 1 there is no variance estimate, and a single expectation flipping moves the total by 0.067. Treat a delta below roughly 0.2 as unresolvable by this benchmark.
- It measures increment over a capable executor, not whether the advice is correct. The
advice in
SKILL.mdis still the advice; what these runs cannot show is how much of it a strong model would have discovered unaided. - It is not evidence that the skill is useless. It is evidence that on well-specified tasks with readable inputs, a strong executor already does the estimation and the chunking. Where the skill plausibly still earns its place: thread and worker bounds, host timeouts and background execution, atomic checkpoints, and refusing to run something that cannot fit.
If you run these, the honest number is more useful than a flattering one. Open an issue with the results either way.
Round 3 (evals/hard/) was run and returned delta +0.000: both arms scored 14/15, and
the one missed item was missed by both. Five pre-registered predictions about which items
would discriminate all failed. Documented in evals/hard/RESULTS.md.
Consequent changes, all of them removing claims rather than adding features:
- The README headline ("stops your coding agent from freezing your laptop") is replaced with a status banner stating the two null results. An unvalidated efficacy claim does not belong in the first line.
SKILL.mdcarries the same banner, because an agent reading it should know what it is holding.- The
descriptionno longer says "trigger whenever compute cost is non-trivial". That instruction was written to maximise triggering, and the measurement did not support it — extra context for no demonstrated effect. It now asks to be used as a reference. - The Usage section splits "as a reference" (supported) from "as a skill" (not validated),
and names
probe_env.pyas the one component with standalone value. - The H2 expectation that both arms failed was, on inspection, written wrong: it asked for chunking along the step axis, but the fixture is component-major with step fastest-varying, so step-axis chunking forces the strided random I/O both arms correctly avoided. Reworded to test the intent.
What was not done: adding features, widening triggers, or re-running the same test a third time. Two null results are enough to stop optimising this, and a third identical run would produce the same null (small stand-in fixtures never approach a real resource limit).
A review of the skill against its own claims found four factual or unit problems, all fixed:
A - Bon scipy sparse does not densify. BothSKILL.mdandreferences/python-numeric.mdasserted it does. Verified against scipy:A - B,A + B,A * B(elementwise) andA.T @ Ball stay sparse between two sparse operands. The operations that do densify are mixing in a dense operand (A * np.ones((n,n))),A.toarray(), and non-integer powers. Advice corrected in all three places, since the old wording would have deterred a correct sparse approach.- The
N²row conflated the trap with the fix. "Pairwise distances or neighbor lists" was listed asN²in one row; a cutoff neighbour list isN × n_neighboursand is the recommended structure. Split into two rows — the trap and the sparse alternative. - Units were ambiguous. The probe computes GiB but labelled everything
_gband printedG; a bare decimal-GB estimate compared against it is off by ~7%, which is enough to cross a threshold on its own.SKILL.md, the cookbook and the README now state the GiB convention, the JSON carries explicit_gibaliases plus aunitsfield, and the human output labels every band. The_gbkeys are kept for backward compatibility. host-notes.mdpresented host caps as facts. They are the most perishable content in the skill. The probing procedure now leads, the numbers are marked illustrative, and the section is labelled "verify before relying on them".
Two additions came out of the same pass:
probe_env.py --selfreports the current process's peak RSS, with the platform split (KiB on Linux, bytes on macOS) handled and the Windows path implemented viaGetProcessMemoryInfo. Theargtypes/restypedeclarations are required — without them ctypes marshals the 64-bit handle as 32-bit and the call returns 0 silently, which is exactly the class of quiet failure this skill exists to prevent. The flag is opt-in, so the default JSON schema is unchanged.- Step 2 and Step 3 carry explicit caveats: the R bands are decision aids, not verdicts, and R between 0.6 and 1.2 should be read as "chunk it" because the estimate itself carries roughly ±30% uncertainty.
MIT. See LICENSE.