Skip to content

About

A reference for writing numerical, simulation and large-data code without freezing the machine. Includes a machine-probe script, per-domain memory formulas, and guard patterns for BLAS threading, chunking and long runs.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

resource-aware-compute

A reference for writing numerical, simulation and large-data code without freezing the machine. Includes a machine-probe script, per-domain memory formulas, and guard patterns for BLAS threading, chunking and long runs.

Read this first: the efficacy claim is not supported

This started as a skill that was supposed to change agent behaviour — "make the agent do the arithmetic first". It was measured twice against its own baseline and failed to show a gain both times:

Round Result
2 (2026-09-16) delta −0.133 — the skill lost
3 (2026-09-18) delta +0.000 — no measurable effect (14/15 vs 14/15)

In round 3, both arms recovered the hidden dimensions, stated the correct footprints, and wrote streaming code — without the skill. Five pre-registered predictions about which items would discriminate all failed. On well-specified tasks with a capable executor, this advice is redundant: the executor already does it unaided.

Use this as a lookup reference, not as a behaviour guard. The technical content is correct and checked; the claim that loading it makes an agent behave differently is not. See evals/hard/RESULTS.md for the full result and what remains untested.

The one component with a demonstrated, standalone use is scripts/probe_env.py — it reports what the machine actually has available, which no amount of model capability substitutes for.

The problem this documents

Most agent hosts apply no memory or CPU quota to the processes they spawn — no cgroup, no job object, no ulimit, no nice. One wrong size estimate is enough to take the machine down: a np.zeros(10**10), a dense pairwise distance matrix over 50 000 points, a read_csv with no chunking on a 40 GB file. The kernel starts paging, the desktop becomes unresponsive, and the run is usually killed anyway.

The host's job timeout protects the session from hanging; it does nothing to stop a child process from consuming all the RAM before that timeout fires. There is no backstop, so any guard has to live in the generated code.

What is in here

  1. A machine probe (scripts/probe_env.py) reporting available RAM, physical and logical cores, swap or commit pressure, and a concrete budget.
  2. Peak-memory formulas as elements × itemsize × simultaneous_copies, with per-domain tables for MD, FEM, CFD, Monte Carlo, dense linear algebra, FFT, imaging, ML, dataframes and graphs.
  3. A strategy table keyed on estimated peak versus available RAM — run, chunk, or refuse.
  4. Guard patterns: BLAS/OpenMP thread caps that actually take effect, a best-effort address-space limit, chunked accumulation, bounded process pools.
  5. Cost reporting: estimated versus measured peak, thread count, and any precision or algorithm fallback applied.

Install

Clone, then copy the directory into whatever skills folder your host reads. SKILL.md is the entry point; the host discovers scripts/ and references/ from there.

git clone https://github.com/Diraw/resource-aware-compute.git

Profer — skills are loaded from a workspace's skills/ directory, and a plain directory copy is enough:

# replace <workspace-slug> with your workspace
mkdir -p ~/.profer/agent-workspaces/<workspace-slug>/skills
cp -r resource-aware-compute ~/.profer/agent-workspaces/<workspace-slug>/skills/

Development builds use ~/.profer-dev/ instead of ~/.profer/. There is no install step to repeat: edits take effect on the next agent run.

Copying the directory is the entire install — do not stop at leaving it anywhere else (for example an E:\agent-skills\ working copy). Profer only reads the workspace's skills/ directory, so a repo sitting elsewhere is invisible: it never triggers, and the failure is silent. After copying, confirm the skill appears in the host's skill list before concluding it works.

The user-global library at ~/.profer/global-skills/user/ is not enough on its own. Every entry there also needs a skill.manifest.json next to SKILL.md, and without it the directory is skipped silently — a plain cp into user/ looks installed and does nothing. Use the host's own "promote to global" action, or copy an entry that already has a manifest, rather than dropping a directory in by hand.

Claude Code — personal skills under ~/.claude/skills/, project skills under .claude/skills/:

cp -r resource-aware-compute ~/.claude/skills/

Any other host — copy the directory into the folder your host documents for skills.

Two warnings that come from testing this rather than assuming it:

  • If your host ships a marketplace CLI (for example npx skills add), check where it actually installs and whether the host reads that directory. Some hosts read one specific folder and ignore everything else, so a successful install can still be invisible.
  • A skill that is not discovered fails silently — no error, no log, it simply never triggers. Confirm it appears in your host's skill list before concluding it works. Running python scripts/probe_env.py on its own is a quick sanity check that the files arrived intact.

Usage

Two ways to use this, and only one of them is validated.

As a reference (recommended). Read SKILL.md and the references/ files when you hit a memory question. This is a lookup, and it is the use the evidence supports.

As a skill (not validated). Installing it and relying on the description to make an agent pick it up is the configuration that was measured twice and showed no gain. It may still trigger on simulation and large-array work, but do not expect it to change what the agent does — in round 3, the arm without it produced the same estimates and the same streaming code. Asking directly still works if you want the workflow applied:

Use the resource-aware-compute reference to check whether this fits in memory before running it.

The probe script is the one part that pays for itself, and it is standalone:

python scripts/probe_env.py          # human-readable report
python scripts/probe_env.py --json   # machine-readable
python scripts/probe_env.py --quiet  # one line, e.g. for pasting into a reply
python scripts/probe_env.py --self   # also report this process's peak RSS
OS        : Windows 11 (win32, Python 3.13.14)
CPU       : 12th Gen Intel(R) Core(TM) i7-12700H
            14 physical / 20 logical cores (source: platform-api, hybrid)
RAM       : 28.04 GiB available of 47.63 GiB total
Commit    : 32.69 GiB used of 74.63 GiB (43.8%)  [commit charge, not disk swap]

Budget for this machine (all figures GiB, 1024^3 bytes)
  run directly below      : 7.01 GiB peak
  chunk / memmap above    : 16.82 GiB peak
  do not exceed           : 25.24 GiB peak
  recommended threads     : 13 (OMP/MKL/OPENBLAS_NUM_THREADS)
  recommended processes   : 13

Compare your own estimate in GiB (bytes / 1024^3) against the bands above.

WARNING: Hybrid CPU (mixed performance/efficiency cores). A thread count near the
logical count will overload the efficiency cores.

All memory figures are GiB (1024³ bytes). Convert your own byte estimate the same way before comparing, or you will be off by ~7% in the direction that makes an over-budget run look safe.

Standard library only, no network access, no writes. Linux reads /proc/meminfo; macOS reads sysctl and vm_stat; Windows reads GlobalMemoryStatusEx and GetLogicalProcessorInformationEx through ctypes. psutil is used when present but is never required. --self reads the current process's peak RSS via getrusage (POSIX) or GetProcessMemoryInfo (Windows), so a run can be checked against the estimate that justified it.

Contents

Path Purpose
SKILL.md The workflow: probe, estimate, choose, guard, run, report. Correct as reference; the behavioural claim is unvalidated.
scripts/probe_env.py Machine budget probe. stdlib only, cross-platform. --self for peak RSS.
references/estimation-cookbook.md Peak-memory formulas per domain.
references/python-numeric.md numpy / pandas / scipy patterns that keep the peak down.
references/parallel-and-gpu.md Thread versus process, oversubscription, VRAM, the OOM ladder.
references/long-running.md Checkpointing, resumable loops, progress, background execution.
references/host-notes.md Host-specific shell timeouts and background support. Optional.
evals/evals.json The round-1/2 test cases. Has a 15/15 ceiling — kept for the record.
evals/hard/ A test set designed to have no ceiling. Contains its own README, fixture generator, validator, and RESULTS.md (round 3: delta +0.000 — it did not discriminate).

What it cannot do

Worth being explicit, because a skill is a prompt, not a sandbox:

  • It cannot enforce anything. No memory limit, no CPU cap, no sandbox. If the model ignores the instructions, nothing stops it. This raises the odds of good behaviour; it does not guarantee it.
  • It cannot watch a running process. It has no way to notice that a job is thrashing and intervene.
  • It is not a substitute for the right algorithm. It will help you notice that a dense N² matrix is the wrong structure — it will not invent the neighbour-list algorithm for you.
  • It is not a job scheduler. It knows how to shape work to survive a kill; it does not queue or retry anything.

If you need real isolation, that belongs in the host (containers, cgroups, Windows job objects) or in the code itself via a bounded runner. The address-space limit in SKILL.md is the closest thing to enforcement that runs entirely in user space, and it does not exist on Windows.

Evaluation

evals/evals.json contains tasks with objectively checkable expectations. Run them in pairs (with and without the skill) and the effect can be measured rather than asserted. None of the prompts uses the words "memory" or "performance".

Results as of 2026-09-18. Round 3 has now been run (see below): delta +0.000. Every number in this section except round 3 was measured against pre-1.0.2 revisions, so those rows neither describe the current revision nor establish that it is better. SKILL.md's Step 2 was rewritten, the sparse/N² errors were corrected, units were disambiguated, and the description was shortened after those measurements were taken. Treat rounds 1–2 as a historical record of what earlier versions scored.

What is still true without re-running: the benchmark has a ceiling (the unskilled arm scored 15/15, so the skill could only lose points), and n = 1 per configuration, so a single expectation flipping moves the total by 0.067. See What this benchmark can and cannot tell you.

Round 3 (2026-09-18): delta +0.000 — no measurable effect

A test set designed to have no ceiling (evals/hard/): unlabelled binaries, dimensions recoverable only via a sidecar, small stand-ins so that loading everything looks harmless. Both arms scored 14/15; the single missed item was the same for both and turned out to be a rubric defect, not an answer defect.

The premise — that a headerless binary would stop a capable executor from recovering the size unaided — is false for this executor. Both arms cat-ed the sidecar, derived the real dimensions, stated the correct footprints, and wrote streaming code. The skill added nothing measurable on these three tasks, and two of the five pre-registered predictions about which items would discriminate failed outright (0 of 5 held).

Full result, prediction scorecard and analysis: evals/hard/RESULTS.md. This is the second measurement in a row where the skill did not beat its baseline (round 2: −0.133). The consistent reading is that on well-specified tasks with a capable executor, the advice is redundant — see the RESULTS file for what that does and does not establish.

Round 2 (2026-09-16): delta −0.133 — the skill lost

Three tasks × two configurations, one run each, graded blind: the two outputs of each task were relabelled A/B with a per-task mapping the grader never saw. Host: Windows 11, 14 physical / 20 logical cores, 48 GB RAM, 27–30 GB available.

Eval with skill without skill
1 — trajectory radius of gyration 4/5 5/5
2 — FEM von Mises post-process 4/5 5/5
3 — Monte Carlo speed-up 5/5 5/5
total 13/15 (0.867) 15/15 (1.000)

Removing the stated data size from the prompts did not rescue the result. The unskilled arm read the format note by itself, found the real dimensions, said "your ~2 GB estimate is off, this is 38.4 GB", and wrote memory-mapped chunked code. So did the FEM run: it read the HDF5 attributes, discovered 300 steps × 2.4 M nodes, and streamed one step at a time.

Both failures in the skill arm are worth reading individually, because they are different kinds of event:

Eval 2 was a real defect, and the skill caused it. The arm with the skill estimated its own runtime peak — "about 1.3 GB" — and never the size of the stored dataset (17 GB, 5.8 GB per component). The arm without the skill stated both. The cause was SKILL.md itself: Step 2 said "Estimate the peak, not the final result", and the model dutifully reported the algorithmic peak. SKILL.md has been rewritten so Step 2 demands two labelled numbers — the data footprint and the implementation peak — and warns explicitly that "peak 1.3 GB" does not answer "how big is my data". This is the one concrete, actionable thing this benchmark produced, and it came from the skill losing.

Eval 1 is a grading artefact. The item asked for results written incrementally "rather than accumulating all frames in a Python list". The losing arm preallocated an np.empty(n_frames) array — 3.2 MB at 400 000 float64 — and wrote the CSV after the loop. The grader flagged it as borderline and failed it on the strict reading. Under a literal reading of the item text that arm passes, making the totals 14/15 vs 15/15, delta −0.067. The item has since been reworded to say what it meant (an O(n_frames) result array is fine; per-frame data is not). Both numbers are reported here; neither is presented as the answer.

Round 1 (2026-09-15): delta +0.067 — kept for the record

Measured against pre-1.0.2. See the version note above.

The same three evals against an earlier revision whose prompts stated the data scale ("400000 frames", "2.4 million nodes, 300 time steps"). Result: 15/15 vs 14/15, delta +0.067, with the only discriminating item being the thread/worker bound on eval 3. A capable executor converts a stated scale into a streaming implementation for free, so those two evals were measuring almost nothing. The prompts were rewritten; round 2 is the result.

That version of the table is kept rather than deleted, because a benchmark repaired after seeing the answer is worth exactly as much as you admit it is.

What this benchmark can and cannot tell you

Applies to rounds 1 and 2 (the evals/ set). Round 3 in evals/hard/ was written specifically to remove the two limits below; its own caveats are in evals/hard/README.md.

  • It has a ceiling. The unskilled baseline scores 15/15. A skill can only lose points here, never win them. If you want to measure the upside, the input has to be less self-describing — an unlabelled binary, no format note, no attrs recording the production size — or the executor has to be weaker. This is what evals/hard/ addresses.
  • One run per configuration. With n = 1 there is no variance estimate, and a single expectation flipping moves the total by 0.067. Treat a delta below roughly 0.2 as unresolvable by this benchmark.
  • It measures increment over a capable executor, not whether the advice is correct. The advice in SKILL.md is still the advice; what these runs cannot show is how much of it a strong model would have discovered unaided.
  • It is not evidence that the skill is useless. It is evidence that on well-specified tasks with readable inputs, a strong executor already does the estimation and the chunking. Where the skill plausibly still earns its place: thread and worker bounds, host timeouts and background execution, atomic checkpoints, and refusing to run something that cannot fit.

If you run these, the honest number is more useful than a flattering one. Open an issue with the results either way.

Revision notes

2026-09-18 — demoted to reference after round 3 came back null

Round 3 (evals/hard/) was run and returned delta +0.000: both arms scored 14/15, and the one missed item was missed by both. Five pre-registered predictions about which items would discriminate all failed. Documented in evals/hard/RESULTS.md.

Consequent changes, all of them removing claims rather than adding features:

  • The README headline ("stops your coding agent from freezing your laptop") is replaced with a status banner stating the two null results. An unvalidated efficacy claim does not belong in the first line.
  • SKILL.md carries the same banner, because an agent reading it should know what it is holding.
  • The description no longer says "trigger whenever compute cost is non-trivial". That instruction was written to maximise triggering, and the measurement did not support it — extra context for no demonstrated effect. It now asks to be used as a reference.
  • The Usage section splits "as a reference" (supported) from "as a skill" (not validated), and names probe_env.py as the one component with standalone value.
  • The H2 expectation that both arms failed was, on inspection, written wrong: it asked for chunking along the step axis, but the fixture is component-major with step fastest-varying, so step-axis chunking forces the strided random I/O both arms correctly avoided. Reworded to test the intent.

What was not done: adding features, widening triggers, or re-running the same test a third time. Two null results are enough to stop optimising this, and a third identical run would produce the same null (small stand-in fixtures never approach a real resource limit).

2026-09-18 — correctness pass

A review of the skill against its own claims found four factual or unit problems, all fixed:

  • A - B on scipy sparse does not densify. Both SKILL.md and references/python-numeric.md asserted it does. Verified against scipy: A - B, A + B, A * B (elementwise) and A.T @ B all stay sparse between two sparse operands. The operations that do densify are mixing in a dense operand (A * np.ones((n,n))), A.toarray(), and non-integer powers. Advice corrected in all three places, since the old wording would have deterred a correct sparse approach.
  • The N² row conflated the trap with the fix. "Pairwise distances or neighbor lists" was listed as N² in one row; a cutoff neighbour list is N × n_neighbours and is the recommended structure. Split into two rows — the trap and the sparse alternative.
  • Units were ambiguous. The probe computes GiB but labelled everything _gb and printed G; a bare decimal-GB estimate compared against it is off by ~7%, which is enough to cross a threshold on its own. SKILL.md, the cookbook and the README now state the GiB convention, the JSON carries explicit _gib aliases plus a units field, and the human output labels every band. The _gb keys are kept for backward compatibility.
  • host-notes.md presented host caps as facts. They are the most perishable content in the skill. The probing procedure now leads, the numbers are marked illustrative, and the section is labelled "verify before relying on them".

Two additions came out of the same pass:

  • probe_env.py --self reports the current process's peak RSS, with the platform split (KiB on Linux, bytes on macOS) handled and the Windows path implemented via GetProcessMemoryInfo. The argtypes/restype declarations are required — without them ctypes marshals the 64-bit handle as 32-bit and the call returns 0 silently, which is exactly the class of quiet failure this skill exists to prevent. The flag is opt-in, so the default JSON schema is unchanged.
  • Step 2 and Step 3 carry explicit caveats: the R bands are decision aids, not verdicts, and R between 0.6 and 1.2 should be read as "chunk it" because the estimate itself carries roughly ±30% uncertainty.

License

MIT. See LICENSE.

About

A reference for writing numerical, simulation and large-data code without freezing the machine. Includes a machine-probe script, per-domain memory formulas, and guard patterns for BLAS threading, chunking and long runs.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages