Skip to content

File.read_into reads a file's bytes straight into an Array<U32> - #1430

Closed
totaland wants to merge 2 commits into
bendlang:mainfrom
totaland:file-read-into
Closed

totaland wants to merge 2 commits into
bendlang:mainfrom
totaland:file-read-into

Conversation

@totaland

@totaland totaland commented Oct 8, 2026 •

Copy link
Copy Markdown

Why

File.read_at returns a List<&2, U32> with one node per byte. A program that needs a file's bytes in an Array<U32>, such as a weight file, an image, or any other binary blob, has to build that list and then pack it into the array four bytes at a time. For large files, the list dominates both time and memory.

Change

This change adds one effect to bend2/base.bend:

def File.read_into(file: File, offset: U32, max: U32, a: Array<U32>, at: U32) ->
  IO(File & (Array<U32> & Result<&1, &1, U32 & String, U32>))

It reads up to max bytes at offset directly into a's slots, starting at slot at. Bytes are stored little-endian, four bytes per slot. It returns the file, the array, and the byte count, or (errno, message).

  • The read stops at the array's end and at the file's end. If at is past the array, it reads nothing.
  • A tail of fewer than four bytes overwrites only the low bytes of its slot. The slot's high bytes stay as they were.
  • If a read fails immediately, as it does with a write-only handle, it returns the errno and leaves the array unchanged.
  • The file position does not move, as with File.read_at.

Backends:

  • C (effs/file_read.c): The IO worker calls pread in a loop and writes directly into the array's block, which it reaches through blk_loc. Each call asks for at most 1 GiB, because macOS pread returns EINVAL for any request past INT32_MAX, even on a small file. A forked handle reaches the same cells, as with Array.set.
  • JS (effs/file_read.js): It calls pread in 16 MiB chunks, loops on short reads, and packs whole words into the JS array. Only unaligned edges and the short tail go byte by byte, so the JS backend and the interpreter produce the same slots as C. A large array over a small file allocates only one chunk.

Verification

tests/io/file_read_into.bend covers the full read, the little-endian packing, the short tail that keeps its high bytes, a slot past the bytes that stays unchanged, an offset read into a later slot, clipping at the array's end, reads at the array's end and the file's end, the error on a write-only handle with the array left unchanged, and a good read issued right after that failure, which must report its own result. All 12 #| lines pass in all three paths:

  • interpreted (bun bend2/main.ts t.bend)
  • C binary (-o t)
  • JS (-o t.js, run under bun)

tests/io/file_binary.bend still passes in all three paths. The test caught a real bug in an earlier draft: when the JS path dropped the last partial slot, the test failed with FAIL short tail keeps high bytes.

The last case caught a bug in a later revision: the C request is reused across a run's effects, and without resetting its error code, the good read reported the earlier Bad file descriptor. With the reset removed, the C lane fails the test; with it, all lanes pass.

Separate checks, all with the same result in C and JS:

  • Array.fork: read_into through one handle, then a read through the other, which saw the bytes.
  • 2 GiB of room (Array.new(U32, 29n, 0), max 4294967295) over a 44-byte file: the first revision printed fail Invalid argument in both C and JS; this one prints read 44.
  • A 33,554,430-byte read from offset 2 across the JS chunk boundary: the byte count and the slots on both sides of the boundary match Python's struct.unpack('<I', ...).

Timing

Benchmark: read a 16 MiB random file into an Array<U32> with 4,194,304 slots (Array.new(U32, 22n, 0)), then print slots 0 and 4,194,303. The baseline used File.read_at(f, 0, 16777216) and packed the list into the array with Array.set. The new path used one File.read_into(f, 0, 16777216, a, 0). The test machine was an Apple M1 Pro running macOS. Runs were whole-process wall times, with 5 alternating rounds per lane after one warm-up. Outputs matched in every run and matched Python's struct.unpack('<I', ...) values for the file. The 1-minute load average was about 7.6.

Lane read_at + pack, median read_into, first revision read_into, this revision Speedup vs read_at
C 102.9 ms 9.1 ms 8.6 ms ~12x
JS (bun) 459.5 ms 53.2 ms 33.3 ms ~14x

The same change to a 1.3 GB quantized-model loader cut load time from about 10 s to about 1.6 s. Those runs were on a busy machine, so take that number as a rough indication only.

Notes / open questions

  • I tested this only on macOS arm64. It relies on pread, as File.read_at already does. I did not run it on Linux, and gates/test.ts needs the cluster, so I could not run it. As a stand-in, I ran all 187 tests in tests/io and tests/base locally on this branch and on 60fa05dc, each through the interpreter, C, and JS. The 186 shared tests printed byte-identical output on both trees in every lane, and file_read_into.bend passes in all three.
  • offset is a U32, the same as File.read_at, so this change does not address files larger than 4 GiB.
  • bun gates/repo.ts passes 54/54 with ttok==0.3, the version pinned in .github/workflows/repo-gate.yml.
  • I marked this PR as a draft. If you prefer different names or argument order, or if forked handles should be rejected instead of written in place, I can change it.

totaland and others added 2 commits October 9, 2026 09:41
File.read_at returns a List of bytes, so a program that wants the bytes
in an Array<U32> (a weight file, an image, any binary blob) builds a
one-node-per-byte List and then packs it into the array. File.read_into
preads up to max bytes at an offset into the array's own slots from a
given slot on, little-endian, four bytes a slot, and returns the array
with the count read. The read stops at the array's end and at the
file's; a tail of under four bytes overwrites only its slot's low bytes;
a read that fails at once (a write-only handle) returns the errno and
leaves the array as it was.

C writes into the array's block in place (through blk_loc, so a forked
handle reaches the same cells, as Array.set does); JS packs the bytes
into the JS array. tests/io/file_read_into.bend covers both lanes and
the interpreter.
…reads JS in 16 MiB chunks

macOS pread fails with EINVAL on any request past INT32_MAX, so a read
into an array with 2 GiB of room failed even on a 44-byte file. The JS
lane allocated the array's whole room up front and read once; it now
reads in 16 MiB chunks, loops on short reads, and packs whole words.
The C effect resets the request's error code, so a read that follows a
failed one on the same run reports its own result (the test covers it).
@nicolas-abril

Copy link
Copy Markdown
Collaborator

Thanks for this, and for the thorough benchmarks. The speedup is real. But I don't think this belongs in Base, so I'm closing it. It would work as a library instead.

It isn't shaped like any Base effect. Base's effects are thin syscall wrappers. Their arguments and results are only U32, String, List, handles and Results, and none of them takes or returns an Array. File.read_into would be the first that writes into a Bend value's memory: the IO helper thread preads straight into the heap through blk_loc. It would also make its rules part of Base's contract: the partial last slot keeps its high bytes, and a forked handle sees the writes.

A library can do exactly the same thing. guide/EFFECTS.md: a user effect's .c is spliced in after the runtime with every runtime symbol in scope, and its .js registers the same way Base's do. So the same def works from any module:

def File.read_into(file: File, offset: U32, max: U32, a: Array<U32>, at: U32) ->
  IO(File & (Array<U32> & Result<&1, &1, U32 & String, U32>)):
  import "./file_read_into.c"
  import "./file_read_into.js"

Your C and JS bodies move over almost unchanged. tests/io/cid_capture_lib.bend is an example of this pattern. The model loader can ship it next to its own code.

The trade-off is that the C side reaches undocumented internals (blk_loc, blk_cls, e.mem), and EFFECTS.md makes no ABI promise, so the library has to track runtime renames. That's the same deal every custom effect has.

If the underlying issue is that File.read_at's one-List-node-per-byte result is too slow for binary data, a Base-shaped fix would be a general byte primitive rather than an effect that fills one array type. If you want to pursue that, please open an issue first.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants