Repository navigation
Conversation
File.read_at returns a List of bytes, so a program that wants the bytes in an Array<U32> (a weight file, an image, any binary blob) builds a one-node-per-byte List and then packs it into the array. File.read_into preads up to max bytes at an offset into the array's own slots from a given slot on, little-endian, four bytes a slot, and returns the array with the count read. The read stops at the array's end and at the file's; a tail of under four bytes overwrites only its slot's low bytes; a read that fails at once (a write-only handle) returns the errno and leaves the array as it was. C writes into the array's block in place (through blk_loc, so a forked handle reaches the same cells, as Array.set does); JS packs the bytes into the JS array. tests/io/file_read_into.bend covers both lanes and the interpreter.
…reads JS in 16 MiB chunks macOS pread fails with EINVAL on any request past INT32_MAX, so a read into an array with 2 GiB of room failed even on a 44-byte file. The JS lane allocated the array's whole room up front and read once; it now reads in 16 MiB chunks, loops on short reads, and packs whole words. The C effect resets the request's error code, so a read that follows a failed one on the same run reports its own result (the test covers it).
|
Thanks for this, and for the thorough benchmarks. The speedup is real. But I don't think this belongs in Base, so I'm closing it. It would work as a library instead. It isn't shaped like any Base effect. Base's effects are thin syscall wrappers. Their arguments and results are only A library can do exactly the same thing. Your C and JS bodies move over almost unchanged. The trade-off is that the C side reaches undocumented internals ( If the underlying issue is that |
Why
File.read_atreturns aList<&2, U32>with one node per byte. A program that needs a file's bytes in anArray<U32>, such as a weight file, an image, or any other binary blob, has to build that list and then pack it into the array four bytes at a time. For large files, the list dominates both time and memory.Change
This change adds one effect to
bend2/base.bend:It reads up to
maxbytes atoffsetdirectly intoa's slots, starting at slotat. Bytes are stored little-endian, four bytes per slot. It returns the file, the array, and the byte count, or(errno, message).atis past the array, it reads nothing.File.read_at.Backends:
effs/file_read.c): The IO worker callspreadin a loop and writes directly into the array's block, which it reaches throughblk_loc. Each call asks for at most 1 GiB, because macOSpreadreturnsEINVALfor any request pastINT32_MAX, even on a small file. A forked handle reaches the same cells, as withArray.set.effs/file_read.js): It callspreadin 16 MiB chunks, loops on short reads, and packs whole words into the JS array. Only unaligned edges and the short tail go byte by byte, so the JS backend and the interpreter produce the same slots as C. A large array over a small file allocates only one chunk.Verification
tests/io/file_read_into.bendcovers the full read, the little-endian packing, the short tail that keeps its high bytes, a slot past the bytes that stays unchanged, an offset read into a later slot, clipping at the array's end, reads at the array's end and the file's end, the error on a write-only handle with the array left unchanged, and a good read issued right after that failure, which must report its own result. All 12#|lines pass in all three paths:bun bend2/main.ts t.bend)-o t)-o t.js, run under bun)tests/io/file_binary.bendstill passes in all three paths. The test caught a real bug in an earlier draft: when the JS path dropped the last partial slot, the test failed withFAIL short tail keeps high bytes.The last case caught a bug in a later revision: the C request is reused across a run's effects, and without resetting its error code, the good read reported the earlier
Bad file descriptor. With the reset removed, the C lane fails the test; with it, all lanes pass.Separate checks, all with the same result in C and JS:
Array.fork:read_intothrough one handle, then a read through the other, which saw the bytes.Array.new(U32, 29n, 0),max4294967295) over a 44-byte file: the first revision printedfail Invalid argumentin both C and JS; this one printsread 44.struct.unpack('<I', ...).Timing
Benchmark: read a 16 MiB random file into an
Array<U32>with 4,194,304 slots (Array.new(U32, 22n, 0)), then print slots 0 and 4,194,303. The baseline usedFile.read_at(f, 0, 16777216)and packed the list into the array withArray.set. The new path used oneFile.read_into(f, 0, 16777216, a, 0). The test machine was an Apple M1 Pro running macOS. Runs were whole-process wall times, with 5 alternating rounds per lane after one warm-up. Outputs matched in every run and matched Python'sstruct.unpack('<I', ...)values for the file. The 1-minute load average was about 7.6.read_at+ pack, medianread_into, first revisionread_into, this revisionread_atThe same change to a 1.3 GB quantized-model loader cut load time from about 10 s to about 1.6 s. Those runs were on a busy machine, so take that number as a rough indication only.
Notes / open questions
pread, asFile.read_atalready does. I did not run it on Linux, andgates/test.tsneeds the cluster, so I could not run it. As a stand-in, I ran all 187 tests intests/ioandtests/baselocally on this branch and on60fa05dc, each through the interpreter, C, and JS. The 186 shared tests printed byte-identical output on both trees in every lane, andfile_read_into.bendpasses in all three.offsetis aU32, the same asFile.read_at, so this change does not address files larger than 4 GiB.bun gates/repo.tspasses 54/54 withttok==0.3, the version pinned in.github/workflows/repo-gate.yml.