The RGBA frame a host hands an ffrwd wasm module, turned into the
tensor a vision model reads: a crop, Pillow's resize, a normalization,
and the little-endian fp32 bytes wasi:nn takes. Plain Rust with one
dependency, so a module compiles it in and its tests run on the host.
use ffrwd_frame::{tensor, Filter, Rect, Rgba, IMAGENET};
let frame = Rgba::new(&bytes, width, height)?;
let input = tensor(&frame, Rect::whole(width, height), 576, 576, Filter::Bilinear, IMAGENET);A module lists the pixel formats it accepts, first choice first, and
the host converts the stream to that before the first frame arrives.
List rgba and ffmpeg's swscale does the colour conversion upstream,
where it reads the stream's own range and matrix rather than guessing
at them. That is still the advice: a module that converts colour
itself has to know whether black is 0 or 16, and the three private
converters this crate's yuv replaces all assumed full range on a
video-range stream, which flattens the picture by about a seventh.
What is left is the part a model is sensitive to. Rect::padded
turns a detector's box into the crop a second model reads, widened by
a fraction of its own size on every side and clamped to the frame.
Filter::Bilinear and Filter::Bicubic are Pillow's BILINEAR and
BICUBIC, antialiased when downscaling, and match PIL.Image.resize to
within one count of 255 on every fixture in tests/, because the
models these modules run were trained on Pillow crops and a different
kernel moves the answer. Norm is the per-channel mean and standard
deviation applied after scaling to 0..1; IMAGENET is the usual one.
planes returns the crop as [3, height, width] floats, tensor the
same as bytes with a leading batch dimension of one, and tensors
several crops of one frame as one [n, 3, height, width] tensor, for a
model that takes a batch.
yuv converts yuv420p to RGBA and RGBA back to yuv420p, for a module
that takes yuv420p on the wire, where a 4K picture is three eighths of
the bytes, or writes it back out. It reads the range and the matrix
from the color-info the host hands over: tv or pc, and bt709,
smpte170m (also bt470bg and bt601), bt2020nc, fcc or
smpte240m. An unknown range is tv and an unknown matrix is BT.601
at every size, as swscale reads them, so a module converting for itself
sees the picture the host would have handed it. Any other matrix is an
error. Where the bytes allow it, asking the host for rgba is still
the better choice.
use ffrwd_frame::yuv::{self, ColorInfo, Colour, Yuv420p};
let colour = Colour::of(Some(info))?;
yuv::to_rgba(&Yuv420p::new(&bytes, width, height)?, colour, &mut rgba)?;Each chroma sample covers its two-by-two block of pixels on the way in,
which is where swscale's default puts it, and a block's chroma is the
mean of its four pixels on the way out. swscale's default flags round
coarsely, so the rgba the host hands a module is up to three counts
from this, and darker on average by half a count or more in tv range.
fast_image_resize is a port of Pillow's resampling and supplies the
kernels, but it runs its vertical pass before its horizontal one for
speed, where Pillow goes horizontal first. Each pass rounds back into
eight bits, and the bicubic kernel undershoots past black at a hard
edge, so the order shows in the answer: nine counts off at an edge
when the two axes go through one call. This crate asks for the two
passes separately in Pillow's order, and skips a pass whose axis does
not change size, as Pillow does. The doc comment on resized_rgb
says so; leave it that way.
The crate builds for wasm32-wasip2 with no target feature. The
resize library takes the simd128 path on wasm whenever it is present,
and the yuv loops ask for simd128 the same way, so the module runs
SIMD instructions and the runtime needs the SIMD proposal, which
wasmtime has on by default.
cargo test
cargo build --target wasm32-wasip2 --release
tests/pillow.rs compares both filters against Pillow 12.3 on an
upscale, a downscale and the identity over two synthetic frames, one
with a hard edge and one of noise; scripts/pillow_fixtures.py
regenerates tests/pillow.bin byte for byte. tests/faceage_bicubic.rs
runs the fixtures the faceage package's own bicubic was checked
against. tests/yuv_ffmpeg.rs holds the yuv conversion to within one
count of ffmpeg for every range and matrix, and skips where there is no
ffmpeg; cargo bench --bench yuv times one 1080p frame each way.
Apache-2.0.