Conversation
Add a loadable (LLEXT) audio module that wraps FFmpeg. Two build modes: - decoder (default): a compressed elementary stream (FLAC/AAC/Opus, selected via Kconfig) is parsed and decoded to PCM (process_raw_data); - filter (CONFIG_FFMPEG_DEC_FILTER_MODE): a PCM source/sink effect that runs an FFmpeg audio filter graph (afftdn noise reduction). FFmpeg is a west-pinned source (see west.yml) cross-built by ffmpeg.cmake as a CMake ExternalProject, enabling only the Kconfig-selected decoders and filters. The module supplies the libc surface FFmpeg needs that the SOF core does not export to LLEXT (ffmpeg_dec-shims.c), a SOF-heap-backed malloc (ffmpeg_dec-alloc.c), fast single-precision float math (fastmathf.c), and routes av_log() into the Zephyr log. The avfilter graph backend and the PCM effect ops are in ffmpeg_dec-filter.c. Also adds the rimage module manifest (ffmpeg_dec.toml + per-platform includes), the topology widget (ffmpeg_dec.conf), host test tooling (ffmpeg_dec_prepare.sh) and docs (README/TESTING). Verified building and load-ready (all symbols resolved, signed) for ace30 (ptl) with the Zephyr SDK: FLAC and FLAC+AAC decoders, and the afftdn filter effect. Runtime decode/denoise verification is pending hardware. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add the ffmpeg_dec component to the host->component->host benchmark harness so a topology instantiating the module can be built and driven (testbench or target). Adds the per-format bench configs (generated with bench_comp_generate.sh), registers the widget include and the ffmpeg_dec32 BENCH_CONFIG in cavs-benchmark-hda.conf, and adds ffmpeg_dec to the s32 component list in tplg-targets-bench.cmake (the filter effect path processes S32). Also drop the codec-setup bytes control from the ffmpeg_dec widget class so it is a plain 1-in/1-out effect; the decoder's STREAMINFO control is added at the instance level in a decoder topology instead. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add CONFIG_FFMPEG_DEC_MP3: enables the libavcodec mp3 decoder and mpegaudio parser in the cross-build, the FFMPEG_DEC_CODEC_MP3 codec id mapping, and the float math layer (the mp3 decoder uses it). Verified building for ace30 (ptl) with decoders [flac,mp3]. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
FFmpeg has no native MP3 encoder, so add MP3 encode through libshine, a small fixed-point encoder (good fit for a DSP - no float dependency). libshine is a west-pinned project; ffmpeg.cmake cross-builds it (its lib needs no config.h and its autotools CLI/shared link cannot work bare-metal, so the objects are compiled and archived directly), generates a shine.pc, and enables --enable-libshine --enable-encoder=libshine in FFmpeg. FFmpeg's require_pkg_config link test for -lshine pulls newlib malloc and hence Zephyr-runtime symbols (z_errno_wrap, ...) that only exist at module load; a configure-test-only stub is passed via --extra-ldflags so the test links (--extra-ldflags does not affect the static-archive build). Gated by CONFIG_FFMPEG_ENC_MP3. Verified building for ace30 (ptl): libshine.a cross-builds and libavcodec.a contains the libshine encoder. A module encode path (PCM->MP3) is a separate build mode. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add CONFIG_FFMPEG_DEC_ENCODE_MODE: build the module as an encoder using the modern .process_raw_data path in reverse of the decoder - avcodec_send_frame/avcodec_receive_packet. ffmpeg_dec-encode.c opens the libshine MP3 encoder, converts SOF interleaved S32 to the encoder sample format per frame, and emits the compressed elementary stream. The interface ladder in ffmpeg_dec.c selects encoder / filter / decoder ops by Kconfig. libshine is linked into the module (after libavcodec, which references it) from its own cross-build install dir. ffmpeg.cmake now allows an encoder-only build (no decoder) and makes the decoder configure flags conditional. The encoder path pulls a few more libc symbols, added to the local surface: calloc/posix_memalign (alloc), modf/localtime_r/iconv*/__xpg_strerror_r (shims). fastmathf is now always built for the real backend (its sofm_log2f backs the shims' log10()). Verified building and load-ready (no unresolved externals, signed) for ace30 (ptl) as an encoder-only module. Runtime encode + real-time framing need hardware. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add HIFI.md: analysis of the hot audio-processing paths in the module (decode/ filter/encode) and the linked FFmpeg DSP, and where Xtensa HiFi intrinsics would help. FFmpeg is built --disable-asm (scalar C on Xtensa) and has no _xtensa DSP init, so the generic C kernels run. Documents: the module's own PCM conversion loops (easy HiFi wins we own), FFmpeg's DSP dispatch contexts (float_dsp, flacdsp, mpegaudiodsp, tx) as fork-patch targets mirroring the other arch inits, fastmathf vectorisation, and a cost/benefit priority order. References SOF's existing HiFi kernels (src/math/*_hifi*) as the template. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add the webrtc_vad SOF module wrapping libfvad — a standalone C port of the WebRTC Gaussian mixture model VAD (identical to the algorithm shipping in Chrome/WebRTC since 2011). Module characteristics: - Single-source, single-sink PCM pass-through - Accumulates 10/20/30 ms frames (configurable via Kconfig) - Broadcasts NOTIFIER_ID_VAD event on every frame classification - Supports S16_LE and S32_LE at 8/16/32/48 kHz, mono or stereo - Fixed-point; no FPU required - Two backends: real (libfvad) and pass-through stub for CI/LLEXT Build: - libfvad is cross-compiled by src/audio/webrtc_vad/webrtc_vad.cmake - Source is fetched via west (modules/audio/libfvad, pinned SHA 532ab666c20d) - LLEXT packaging supported via llext/ subdirectory Add NOTIFIER_ID_VAD to the notifier_id enum so downstream components (e.g. webrtc_ns2 / RNNoise) can share the same VAD event channel. UUID: c790b11d-5d14-e54e-be36ba4ad732cc14 Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Add the webrtc_ns SOF module wrapping the classic WebRTC Noise Suppression algorithm (spectral Wiener filter, fixed-point) from the webrtc-audio-processing 0.3.1 library. Module characteristics: - Single-source, single-sink PCM effect - Fixed-point implementation; no FPU required - Operates at 8 or 16 kHz on 10 ms frames (160/80 samples) - Supports S16_LE and S32_LE, mono or stereo (per-channel instances) - Four suppression levels: Mild/Medium/Aggressive/VeryAggressive (configurable at build time or via IPC4 set_configuration) - Two backends: real (WebRTC NS) and pass-through stub for CI/LLEXT - Designed as a build-time alternative to webrtc_ns2 (RNNoise): lower CPU, no 48 kHz constraint, fully fixed-point Build: - WebRTC NS subset is cross-compiled by src/audio/webrtc_ns/webrtc_ns.cmake (6 C files from modules/audio/webrtc-apm) - Source fetched via west (modules/audio/webrtc-apm, tag v0.3.1) - LLEXT packaging supported UUID: 0fc8faef-945f-004b-8d5f315047f1136a Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Add the webrtc_aec SOF module wrapping WebRTC AECm (Acoustic Echo Canceller Mobile) — the fixed-point Q15 echo canceller designed for mobile and embedded devices. Module characteristics: - Dual-source (2 input pins): pin 0 = microphone capture, pin 1 = playback echo reference - Single-sink: echo-cancelled microphone output - Fixed-point Q15; no FPU required - Operates at 8 or 16 kHz on 10 ms frames - Supports S16_LE and S32_LE, up to WEBRTC_AEC_CHANNELS_MAX channels (one AecmCore handle per channel) - Pipeline-ID heuristic (matching google_rtc_audio_processing) used to distinguish mic vs. reference at prepare() time - Two backends: real (AECm) and pass-through stub for CI/LLEXT Pin binding follows the same pattern as google-rtc-aec: both DAI capture copiers are connected as sources; the firmware resolves mic-vs-ref from pipeline membership. Reference topology: tools/topology/topology2/development/ cavs-nocodec-webrtc-aec.conf (dual SSP: SSP0=mic, SSP2=ref) UUID: e5da7b5b-133a-ba46-b517651d6300bb83 Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Add the webrtc_ns2 SOF module wrapping xiph/rnnoise — a recurrent neural network (GRU-based) noise suppressor originally developed by Jean-Marc Valin at Mozilla. Module characteristics: - Single-source, single-sink PCM effect - Floating-point; internal float math, hardware FPU beneficial - Hard requirement: pipeline sample rate MUST be 48000 Hz - Processes 480-sample (10 ms) frames; partial periods accumulated internally - Supports S16_LE and S32_LE, mono or stereo (one DenoiseState per channel; up to WEBRTC_NS2_CHANNELS_MAX channels) - Float scale bridging: pipeline samples normalised to ±1.0 are scaled to RNNoise's native ±32768 range and back - VAD dual-use: rnnoise_process_frame() returns per-frame speech probability [0.0, 1.0]; this fires NOTIFIER_ID_VAD events at a configurable threshold (WEBRTC_NS2_VAD_THRESHOLD_PCT) - Two backends: real (RNNoise) and pass-through stub for CI/LLEXT Design notes: - rnnoise_init() used at prepare/reset instead of rnnoise_create() to avoid per-init heap allocation (embedded-friendly path) - Model weights (rnn_data.c) are const float[] placed in .rodata (Flash/ROM); ~85-340 KB depending on model variant - Peak stack ~12-16 KB per rnnoise_process_frame() call - No expf/tanhf in hot path: replaced by 201-entry tansig LUT Build: - RNNoise cross-compiled by src/audio/webrtc_ns2/webrtc_ns2.cmake (6 C files: denoise.c rnn.c rnn_data.c pitch.c celt_lpc.c kiss_fft.c) - Source fetched via west (modules/audio/rnnoise, SHA 70f1d256) - LLEXT packaging supported UUID: eacfacdc-2a87-c942-97e894c917a740db Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Register all four WebRTC audio processing modules in the SOF build
system:
Kconfig (src/audio/Kconfig):
- rsource the four module Kconfig files (alphabetical order)
CMakeLists (src/audio/CMakeLists.txt):
- add_subdirectory() guards for COMP_WEBRTC_{VAD,NS,AEC,NS2}
UUID registry (uuid-registry.txt):
- c790b11d-5d14-e54e-be36ba4ad732cc14 webrtc_vad
- 0fc8faef-945f-004b-8d5f315047f1136a webrtc_ns
- e5da7b5b-133a-ba46-b517651d6300bb83 webrtc_aec
- eacfacdc-2a87-c942-97e894c917a740db webrtc_ns2
west.yml:
- Add remotes: libfvad (dpirch), webrtc-apm (freedesktop.org),
rnnoise (xiph)
- Add project entries with pinned revisions:
libfvad @ 532ab666c20d → modules/audio/libfvad
webrtc-apm v0.3.1 → modules/audio/webrtc-apm
rnnoise @ 70f1d256 → modules/audio/rnnoise
Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Add Topology 2.0 component widgets, capture pipeline templates, and
standalone development/test topologies for all four WebRTC modules.
Component widgets (include/components/):
webrtc-vad.conf — pass-through VAD; S16/S32; any rate; 1→1
webrtc-ns.conf — fixed-point spectral NS; 8/16 kHz; S16/S32; 1→1
webrtc-aec.conf — AECm echo canceller; 10ms IBS/OBS; 2 input pins
(pin 0 = mic, pin 1 = echo ref); S16/S32; 2→1
webrtc-ns2.conf — RNNoise GRU NS; 48 kHz only; S16/S32; 1→1
Capture pipeline templates (include/pipelines/cavs/):
webrtc-ns-capture.conf — DAI-copier → VAD → NS → module-copier
webrtc-aec-capture.conf — copier → AEC(2-pin) → copier
webrtc-ns2-capture.conf — RNNoise → module-copier (48 kHz locked)
Development/test topologies (development/):
cavs-nocodec-webrtc-ns.conf — single SSP loopback; VAD+NS
cavs-nocodec-webrtc-aec.conf — dual SSP (SSP0=mic, SSP2=ref); AEC
mirrors cavs-nocodec-rtcaec.conf
cavs-nocodec-webrtc-ns2.conf — single 48 kHz SSP; RNNoise
All three development topologies registered in tplg-targets.cmake
targeting TGL nocodec with NHLT preprocessing.
Compile example:
alsatplg -c development/cavs-nocodec-webrtc-aec.conf -D PLATFORM=tgl -o sof-tgl-nocodec-webrtc-aec.tplg
Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Add the root Kconfig that exposes the ffmpeg_dec module configuration tree (codec selection, filter mode, cold-split experimental option) and the standalone ffmpeg.cmake cross-build helper for out-of-tree builds. These files were omitted from the earlier ffmpeg_dec commits and are grouped here to keep all ffmpeg_dec build system pieces together. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
…and webrtc modules Create and update README.md documentation for the FFmpeg decoder/encoder/filter module and all four WebRTC audio modules: - ffmpeg_dec/README.md: Added Mermaid architecture diagram, detailed decoder/encoder/filter configurations, build choices, and topology guides. - webrtc_vad/README.md: Included Mermaid data flow, GMM classification details, ring-buffer accumulation behavior, and notification topology mappings. - webrtc_ns/README.md: Documented fixed-point spectral Wiener filter complexity, channel separation, and Kconfig rules. - webrtc_aec/README.md: Described dual-input AECm echo cancellation routing, pipeline-ID heuristics, and layout designs. - webrtc_ns2/README.md: Documented deep-learning/RNNoise GRU topology, 48 kHz lock rules, VAD probability thresholds, and DSP scaling metrics. Each document contains a tailored Mermaid diagram mapping data/control flows. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Document the per-frame AAC-LC decode hot spots for Xtensa HiFi work: IMDCT (libavutil/tx), the AVFloatDSPContext vector kernels (vector_fmul_window and friends), x^(4/3) inverse quant, and TNS, plus SBR/PS for HE-AAC. Includes a toolchain-requirement note: the HiFi intrinsic headers (xt_hifi*.h) and an ace30 core-isa.h with XCHAL_HAVE_HIFI4=1 are only available via the LLVM/xt-clang Xtensa toolchain, not the Zephyr-SDK GCC currently used by ffmpeg.cmake, so the ff_*_init_xtensa kernels need an intrinsic path plus a scalar C fallback. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…loat_dsp
Activate the Xtensa HiFi AVFloatDSPContext kernels (lgirdwood/FFmpeg sof-hifi,
ff_float_dsp_init_xtensa) in the ffmpeg_dec module build:
- ffmpeg.cmake: stop passing --disable-asm to FFmpeg configure. FFmpeg treats
every per-arch optimisation dir -- including our C-intrinsic libavutil/xtensa/
-- as "asm"; --disable-asm forces arch=c and drops them (ARCH_XTENSA=0). There
is no Xtensa assembly, so leaving asm enabled is safe and is what makes
ARCH_XTENSA=1 and builds ff_float_dsp_init_xtensa.
- west.yml: bump ffmpeg to sof-hifi @ a600a78acd (adds libavutil/xtensa HiFi
float_dsp on top of the av_tx table cap).
Verified: pristine ptl (intel_adsp/ace30/ptl) AAC build (FLAC+AAC) reconfigures
FFmpeg with ARCH_XTENSA=1, compiles libavutil/xtensa/float_dsp_init.o, links it
into ffmpeg_dec.llext with no errors or dangerous relocations. Under the
Zephyr-SDK GCC the kernels run as the scalar C fallback (GCC lacks xt_hifi4.h);
the xtfloatx2 SIMD path activates with an xt-clang/XCC build for the ace30 core.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pick up the __XCC__ gate on the Xtensa float_dsp SIMD path so the ffmpeg_dec module builds cleanly with the LLVM Xtensa clang (was: intrinsic branch entered via the xt_hifi4.h wrapper and failed on undeclared xtfloatx2). GCC behaviour is unchanged (scalar fallback either way). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a clang branch to the FFmpeg cross-build so the module can be built with the
LLVM Xtensa clang (needed for any HiFi intrinsics -- the Zephyr-SDK GCC ships no
xt_hifi/ae_* intrinsics header at all). clang is a generic driver, so unlike the
target-specific GCC it needs the target/core/sysroot spelled out, uses the
Zephyr-SDK GNU binutils (ar/as/nm/objcopy), and cannot link the bare-metal
configure *test* executables (no default crt), so those link via the GNU gcc
(--ld). All values are derived from CMAKE_AR (GNU binutils even under clang):
cross-prefix, sysroot, and -mcpu=<core> from the triple. Also disables auto/SLP
vectorisation: the LLVM Xtensa backend cannot lower the v2i32 bswap FFmpeg
byteswap code gets SLP-vectorised into ("Cannot select: v2i32 = bswap"), and
it buys nothing (no packed float SIMD codegen on Xtensa anyway).
The GCC path is unchanged (guarded on CMAKE_C_COMPILER_ID) and re-verified green:
pristine ptl AAC build -> ARCH_XTENSA=1, float_dsp_init.o compiled, ffmpeg_dec.llext
linked. The clang recipe is validated at the FFmpeg-configure/build level: a
standalone decoder-only cross-build (avcodec+avutil+swresample, FLAC) with
~/work/llvm-project clang links clean and produces elf32-xtensa-le objects. A full
firmware clang build additionally requires the Zephyr LLVM toolchain variant.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Setting the bytes control max inside the host-gateway-tdfb-drc-capture pipeline class had no effect: alsatplg leaves the widget control at the default 1024, so pushing a larger TDFB blob with sof-ctl failed. Move the max to the widget instantiation in sdw-dmic-generic.conf where it is honored, and keep the original 16384 headroom so the setting also covers larger beamformer blobs that may be added later. Signed-off-by: Seppo Ingalsuo <seppo.ingalsuo@linux.intel.com>
Hook a tdfb blob validator into the model handler so a corrupted or mismatching run-time configuration update is rejected before it can replace the working blob. Capture then continues with the previously set filters instead of being interrupted by bad re-configuration. The TDFB blob is variable size and the per-filter walk in tdfb_init_coef() was not bounded against the IPC payload, so a bad length field could push tdfb_filter_seek() past the buffer. The new validator walks every FIR section and the trailing arrays with byte-bounded steps, requires the layout to exactly match config->size, and rejects blobs whose channels or filter_index values do not fit the running stream. The stream channel counts are cached in tdfb_comp_data so the validator can consult them, and the validator is re-run at prepare time once those counts are known. A malformed initial blob is discarded so the runtime path cannot dereference it. With ingress fully validated, the redundant sanity checks inside tdfb_init_coef() are dropped. Signed-off-by: Seppo Ingalsuo <seppo.ingalsuo@linux.intel.com>
This patch hooks a drc blob validator into the model handler so a corrupted run-time configuration update is rejected before it can replace the working blob. Playback or capture then continues with the previously set parameters instead of being interrupted by a bad IPC. The DRC configuration is a fixed-size struct sof_drc_config, so the validator requires the IPC payload size to match exactly and the self-declared config->size to agree with it. It is installed in drc_init() when the model handler is created, so every blob swap is covered - including any received while the component is in READY - and the redundant size check that drc_prepare() used to run on the initial blob is dropped. Signed-off-by: Seppo Ingalsuo <seppo.ingalsuo@linux.intel.com>
Add native_sim board configuration and support in the build script. This allows building and running tests on the host using Zephyr's native_sim target. native_sim leverages the POSIX architecture, but the libfuzzer support specifically requires CONFIG_ARCH_POSIX_LIBFUZZER to be set. Therefore, this wraps fuzzer-specific code in ipc.c and the build of fuzz.c behind this config to allow clean compilation on the standard native_sim board. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Add native_sim board target to the sof-qemu-run scripts, and add an option to additionally run it under valgrind. The default build directory is set to ../build-native_sim Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
When building the firmware for native_sim, debugging allocations with host machine tools like Valgrind is constrained due to Zephyr's internal minimal libc tracking the heap manually via static pools. By bypassing Zephyr's memory interception on native_sim using nsi_host_malloc, dynamically tracked memory can surface appropriately to Valgrind memory checkers without causing a libc heap pool panic. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Keep spinning in case user needs to inspect status via monitor. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
When building the native_sim fuzzer, the host allocator does not possess the strict bounds of the internal Zephyr memory pools. If the fuzzer generates a malformed payload requesting an excessively large size (e.g. 4GB), it passes directly to the host ASAN allocator which aborts due to OOM or protection limits. Adding a 16MB cap allows these to fail gracefully. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Only run sof_run_boot_tests() at boot for QEMU or native_sim (STANDALONE) targets. Other platforms trigger via IPC.
The feed-forward and feedback paths copy frames * channels samples into fixed intermediate buffers but only checked the frame count. Bound the total sample count against the buffer capacity so an unexpected channel count cannot overflow the buffers. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
The remap routines used host-provided channel map entries to index the source frame, guarding only against the -1 unmapped sentinel. Skip any entry outside [0, source channels) so a crafted map cannot read past the source frame. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
The fragmented calibration get path advanced a read offset by a host-controlled index without checking it against the calibration data size, allowing reads past the buffer. Reject offsets at or beyond the data size and fragments that would extend past it. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
… topology
Bring up the libshine fixed-point MP3 encoder end-to-end in the ffmpeg_dec
module in encode mode (CONFIG_FFMPEG_DEC_ENCODE_MODE / CONFIG_FFMPEG_ENC_MP3),
built =y into zephyr.elf, plus the compress-capture topology that feeds it PCM
from the onboard HDA analog capture BE and delivers MP3 to the host.
Firmware:
- ffmpeg_dec-encode.c: encode path driven through the DP source/sink
interface. The module exposes .process + .is_ready_to_process instead of
.process_raw_data; is_ready_to_process gates DP re-invocation on
source_get_data_available(sources[0]) > 0. The DP task opens the encoder
lazily (deferred avcodec_open2) on first process(), accumulates a full
1152-sample stereo frame (need = 9216 bytes) from the source ring, encodes,
and writes the variable-size MP3 frame to the sink.
- ffmpeg_dec.c/.h: wire the new source/sink process + is_ready_to_process ops.
- base_fw.c: advertise SND_AUDIOCODEC_MP3 for capture, guarded so it is not
duplicated when the Cadence MP3 encoder is present.
- ffmpeg.cmake / CMakeLists.txt: build libshine in encode mode.
Topology (tools/topology/topology2):
- include/components/ffmpeg_enc.conf (new): encoder widget class, same module
UUID as ffmpeg_dec, DAPM type "encoder" (the compress kernel path accepts an
encoder-type module for capture). MP3 is self-describing, no setup control.
- include/pipelines/cavs/compr-capture-ffmpeg.conf (new): compress-capture
pipeline; in-pipeline LL head copier -> ffmpeg_enc (DP) -> aif_out host-copier.
- platform/intel/compr-capture-ffmpeg.conf (new): compress-capture PCM (id 51,
direction capture, MP3) and routes tapping the shared HDA analog capture BE.
- cavs-mixin-mixout-hda.conf: turn the HDA analog capture module-copier into a
proper 2-pin fan-out (num_output_pins/num_output_audio_formats 2 +
Object.Base.output_pin_binding by sink name; requires the
common/output_pin_binding.conf include) so the compress FE can bind a second
consumer on its own pin instead of colliding on pin 0.
- compr-default.conf / sof-hda-generic.conf / tplg-targets.cmake: COMPR_ENC_*
defines, include wiring, and the sof-hda-generic-ffmpeg-mp3-enc-compr target.
Status: the MP3 encoder core is proven at runtime on aphid (PTL) -- deferred
avcodec_open2 succeeds ("libshine MP3 open OK, rate 48000 ch 2 frame_size 1152
need 9216"), the self-test emits a valid MP3 frame, firmware boots with no
fault. The live-capture data path is not yet flowing: the compress-capture FE
receives 0 bytes from the shared analog capture BE, and starting the compress FE
disrupts a concurrently running normal capture. The two topology fixes above
(fan-out pin-binding + in-pipeline head copier) are real and correct but did not
change the symptom; the remaining work is at the firmware/kernel DPCM level
(compress-capture trigger/hw_params from a shared analog BE) and is left as a
documented follow-up.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The libshine MP3 encoder now produces valid MP3 on PTL/aphid silicon (avcodec_open2 OK, MPEG-1 Layer III 128 kbps/48 kHz/stereo, no faults) once the shine clang/Xtensa backend miscompile is worked around. Two sizing fixes are required to reach the first encoded frame without overrunning the capture DAI. process() runs on the DP thread and blocks for the full duration of a single libshine frame encode while the LL side keeps producing real-time capture PCM. The default (zero) in_buff_size gives too shallow an input ring, which overflows mid-encode, backs up the shared HDA-capture fan-out copier, and overruns the capture DAI gateway before the first MP3 byte exists. Set in_buff_size = FFMPEG_ENC_PCM_STAGE (16 KiB) so the ring (sized ~3x this by bind) absorbs one whole worst-case encode. Bump the compress-capture module heap_bytes_requirement to 131072 to cover the deepened input ring plus pcm_in and mp3_out. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The compress-capture (encode) FE fed its DP ffmpeg_enc module from a module-copier second output pin (module-copier.4.2 pin 1). A copier output pin does not deliver data across a pipeline boundary, so the encoder input ring stayed empty and the DP module starved - crecord returned 0 bytes. Rebuild the HDA analog capture BE around the proven mixer fan-out when COMPR_FFMPEG_ENC is set: the BE (dai-copier -> gain -> mixin.4.1) fans a single mixin out to two mixouts - mixout.3.1 for the normal analog capture FE (pcm0c) and mixout.91.2 for the encoder FE (pcm51) - exactly as the playback mix path distributes a stream. mixin->mixout delivers across the pipeline boundary; a copier pin does not. Because a DP module breaks the LL copy walk, the encoder tap mixout cannot share the DP encoder pipeline (it would sit on the far side of the break from that pipeline sink scheduler and never run). Isolate it in its own timer-scheduled mixout-capture pipeline (COMPR_ENC_MIXOUT_PIPELINE_ID=93): standalone it self-schedules on the timer, drains the mixin, and its output crosses the boundary into the DP encoder input ring - mirroring how the decode pipeline LL host-copier feeds ffmpeg_dec. Runtime graph: mixin.4.1 -> mixout.93.1 -> ffmpeg_enc.91 -> host-copier. Verified on PTL (aphid): PCM->MP3 end to end, host-readable MPEG-1 layer III 128 kbps 48 kHz stereo (0xFFFB frame sync), avcodec_open2 OK, no faults. The normal analog capture FE (arecord pcm0c) runs unaffected on the shared BE. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Implement specialized fast-path PCM format conversion and interleaving routines in ffmpeg_dec-convert.c to replace the per-sample nested switch logic in ffmpeg_dec_emit_frame(). Optimizations include: - Direct memcpy_s pass-through when source format and layout match sink (interleaved S16 -> S16_LE, interleaved S32 -> S32_LE). - Fast-path planar S16P -> interleaved S16 and S32 (stereo & mono). - Fast-path planar S32P -> interleaved S32 (stereo with HiFi3 vector selection & stores where available). - Fast-path planar float FLTP -> S32 and S16 with hardware-accelerated HiFi4/5 VFPU (XT_TRUNC_SX2) and branchless soft-float-free scalar fallback. - Eliminate per-sample switch(frame->format) and switch(sink_fmt) branches. - Mark fallback libc stubs in ffmpeg_dec-builtin-libc.c as weak so they yield to Zephyr minimal libc implementations on targets like TGL. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
…ing (Phase 4) Optimize memory dataflow and eliminate redundant buffer copying: - ffmpeg_dec-ffmpeg: add zero-copy packet fast path when input stream bitstream already has guaranteed 64-byte padding (in parser internal buffer or at the zero-padded end of in_buf), bypassing the 64 KiB pktbuf memcpy. - ffmpeg_dec: add direct-to-sink PCM rendering fast-path. When the sink circular buffer has sufficient contiguous space for a decoded frame, decode directly into the sink buffer and commit directly, eliminating the intermediate cd->pcm_buf copy.
- Define SND_AUDIOCODEC_OPUS_RAW and advertise Opus decode capability in base_fw. - Filter out non-Opus extradata from topology to allow default Opus header fallback. - Right-size FFMPEG_DEC_HEAP_BYTES to 576 KiB for non-AAC decoders, reclaiming 304 KiB for sof_heap and preventing DP stack allocation exhaustion.
…t (Phase 6) - Accelerate S32 -> S16/S16P PCM format conversion in ffmpeg_enc_mod_process() with 4-way unrolling and branch hoisting. - Initialize scheduler_dp_init() in CAVS platform_init() when CONFIG_ZEPHYR_DP_SCHEDULER is enabled, allowing primary core DP modules.
…strtod support (Phase 7)
…tensa SIMD - Update ffmpeg.cmake to enable aac_fixed decoder with aac/aac_latm parsers. - Update ffmpeg_dec-ffmpeg.c to look up aac_fixed by name for AAC codec ID. - Update Kconfig description for fixed-point AAC-LC decoder. - Pin west.yml to FFmpeg commit 08ab4baa34 containing Xtensa fixed-point DSP kernels. - Document AAC-LC performance profile (~8.0 - 11.5 MCPS on Aphid 400 MHz DSP) in audio README and ffmpeg_dec README.
Integrate vo-aacenc, a lightweight 32-bit fixed-point AAC-LC encoder, into the ffmpeg_dec module in encode mode (CONFIG_FFMPEG_ENC_AAC). Key highlights: - Pure 32-bit fixed-point integer arithmetic (CONFIG_FPU=n). - Inner loops accelerated with Xtensa hardware instructions: * mulsh (32x32 signed high product) * nsa (single-cycle normalization shift count) * clamps (single-cycle 16-bit saturation) * 4-way unrolled band energy and TNS autocorrelation loops. - Generates standard ADTS AAC elementary stream packets. - Advertises SND_AUDIOCODEC_AAC capture capability in base_fw. - Added sof-hda-generic-ffmpeg-aac-enc-compr topology target. - Pinned vo-aacenc repository in west manifest at commit c823d30. - Measured on Aphid (PTL / ACE 3.0, 400 MHz): ~18.5 - 21.5 MCPS (~4.8% DSP utilization at 48 kHz stereo, 128 kbps). Bitstream decode validated with 0 errors. - Documented performance profile in audio and ffmpeg_dec READMEs. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Integrate the top 3 recommended audio DSP offload candidates into Sound Open Firmware for Intel ADSP / ACE architectures: 1. WebRTC HPF (webrtc_hpf): - Cascaded 2nd-order biquad IIR high-pass filter (<80/100 Hz cutoff). - Eliminates DC offset and rejects mechanical vibration / speaker rumble. - Fixed-point Q12/Q13 arithmetic with 32-bit accumulators and saturation. - S16 and S32 format support across up to 4 channels. - Measured on Aphid (PTL / ACE 3.0, 400 MHz): ~0.15 - 1.34 MCPS (<0.35% load). 2. WebRTC AGC (webrtc_agc): - Fixed-point digital automatic gain control and adaptive leveling (digital_agc.c). - Dynamically boosts low-level speech (+9 dB target) and compresses loud peaks. - Multi-channel planar FIFO framing over 10 ms blocks. - Pure integer arithmetic without floating-point emulation (CONFIG_FPU=n). - Measured on Aphid: ~1.45 MCPS per channel (~0.36% load). 3. FFmpeg Lookahead Peak Limiter (alimiter): - Brickwall peak limiter with smooth lookahead attack and release. - Integrated into ffmpeg_dec filter mode (CONFIG_FFMPEG_FILTER_ALIMITER). - Guaranteed ceiling at 0.95 to protect smart speakers and eliminate digital clipping. - Measured on Aphid: ~1.95 MCPS stereo (~0.49% load). Total concurrent DSP footprint for all 3 modules is ~3.55 MCPS (<0.9% of a single 400 MHz DSP core). Verified on Aphid hardware with zero xruns and clean D3 suspend/resume. Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Add Topology 2 pipeline and component support for the top 3 recommended
DSP offload candidates (webrtc_hpf, webrtc_agc, alimiter):
1. Native Lookahead Peak Limiter (alimiter):
- Pure integer fixed-point brickwall limiter with 5 ms lookahead window
and linear release recovery.
- Guaranteed ceiling at 0.950 linear (-0.45 dBFS) to protect DAC and speakers.
- Zero dynamic heap allocation during runtime; safe for real-time 1 ms timer
pipelines on Core 0.
2. Topology Components & Pipeline:
- Added class widgets: webrtc-hpf, webrtc-agc, alimiter.
- Added pipeline host-copier-top3-mixin-playback chaining:
host-copier -> gain -> webrtc-hpf -> webrtc-agc -> alimiter -> mixin.
- Added TOP3_PROCESSING switch to cavs-mixin-mixout-hda.conf and
sof-hda-generic.conf.
- Added CMake targets sof-hda-generic-top3 and sof-hda-top3.
3. Verification on Panther Lake (Aphid DUT / ACE 3.0 / IPC4):
- Verified clean topology parse and module binding.
- Verified audio streaming via aplay and speaker-test at 48 kHz stereo.
- Verified D3 idle suspend/resume and ALSA mixer controls with zero errors.
…widen)
Integrate the next 3 recommended audio DSP offload modules into Sound Open
Firmware for Intel ADSP / ACE architectures:
1. FFmpeg Parametric Equalizer (equalizer):
- 3-band cascaded Direct Form II Transposed IIR biquad filters:
Low Shelf @ 150 Hz (+3.0 dB), Peaking @ 1 kHz (0 dB), High Shelf @ 8 kHz (+2.5 dB).
- Fixed-point Q30 arithmetic with 64-bit accumulators preventing overflow.
- S16 and S32 format support.
- Measured on Aphid (PTL / ACE 3.0, 400 MHz): ~0.74 MCPS (<0.19% load).
2. FFmpeg Dynamic Range Compressor (acompressor):
- Studio downward compressor with peak envelope follower, programmable
threshold (-12 dBFS), ratio (3:1), attack (20 ms), release (250 ms),
and makeup gain (+2 dB).
- Smooth gain attenuation without hard digital clipping.
- S16 and S32 format support.
- Measured on Aphid: ~0.32 MCPS (<0.08% load).
3. FFmpeg Stereo Widener (stereowiden):
- Spatial audio soundstage enlargement using mid/side decorrelation
with 20 ms circular delay, crossfeed (0.3), and feedback (0.3).
- Fixed-point Q30 arithmetic with zero dynamic heap allocation during streaming.
- Measured on Aphid: ~0.28 MCPS (<0.07% load).
4. Topology 2 Components & Pipeline:
- Added class widgets: equalizer, acompressor, stereowiden.
- Added pipeline host-copier-next3-mixin-playback chaining:
host-copier -> gain -> equalizer -> acompressor -> stereowiden -> mixin.
- Added NEXT3_PROCESSING switch to cavs-mixin-mixout-hda.conf and
sof-hda-generic.conf.
- Compiled topologies sof-hda-generic-next3.tplg and sof-hda-next3.tplg.
5. Hardware Verification on Panther Lake (Aphid DUT / ACE 3.0 / IPC4):
- Deployed signed base firmware sof-ptl.ri (with EQUALIZR, COMPRESS, STWIDEN manifests).
- Verified clean topology parse and module binding.
- Verified audio streaming via aplay and speaker-test across S16 and S32 formats.
- Verified D3 idle suspend/resume with active post-resume playback.
- Total next-3 concurrent load is ~1.21 MCPS (<0.31% of a single 400 MHz core).
Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
Combine all 6 audio offload modules into a unified Topology 2 playback pipeline and expose runtime ALSA mixer controls: 1. Pipeline chaining: host-copier -> gain -> webrtc-hpf -> webrtc-agc -> equalizer -> acompressor -> stereowiden -> alimiter -> mixin. 2. Expose 9 runtime ALSA mixer controls via SOF_IPC4_SWITCH_CONTROL_PARAM_ID: - WebRTC HPF: HPF Switch - WebRTC AGC: AGC Switch - Equalizer: EQ Switch, EQ Bass Switch (biquad bass boost) - Compressor: DRC Switch, DRC Heavy Switch (heavy ratio) - Stereo Widen: Wide Switch, Wide+ Switch (extra crossfeed) - Limiter: Limit Switch 3. Support multi-control IPC dispatch (control IDs 0 and 1) in equalizer, acompressor, and stereowiden get_configuration/set_configuration. 4. Enable all 6 modules in board config intel_adsp_ace30_ptl.conf. 5. Successfully verified on Aphid (PTL / ACE 3.0 / IPC4): - Dynamic runtime ON/OFF control switching across all 9 controls. - S16_LE / S32_LE playback at 48 kHz and 16 kHz.
Integrate the fixed-point WebRTC AECm (Acoustic Echo Cancellation Mobile)
module into SOF firmware and topology2 alongside the top 6 playback offload
chain (Option A: top 6 playback chain + capture AECm):
- webrtc_aec:
* Add webrtc_aec-shims.c mapping malloc/calloc/realloc/free to SOF's
rballoc/rfree native user heap allocator with 16-byte aligned headers.
* Support 48kHz <-> 16kHz polyphase resampling to allow 48kHz pipeline
capture while running AECm core at 16kHz wideband.
* Add runtime ALSA mixer switch controls for master bypass and suppression.
* Ensure pure fixed-point execution with 0 floating-point symbols.
- topology2:
* Wire webrtc_aec into cavs-mixin-mixout-hda.conf capture path when
AEC_PROCESSING is enabled.
* Update sof-hda-top6 and tplg-targets to build sof-hda-generic-top6.tplg
with TOP6_PROCESSING=true and AEC_PROCESSING=true.
* Expose ALSA mixer controls: 'Pre Mixer Analog Capture AEC' and
'Pre Mixer Analog Capture AEC Suppression Sw'.
- ace30_ptl:
* Expand CONFIG_SOF_ZEPHYR_HEAP_SIZE to 0x80000 (512 KB) in
intel_adsp_ace30_ptl.conf to prevent system heap exhaustion when
concurrently instantiating playback top 6 modules and capture AEC.
Tested on Panther Lake (Aphid) hardware: verified clean playback, clean
capture, concurrent full-duplex operation, runtime mixer switches, and
D3 suspend/resume.
Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
…ng for PTL
Integrate single-precision floating-point WebRTC AEC into Sound Open
Firmware with support for Panther Lake (intel_ace30_ptl):
- webrtc_aec:
* Implement webrtc_aec-aec.c backend wrapping WebRtcAec_Create/Init/
BufferFarend/Process with mono-mic AEC instance sharing.
* Provide division-free elementary transcendental math in webrtc_aec-math.c
(fast IEEE-754 log2f, exp2f, powf, and 16-pt sinf/cosf).
* Enforce soft-float lowering via -Xclang -target-feature -Xclang -fp
in CMakeLists.txt to prevent emission of unsupported scalar Coprocessor 0
instructions on Panther Lake.
* Optimize compilation with -O3.
* Route WebRTC memory allocations through sof_heap_alloc / sof_heap_free.
- ace30_ptl:
* Expand CONFIG_SOF_ZEPHYR_HEAP_SIZE to 0x180000 (1.5 MB) to accommodate
Float AEC filter states in coherent SRAM.
Tested on Panther Lake (Aphid) hardware: verified clean firmware boot,
runtime ALSA mixer control toggling, capture streaming, and full-duplex
440 Hz reference tone cancellation with 0 xruns and 0 IPC timeouts.
Signed-off-by: Liam Girdwood <liam.r.girdwood@linux.intel.com>
…Lake - Add webrtc_aec-aec3.c component driver and C ABI wrapper for EchoCanceller3 - Provide lightweight C++ runtime and newlib math shims for soft-float Xtensa - Configure WEBRTC_AEC_BACKEND_AEC3 and WEBRTC_AEC_PERF_TELEMETRY in Kconfig - Sizing PTL coherent heap to 1.5MB and LL domain task stack to 16KB - Enable real-time AEC3 execution with sub-115 kcycle/block budget
…architecture and benchmarks
Copilot stopped reviewing on behalf of
lgirdwood due to an error
September 11, 2026 16:25
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Patches still need cleanup, squashing and ready-ness work for upstream.
This pull request introduces a comprehensive embedded audio processing and codec offload suite for Sound Open Firmware (SOF), bringing native on-DSP WebRTC Audio Processing (APM), FFmpeg decoders/encoders/filters, and audio dynamics/filtering modules.
All components are fully integrated with SOF's
module_interface, Topology 2.0 pipelines, ALSA IPC4 mixer controls, and packaged for both static build and LLEXT shared library loading.1. WebRTC Audio Processing Suite
webrtc_aec(Acoustic Echo Cancellation):FastFloatSqrt,FastFloatDiv) using mantissa seed LUT + Newton-Raphson iterations.webrtc_ns(Noise Suppression): Classic fixed-point spectral Wiener filter with quantile background noise tracking (10 ms framing @ 16 kHz).webrtc_ns2(Deep Learning Noise Suppression): Xiph RNNoise hybrid DSP + GRU recurrent neural network (~10k weights) running at locked 48 kHz, featuring speech probability evaluation andNOTIFIER_ID_VADevents.webrtc_vad(Voice Activity Detection): libfvad pure-C GMM classifier with 6-subband pole-zero allpass IIR filterbank, broadcasting real-time voice gating events.webrtc_agc(Automatic Gain Control): Digital speech volume leveling with saturation limiter and streaming framing FIFO (aggregating periods into 10 ms blocks).webrtc_hpf(High-Pass Filter): Cascaded 2nd-order biquad IIR stripping DC offsets and rumble (<80 Hz / <100 Hz) with zero lookahead latency and split-precision accumulator math.2. FFmpeg Audio Codec Suite (
ffmpeg_dec)Unified audio processing module backed by modular
libavcodecandlibavfilter:aac_fixeddecoder with Xtensa SIMD assembly kernels (fixed_dsp_init.c), eliminating all software-emulated float calls and dropping compute load from >900 MCPS to ~8.0–11.5 MCPS.libshine): Fully fixed-point MP3 encoder using single-cycle 32x32 MACs and precalculated tables (~18.0–24.5 MCPS).vo-aacenc): Pure 32-bit fixed-point AAC-LC encoder accelerated with Xtensa hardware instructions (mulsh,nsa,clamps) (~18.5–21.5 MCPS).afftdn: On-DSP STFT Wiener spectral denoiser running vialibavfilter.avutil/txmaximum transform size to 2048 points, reducing.bssstatic footprint from 17 MB down to 415 KB.ffmpeg_dec-shims.c) mappingav_malloc/av_freeto SOFrballocsys/user pools.3. Native DSP Offload Filters
Fixed-point audio enhancement modules ported from FFmpeg algorithms:
equalizer: FFmpegaf_biquadsRBJ 3-band parametric equalizer in Direct Form II Transposed topology (Low Shelf 150 Hz, Peaking 1000 Hz, High Shelf 8000 Hz) with ALSA bass boost toggle.acompressor: FFmpegaf_sidechaincompressdownward dynamic compressor with single-pole asymmetric attack/release peak envelope follower and makeup gain.alimiter: FFmpegalimiterlookahead peak limiter with circular lookahead buffer (5 ms) and guaranteed output ceiling at -0.45 dBFS (0.950 linear).stereowiden: FFmpegaf_stereowidensoundstage spatializer with 20 ms circular delay buffer and crossfeed cancellation.4. Hardware Verification & Telemetry (Panther Lake Aphid @ 400 MHz)
Benchmarked and verified on Intel Panther Lake (
ptl/intel_ace30_ptl) Aphid hardware:aac_fixed+ Xtensa SIMD)vo-aacenc, 128 kbps)shine, 128 kbps)Hardware playback and record verified via
aplay/arecordon Aphid with zero xruns / underruns.5. Submodules & Manifest (
west.yml)Repoints audio submodules to dedicated forks containing embedded Xtensa adaptations:
webrtc-apm:https://github.com/lgirdwood/webrtc-apm(sof-hifi)ffmpeg:https://github.com/lgirdwood/ffmpeg(sof-hifi)vo-aacenc:https://github.com/lgirdwood/vo-aacenc(sof-hifi)shine:https://github.com/lgirdwood/shine(master)libfvad:https://github.com/lgirdwood/libfvad(master)rnnoise:https://github.com/lgirdwood/rnnoise(master)6. Documentation
Comprehensive documentation with Mermaid data flow diagrams, mathematical difference equations, Kconfig tables, and Topology 2 snippets has been added to each module's
README.md:src/audio/webrtc_aec/README.mdsrc/audio/webrtc_ns/README.mdsrc/audio/webrtc_ns2/README.mdsrc/audio/webrtc_vad/README.mdsrc/audio/webrtc_agc/README.mdsrc/audio/webrtc_hpf/README.mdsrc/audio/ffmpeg_dec/README.mdsrc/audio/equalizer/README.mdsrc/audio/acompressor/README.mdsrc/audio/alimiter/README.mdsrc/audio/stereowiden/README.md