0.57.1: streaming real-time budget, stage stats, pre-roll - #460
Merged
Merged
Conversation
On an ARM Android device with a Mali GPU the streaming decode does not keep up with the microphone. Per 0.88 s hop it spends about 350 ms on the encoder window and another 350 ms on a prefill plus two greedy steps, so the stream runs at roughly 1.2-1.45x real time. Nothing absorbs that: the backlog is still there to be worked off at the instant the user stops speaking, and it is the first of several seconds they wait for an answer. The two halves of that work are not equally necessary. The frontend, encoder and adapter build the cross memory the final transcript is decoded from. The decode hops only produce text to show while the user is still talking, and finish() re-decodes from the memory regardless -- so dropping a hop costs a partial and provably cannot change the final result. So measure the lag (wall clock since the first sample, minus the audio that has arrived) and, past MOONSHINE_LAG_BUDGET_MS, run the window and skip the decode. The decode state is untouched; the next affordable hop continues from it, and a restart hop rebuilds it anyway. The budget is off when MOONSHINE_FAST_FINISH is on, where the incremental result *is* the answer and the hops are therefore not optional. finish() now also logs how many decodes were dropped and what the residual lag was, so the effect is read from the log rather than inferred.
…only logging them The streaming runtime measured where recognition spent its time and then wrote it to logcat and nowhere else. A caller that wanted the flush/decode split of the final decode, how far the stream had fallen behind the microphone, or how many partial decodes the real-time budget gave up had to scrape the device log for "moonshine-timing" lines -- which is not something a host can build a status screen or a trace span out of. nativeStats(handle) returns the counters of the last finish() as key=value pairs, and IreeMoonshineStream.stats() parses them into a map. They are snapshotted before the reset that clears them, so they stay readable after the transcript is returned, and are the same figures the log line prints -- one source, two outlets. Two counters are new rather than only newly exposed: how many encoder windows ran (the mandatory work) and how many partial decodes ran (the optional work). Together with droppedDecodes they say whether a slow turn was the model or the budget. A string map rather than a typed record on purpose: this travels through the cartridge and the pipeline before anything consumes it, and adding a counter must not change a signature at either stop.
nativeStats gave the totals of the last finish(), which is enough to say what recognition cost but not where the cost sat. A caller that wants to show the shape of an utterance -- which encoder windows ran when, where the early peek fell, which partial decodes the real-time budget gave up -- had to go back to reading the log. Add a `timeline` entry: `<kind>:<atMs>:<a>:<b>:<c>` records, kind `w` for an encoder window with its frontend/encoder/adapter costs, `p` for the early peek, `h` for a partial decode and `d` for one dropped to the budget. Offsets share their zero with finishAtMs, the first fed sample. Capped at MAXEVT (24) entries: the counters keep growing past it and the timeline simply stops. A status surface wants the shape of a command-and-control utterance, not an unbounded event log, and the buffer has to stay a fixed size. Measured on the box, a 1.95 s utterance: w:926:249:126:9 | p:926:726:0:0 | w:1846:205:118:5 | d:1846:330:0:0 | w:3037:242:127:5 which reads directly: three windows at ~380 ms each, one 726 ms peek, and every subsequent partial decode dropped -- the mandatory path is well inside real time, the optional one is what used to push it past.
…elease The line read 'a transformers X.Y.Z ships against engine X.Y.Z', which reads as an absolute lock-step and makes a transformers-only fix look like it needs an engine release to go out. The actual rule is looser in exactly one direction: a release carries at least the engine's X.Y, and the patch number may advance on its own.
michalharakal
force-pushed
the
asr/streaming-real-time-budget
branch
from
September 27, 2026 06:19
05ac545 to
a071ed7
Compare
This library is generic; how a caller's users happen to trigger push-to-talk is not its business. One comment described the measurement set as remote-control recordings.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three changes to the Android IREE streaming Moonshine path, measured on an ARM Android device with a Mali GPU. No engine API is touched, so this ships against engine 0.57.0 as a patch.
A real-time budget for the partial decode
The stream did not keep up with the microphone. Per 0.88 s hop it spent ~350 ms on the encoder window and another ~350 ms on a prefill plus greedy steps — RTF 1.2–1.45 — and the whole backlog was still owed at the instant the user stopped speaking.
The two halves of that work are not equally necessary. The frontend, encoder and adapter build the cross memory the final transcript is decoded from. The decode hops only produce text to show while the user is still talking, and
finish()re-decodes from that memory regardless, so dropping a hop costs a partial and provably cannot change the final result.Past
MOONSHINE_LAG_BUDGET_MS(default 250) the window runs and the decode is skipped. Disabled whenMOONSHINE_FAST_FINISHis on, where the incremental result is the answer and the hops are therefore not optional.Measured: RTF 1.23 → 0.67, drain 656 ms → 6 ms.
Verified against a 444-utterance German evaluation set, 311 comparable rows (both sides completed a session):
The one differing row is a runaway repetition loop on both sides, truncated at a different point; it resolved to the same downstream result.
Two hypotheses checked and discarded before landing on the cause: GPU contention with a resident LLM (a cold process with no model loaded gives identical timings) and a regression in this file (the peek and mid-hop slices are in the original commit). The cost that was hardest to see is the mid-hop slice — it decodes two more tokens on every feed call with no log line, which is why 1.4 s of a cold run sat between two logged events.
stats(): what the last finish costThese numbers only ever went to logcat, so a caller that wanted to know where recognition spent its time had to scrape the device log — no use for a status surface or a trace span.
nativeStatsreturns them askey=valuepairs andIreeMoonshineStream.stats()parses them: the flush/decode split, how far the stream was behind the microphone, dropped decodes, encoder windows and decode hops run, and the utterance's size. Atimelineentry carries each event with its timestamp and cost, so the shape of an utterance can be drawn rather than only totalled.A string map rather than a typed record on purpose: these values cross two more artifacts before anything consumes them, and adding a counter must not change a signature on the way.
Audio before the run is held, not dropped
sendAudioChunkwasactiveRun?.feed(audioData): with no run open the chunk vanished and nothing said so. That looks harmless because a host should record after the engine is ready — but opening a run costs 229 ms of the 259 ms measured between a push-to-talk trigger and the first recorded sample, and a host that waits loses that much speech off the front.It was enough to swallow the first word. Sixteen recordings made through a push-to-talk button all began without any leading silence — i.e. mid-utterance — and each one lost the same leading word, which then reappeared in the transcript as a fragment of the second. The ASR was transcribing faithfully; it received the second half.
Chunks now go into a pre-roll and are fed in order once the run opens, ahead of everything else. Bounded twice: one second of size keeping the newest audio, and two seconds of age after which the buffer is discarded entirely — a host that leaves its microphone running between utterances must not have stale speech prepended to the next one.
Measured: 259 ms → 68 ms, and an utterance whose first word was previously lost now transcribes in full.
Also
The CHANGELOG's lock-step sentence read as absolute and made a transformers-only fix look like it needed an engine release. Refined: a release carries at least the engine's
X.Y, and the patch may advance on its own.