Skip to content

0.57.1: streaming real-time budget, stage stats, pre-roll - #460

Merged
michalharakal merged 6 commits into
developfrom
asr/streaming-real-time-budget
Sep 27, 2026
Merged

michalharakal merged 6 commits into
developfrom
asr/streaming-real-time-budget

Conversation

@michalharakal

@michalharakal michalharakal commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Three changes to the Android IREE streaming Moonshine path, measured on an ARM Android device with a Mali GPU. No engine API is touched, so this ships against engine 0.57.0 as a patch.

A real-time budget for the partial decode

The stream did not keep up with the microphone. Per 0.88 s hop it spent ~350 ms on the encoder window and another ~350 ms on a prefill plus greedy steps — RTF 1.2–1.45 — and the whole backlog was still owed at the instant the user stopped speaking.

The two halves of that work are not equally necessary. The frontend, encoder and adapter build the cross memory the final transcript is decoded from. The decode hops only produce text to show while the user is still talking, and finish() re-decodes from that memory regardless, so dropping a hop costs a partial and provably cannot change the final result.

Past MOONSHINE_LAG_BUDGET_MS (default 250) the window runs and the decode is skipped. Disabled when MOONSHINE_FAST_FINISH is on, where the incremental result is the answer and the hops are therefore not optional.

Measured: RTF 1.23 → 0.67, drain 656 ms → 6 ms.

Verified against a 444-utterance German evaluation set, 311 comparable rows (both sides completed a session):

transcripts byte-identical 310 / 311
dispatched actions changed 0
total turn time, median 7876 ms → 7006 ms

The one differing row is a runaway repetition loop on both sides, truncated at a different point; it resolved to the same downstream result.

Two hypotheses checked and discarded before landing on the cause: GPU contention with a resident LLM (a cold process with no model loaded gives identical timings) and a regression in this file (the peek and mid-hop slices are in the original commit). The cost that was hardest to see is the mid-hop slice — it decodes two more tokens on every feed call with no log line, which is why 1.4 s of a cold run sat between two logged events.

stats(): what the last finish cost

These numbers only ever went to logcat, so a caller that wanted to know where recognition spent its time had to scrape the device log — no use for a status surface or a trace span. nativeStats returns them as key=value pairs and IreeMoonshineStream.stats() parses them: the flush/decode split, how far the stream was behind the microphone, dropped decodes, encoder windows and decode hops run, and the utterance's size. A timeline entry carries each event with its timestamp and cost, so the shape of an utterance can be drawn rather than only totalled.

A string map rather than a typed record on purpose: these values cross two more artifacts before anything consumes them, and adding a counter must not change a signature on the way.

Audio before the run is held, not dropped

sendAudioChunk was activeRun?.feed(audioData): with no run open the chunk vanished and nothing said so. That looks harmless because a host should record after the engine is ready — but opening a run costs 229 ms of the 259 ms measured between a push-to-talk trigger and the first recorded sample, and a host that waits loses that much speech off the front.

It was enough to swallow the first word. Sixteen recordings made through a push-to-talk button all began without any leading silence — i.e. mid-utterance — and each one lost the same leading word, which then reappeared in the transcript as a fragment of the second. The ASR was transcribing faithfully; it received the second half.

Chunks now go into a pre-roll and are fed in order once the run opens, ahead of everything else. Bounded twice: one second of size keeping the newest audio, and two seconds of age after which the buffer is discarded entirely — a host that leaves its microphone running between utterances must not have stale speech prepended to the next one.

Measured: 259 ms → 68 ms, and an utterance whose first word was previously lost now transcribes in full.

Also

The CHANGELOG's lock-step sentence read as absolute and made a transformers-only fix look like it needed an engine release. Refined: a release carries at least the engine's X.Y, and the patch may advance on its own.

@michalharakal michalharakal changed the title Asr/streaming real time budget 0.57.1: streaming real-time budget, stage stats, pre-roll Sep 27, 2026
On an ARM Android device with a Mali GPU the streaming decode does not keep up with
the microphone. Per 0.88 s hop it spends about 350 ms on the encoder window and another
350 ms on a prefill plus two greedy steps, so the stream runs at roughly 1.2-1.45x real
time. Nothing absorbs that: the backlog is still there to be worked off at the instant
the user stops speaking, and it is the first of several seconds they wait for an answer.

The two halves of that work are not equally necessary. The frontend, encoder and adapter
build the cross memory the final transcript is decoded from. The decode hops only produce
text to show while the user is still talking, and finish() re-decodes from the memory
regardless -- so dropping a hop costs a partial and provably cannot change the final result.

So measure the lag (wall clock since the first sample, minus the audio that has arrived)
and, past MOONSHINE_LAG_BUDGET_MS, run the window and skip the decode. The decode state is
untouched; the next affordable hop continues from it, and a restart hop rebuilds it anyway.
The budget is off when MOONSHINE_FAST_FINISH is on, where the incremental result *is* the
answer and the hops are therefore not optional.

finish() now also logs how many decodes were dropped and what the residual lag was, so the
effect is read from the log rather than inferred.
…only logging them

The streaming runtime measured where recognition spent its time and then wrote it to
logcat and nowhere else. A caller that wanted the flush/decode split of the final
decode, how far the stream had fallen behind the microphone, or how many partial
decodes the real-time budget gave up had to scrape the device log for
"moonshine-timing" lines -- which is not something a host can build a status screen or
a trace span out of.

nativeStats(handle) returns the counters of the last finish() as key=value pairs, and
IreeMoonshineStream.stats() parses them into a map. They are snapshotted before the
reset that clears them, so they stay readable after the transcript is returned, and are
the same figures the log line prints -- one source, two outlets.

Two counters are new rather than only newly exposed: how many encoder windows ran (the
mandatory work) and how many partial decodes ran (the optional work). Together with
droppedDecodes they say whether a slow turn was the model or the budget.

A string map rather than a typed record on purpose: this travels through the cartridge
and the pipeline before anything consumes it, and adding a counter must not change a
signature at either stop.
nativeStats gave the totals of the last finish(), which is enough to say what recognition
cost but not where the cost sat. A caller that wants to show the shape of an utterance --
which encoder windows ran when, where the early peek fell, which partial decodes the
real-time budget gave up -- had to go back to reading the log.

Add a `timeline` entry: `<kind>:<atMs>:<a>:<b>:<c>` records, kind `w` for an encoder window
with its frontend/encoder/adapter costs, `p` for the early peek, `h` for a partial decode
and `d` for one dropped to the budget. Offsets share their zero with finishAtMs, the first
fed sample.

Capped at MAXEVT (24) entries: the counters keep growing past it and the timeline simply
stops. A status surface wants the shape of a command-and-control utterance, not an
unbounded event log, and the buffer has to stay a fixed size.

Measured on the box, a 1.95 s utterance:
  w:926:249:126:9 | p:926:726:0:0 | w:1846:205:118:5 | d:1846:330:0:0 | w:3037:242:127:5
which reads directly: three windows at ~380 ms each, one 726 ms peek, and every subsequent
partial decode dropped -- the mandatory path is well inside real time, the optional one is
what used to push it past.
…elease

The line read 'a transformers X.Y.Z ships against engine X.Y.Z', which reads as an
absolute lock-step and makes a transformers-only fix look like it needs an engine
release to go out. The actual rule is looser in exactly one direction: a release
carries at least the engine's X.Y, and the patch number may advance on its own.
@michalharakal
michalharakal force-pushed the asr/streaming-real-time-budget branch from 05ac545 to a071ed7 Compare September 27, 2026 06:19
This library is generic; how a caller's users happen to trigger push-to-talk is not
its business. One comment described the measurement set as remote-control recordings.
@michalharakal
michalharakal merged commit 2f80182 into develop Sep 27, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant