Skip to content

HYBRID + DTX: CELT is encoded then discarded on silent frames in 1.6.1 (regression vs 1.5.2) #477

Description

@shuanzhu

Summary

We are migrating our voice application from libopus 1.5.2 to 1.6.1. With the same encoder configuration (HYBRID, SWB, DTX + InbandFEC, DRED disabled), 1.6.1 uses noticeably more CPU on prolonged silence / muted input than 1.5.2.

The reason looks straightforward from the source: on each silent HYBRID frame 1.6.1 still runs the full CELT encode and only then calls decide_dtx_mode, which discards the CELT bitstream and emits a 1-byte DTX packet. In 1.5.2 the same frames stopped right after SILK via if (nBytes==0) return 1; and never entered CELT.

This appears to be a side effect of unifying the generalized DTX path across float / fixed / low-complexity builds, not an intentional design decision about CELT. We'd like to confirm with the maintainers before sending a patch.

Repro

opus_demo-equivalent invocation we use (DRED off):

opus_demo voip 48000 1 60000 \
    -bandwidth SWB -complexity 5 \
    -inbandfec -dtx \
    -dec_complexity 0 -loss 100 \
    Silence.pcm Silence_out.pcm
  • Input: 60 s of digital silence, mono, 48 kHz, 16-bit PCM.
  • With this config the encoder selects MODE_HYBRID.
  • We measure encoder-side CPU averaged over the run.
Build (fixed-point) CPU on silence (relative) DTX path actually taken
1.5.2 1.0× SILK-internal DTX, early return 1 right after silk_Encode
1.6.1 substantially higher (machine-dependent) Generalized DTX: SILK + full CELT encode, then drop the packet

The point is that on these frames 1.6.1's work is a strict superset of 1.5.2's.

Where the change came from

Two commits landed together on 2025-11-12:

  • ad854445"Fix activity when no analysis (fixes DRED and DTX)". Adds the VAD_NO_DECISION fallback that derives activity from silk_mode.signalType after silk_Encode, and removes the #ifndef DISABLE_FLOAT_API guards around decide_dtx_mode and the activity/peak-energy block.
  • 285ef69e"Use silence detection even with float_api off". Extends is_silence and useDTX = use_dtx && !is_silence to the DISABLE_FLOAT_API build.

The stated goal (making generalized DTX work for builds and complexity levels where analysis_info is unavailable) makes sense. The side effect on the encode-on-silence cost just doesn't appear to have been discussed in either commit message.

What the regression looks like in the code

Current src/opus_encoder.c on main (same as 1.6.1):

  1. silk_Encode runs (around line 2213). For HYBRID at low SILK rate over long silence, useDTX is now gated by is_silence, so silk_mode.useDTX ends up false and SILK does not stop early; nBytes != 0, so the early return 1 just after silk_Encode is not taken.
  2. The full CELT pipeline runs (analysis, prefilter, MDCT, bit allocation, range encoding).
  3. Only after CELT, around line 2566, the generalized DTX decision is made:
/* DTX decision */
if (st->use_dtx && !st->silk_mode.useDTX)
{
   if (decide_dtx_mode(activity, &st->nb_no_activity_ms_Q1,
                       2*1000*frame_size/st->Fs))
   {
      st->rangeFinal = 0;
      data[0] = gen_toc(st->mode, st->Fs/frame_size,
                        curr_bandwidth, st->stream_channels);
      RESTORE_STACK;
      return 1;
   }
}

When this triggers, the CELT bitstream produced above is thrown away.

In 1.5.2 fixed-point, silk_mode.useDTX = st->use_dtx; unconditionally, so SILK's own DTX returned nBytes == 0 and the encoder exited before CELT.

Questions for the maintainers

  1. Was this extra encode-on-silence cost (full CELT per silent HYBRID frame, only to be dropped) considered when the DTX/activity path was unified in ad854445 / 285ef69e? The commit messages don't mention it, so we'd like to confirm whether it's intentional or just a consequence of the refactor.

  2. Is there a reason decide_dtx_mode must be called after CELT? From reading the diffs the only hard data dependency is on activity, which is now finalized right after silk_Encode (via the VAD_NO_DECISION fallback). It looks like the decision could be hoisted to just after that fallback, with an early return 1 short-circuiting the CELT block.

  3. If you would accept a patch that:

    • calls decide_dtx_mode immediately after the post-SILK activity fixup, and
    • on a positive DTX decision, skips the CELT encode and returns the same 1-byte packet,

    are there CELT-internal state concerns we should worry about for the first non-silent frame after a long silence (prefilter memory, MDCT overlap, energy prediction, mode-transition behavior)? We can run subjective and objective tests against the standard testvectors, but it would help to know upfront if there's a known transient artifact this would expose.

Happy to send a draft patch and benchmark numbers if there's interest.

Thanks!

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions