Skip to content

Speed up stats computation for the whole video. - #60

Open
copybara-service[bot] wants to merge 1 commit into
mainfrom
test_995801615
Open

copybara-service[bot] wants to merge 1 commit into
mainfrom
test_995801615

Conversation

@copybara-service

Copy link
Copy Markdown

Speed up stats computation for the whole video.
E.g. on the motion_floor_to_sky_av1_hdr10p.mp4 it went down from ~1 minute to ~7 seconds.

This is achieved in two ways:

  • speed up frame decoding by using the WebCodecs API (if available) to decode frame by frame, rather than repeatedly seeking to the requested time which gets slower and slower the farther away you are from a keyframe
  • speed up the stats computation itself by precomputing LUTs for math functions called many times, and a bunch of other optimizations

Detailed changes:

  1. WebCodecs Hardware Video Decoding (video_decoder.ts):

    • Added sequential hardware decoding of demuxed MP4 samples via WebCodecs
      VideoDecoder, eliminating O(N^2) GOP seek times in HTML5 video.
    • Added backpressure handling via decodeQueueSize and ondequeue.
    • Added Annex E.3 HEVC and AV1 codec string builders (with O(1) bitwise
      compatibility flag reversal) and box serializers.
    • Preserved fallback to HTML5 video seek if WebCodecs is unsupported or fails.
  2. Color Transfer Lookup Tables (image_stats.ts):

    • Precomputed 64K-entry Float32Array LUTs for transfer functions (HLG, PQ,
      sRGB) and HLG OOTF, replacing millions of Math.pow calls with fast table
      lookups.
  3. VideoFrame & WebGL Pipeline Optimizations (image_stats.ts, app.ts):

    • Allowed ImageStats to accept VideoFrame directly, avoiding createImageBitmap
      allocations.
    • Consolidated and simplified WebGL2 resource management, reusing a shared
      context, texture, and framebuffer across frames to eliminate GPU resource
      recreation overhead.
    • Added keepEncoded: false and shared scratch buffer reuse to avoid allocating
      tens of gigabytes of temporary Float32Arrays during batch processing.
    • Optimized getInverseDistribution by reusing bin boundaries, guarded
      getPercentile against division-by-zero, and optimized averageStats math.
  4. UI:

    • Throttled panel re-rendering during dynamic metadata batch calculation.
  5. Tests:

    • Added unit tests for video_decoder (codec strings, bit reversal).
    • Added unit tests for image_stats (ImageBitmap, scratch buffer, averaging).

E.g. on the motion_floor_to_sky_av1_hdr10p.mp4 it went down from ~1 minute to ~7 seconds.

This is achieved in two ways:
- speed up frame decoding by using the WebCodecs API (if available) to decode frame by frame, rather than repeatedly seeking to the requested time which gets slower and slower the farther away you are from a keyframe
- speed up the stats computation itself by precomputing LUTs for math functions called many times, and a bunch of other optimizations

Detailed changes:

1. WebCodecs Hardware Video Decoding (video_decoder.ts):
   - Added sequential hardware decoding of demuxed MP4 samples via WebCodecs
     VideoDecoder, eliminating O(N^2) GOP seek times in HTML5 video.
   - Added backpressure handling via decodeQueueSize and ondequeue.
   - Added Annex E.3 HEVC and AV1 codec string builders (with O(1) bitwise
     compatibility flag reversal) and box serializers.
   - Preserved fallback to HTML5 video seek if WebCodecs is unsupported or fails.

2. Color Transfer Lookup Tables (image_stats.ts):
   - Precomputed 64K-entry Float32Array LUTs for transfer functions (HLG, PQ,
     sRGB) and HLG OOTF, replacing millions of Math.pow calls with fast table
     lookups.

3. VideoFrame & WebGL Pipeline Optimizations (image_stats.ts, app.ts):
   - Allowed ImageStats to accept VideoFrame directly, avoiding createImageBitmap
     allocations.
   - Consolidated and simplified WebGL2 resource management, reusing a shared
     context, texture, and framebuffer across frames to eliminate GPU resource
     recreation overhead.
   - Added keepEncoded: false and shared scratch buffer reuse to avoid allocating
     tens of gigabytes of temporary Float32Arrays during batch processing.
   - Optimized getInverseDistribution by reusing bin boundaries, guarded
     getPercentile against division-by-zero, and optimized averageStats math.

4. UI:
   - Throttled panel re-rendering during dynamic metadata batch calculation.

5. Tests:
   - Added unit tests for video_decoder (codec strings, bit reversal).
   - Added unit tests for image_stats (ImageBitmap, scratch buffer, averaging).

PiperOrigin-RevId: 995801615

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant