Skip to content

cvcGL: fast-draw low-memory mapper (~2/3 fewer GL calls per frame on WebGL2) + reship cvc.16 - #522

Merged
transfix merged 21 commits into
masterfrom
perf/cvcgl-lowmem-draw
Oct 2, 2026
Merged

transfix merged 21 commits into
masterfrom
perf/cvcgl-lowmem-draw

Conversation

@transfix

@transfix transfix commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

LowMemoryPolyDataMapper — a cvcGL subclass of VTK 9.5's vtkOpenGLLowMemoryPolyDataMapper (the mapper VTK uses on WebGL2/wasm) that removes per-draw GL work which cannot change a pixel:

  • Skip empty cell-type agents. VTK 9.5 runs PreDraw (program/uniform/state setup) for every cell-type agent — verts, lines, polys, strips — even when that type has no cells; ~75% of a draw's ~150–180 GL calls were these. Skipped only where provably safe (falls back to VTK's own draw for picking, vertex visibility and depth-offset layouts where a skip would change coincident-topology uniforms).
  • Re-bind the cached shader program while the shaders have not been rebuilt, instead of the per-draw shader-cache lookup (MD5 of the sources).
  • Selected by newPolyDataMapper() for GeometryNode, BBoxNode, GridNode and GeometryShape wherever VTK would hand out the low-memory mapper (wasm); CVCGL_LOWMEM_MAPPER=force uses it on desktop; LowMemoryPolyDataMapper::setDrawPath() / CVCGL_LOWMEM_DRAW=stock|all|fast for A/B. StreamingLowMemoryPolyDataMapper (cvcGL: streaming overlay geometry with partial uploads (StreamingGeometryNode, RibbonNode, terrain-hugging links) #512) now derives from it.
  • Fix carried along: late shader replacements (added after the first draw, or cleared) now reach the GPU on the low-memory path for every node (VTK 9.5's low-memory mapper ignores shader-property changes after the first build; cvcGL: streaming overlay geometry with partial uploads (StreamingGeometryNode, RibbonNode, terrain-hugging links) #512 had fixed it only for streaming nodes).

Measured (native, low-memory mapper forced on desktop; bench scene: textured ground, 20k-box merged city, 6 vehicles, 4 translucent ribbons, 4 line trails, 960×540):

Frame GL calls/frame stock → fast draw-stage CPU ms frame CPU ms
shadow bake 4371 → 1437 (−67%) 3.02 → 1.94 6.68 → 5.52
steady 2532 → 831 (−67%) 1.65 → 0.41 2.33 → 1.05
shadows off 2227 → 778 (−65%) 1.44 → 0.36 2.21 → 1.10

Pixels byte-identical across stock / all-cell-types / fast. In Firefox the GL-call reduction also cuts the out-of-process command volume that triggers WebGL flushes (async present is lost past 10 flushes between presents).

Correctness evidence (cvcgl_lowmem_fastdraw, 288 checks, NVIDIA + llvmpipe): an ordered GL trace (entry point, bound program / texture unit, argument bytes, uniform values) of the all-cell-types path equals stock's exactly; the fast path's trace equals that minus the empty blocks, with the expected removals derived independently by replaying uniform state (a wrong skip-safety decision fails 30+ checks — mutation-tested); scenes cover polys/lines/points/strips, wireframe and points representations, edge visibility, cell and point colours, texture coordinates, translucency, shadow bake, the hardware selector, and the depth-offset guard. Full cvcGL suite (48 binaries) passes under default, forced-stock and forced-fast.

VTK version guard: the replicated agent draw code compiles only against exactly VTK 9.5.0 and switches itself off when a patched VTK defines VTK_CVC_LOWMEM_PATCHLEVEL or when the installed headers differ from pristine 9.5.0 (configure-time hash, re-run on header change). Header notes record that VTK 9.7.0 has upstream equivalents of the program re-bind and adds 4 synchronous glCheckFramebufferStatus per frame (d1cde2af1a8) — relevant before any VTK upgrade for WebGL.

Licensing: ~60 lines are derived from VTK 9.5.0 (BSD-3-Clause), marked with BEGIN/END blocks and the full notice. New THIRD_PARTY_NOTICES.md lists every third-party component found in the tree (VTK, stb, PocketFFT, XmlRpc++, the contour library, kazlib, sdf, mtxlib, ABYSSAL for the ocean-FFT example, DTU vtk_io); cvcgl/cvcgl-cuda declare LGPL-2.1-only AND BSD-3-Clause, cvcgl-examples LGPL-2.1-only AND BSD-3-Clause AND MIT, and the examples packages (native + wasm web payload) now ship the notices.

Reship: cvc_revision libcvc 15 → 17 (+cvc.16 went out as a native publish of master without this), cvcgl 18 → 19, cvcgl-cuda 3 → 4, cvcgl-examples 17 → 18. Each is one above master's wasm app contract bump (#537: cvcgl 18, cvcgl-cuda 3, cvcgl-examples 17) and above every 3.4.0 revision already published on any platform (cvcgl +cvc.16, libcvc +cvc.16, cvcgl-examples +cvc.15; no cvcgl-cuda yet); the wasm publish packs the recipe revision verbatim, so it must clear every platform.

Apple's renderer: macOS CI draws on Apple's software renderer, which rounds a few bytes 1 LSB apart between frames of an identical GL stream (measured 1–10 bytes, max delta 1, in one window and across windows). There, and only there (detected from the GL renderer/version string), every pixel check in cvcgl_lowmem_fastdraw is judged by one fixed LSB envelope: byte-identical, or no byte off by more than 1 and at most 64 bytes. That covers AllCellTypes and Fast against Stock, Stock right after Fast, frames across a recompile and across a new window, and the sampled pairs of Stock frames. The envelope belongs to the renderer, not the run, so no frame can widen it or excuse itself, and a shading change or a changed region still fails. Every other renderer (NVIDIA, llvmpipe) stays byte for byte, the GL traces stay the exact oracle everywhere, point picking on Apple allows at most 1 point per prop, and cell picking stays exact.

Independently reviewed twice (adversarial); all findings fixed.

… per-draw shader lookup

On GLES3/WebGL2 every cvcGL mesh draws through VTK 9.5.0's
vtkOpenGLLowMemoryPolyDataMapper, whose RenderPieceDraw runs four cell-type
agents (verts, lines, polys, strips) whether or not they have anything to
draw. Each empty one still re-binds every array texture, re-sends every
camera/material/shadow/agent uniform and binds+unbinds the VAO: ~110 of the
~160 GL calls a lit mesh costs per draw. Every draw also re-resolves its
program through the shader cache (5 source copies, re-substitution, an MD5
of ~16 KB) to find the program it already has.

cvc::gl::LowMemoryPolyDataMapper replicates the drawn agent's
PreDraw/Draw/PostDraw from VTK 9.5.0 (the agents are private, hidden VTK
classes), skips the cell types that cannot render, and re-binds the cached
program while ShaderBuildTimeStamp and the window's shader cache are
unchanged. The drawn type gets the same uniforms, textures and draw call in
the same order, so frames are byte-identical. The stock draw is kept for
picking, vertex visibility, and coincident-offset layouts where skipping a
type would change the depth offset a drawn type (or the next draw of the
same program) sees. The replica compiles only against exactly 9.5.0;
elsewhere the class forwards to VTK.

It also zero-initialises CoordinateShiftAndScaleInUse/ShiftValues/
ScaleValues, which 9.5.0 leaves uninitialised.

newPolyDataMapper() picks the mapper for every cvcGL node (GeometryNode,
BBoxNode, GridNode, GeometryShape): CVCGL_LOWMEM_MAPPER=auto (default:
LowMemoryPolyDataMapper exactly where VTK's factory hands out the low-memory
mapper, i.e. wasm), force (also on desktop GL, for native tests and
measurement) or off. DrawPath Stock / AllCellTypes / Fast (default; initial
value from CVCGL_LOWMEM_DRAW) is a process-wide switch for A/B.
…t; add a bench

cvcgl_lowmem_fastdraw forces the low-memory mapper natively and runs cvcGL's
real pipeline (shadow baker -> shadow map -> translucent -> volumetric ->
overlay, and shadows off) over every GeometryNode kind plus raw actors whose
mapper tallies the GL calls of its own draws. test/gl_call_counter swaps
VTK's glad pointers for counting wrappers (with uniform-redundancy tracking).

It checks the mapper policy; that DrawPath::AllCellTypes issues exactly VTK's
GL calls per entry point on bake and steady frames; that Stock, AllCellTypes
and Fast frames are byte-identical RGBA; the per-draw saving (one VAO
bind/unbind, no sync query, no shader-cache lookup on a steady frame);
lifecycle (shadow-bake rebuilds, a mapper's released resources, programs
released in the shader cache, the scene re-opened in a new window); and the
guards (offset layouts that need all four agents, picking, vertex
visibility). Frames render without MSAA: VTK's default 8x is not
frame-to-frame deterministic on NVIDIA (1 LSB at a few edge pixels for an
identical GL stream).

cvcgl_lowmem_bench (not a test) prints GL calls and draw-stage CPU per frame
and per draw, Stock vs Fast, for a demo-style scene with shadows on and off.
…aMapper

#512 and the fast-draw mapper landed independently; re-parent the streaming
low-memory mapper so streaming overlays (tracks, ribbons, draped links) get
the per-draw cuts too. Its draw-range RenderPieceDraw narrows cell group 0
and then calls the base, which draws exactly that group. Its constructor's
shift/scale initialisation moves to the base (now applied on every VTK 9.x:
9.5.1+ initialise only the bool).
…ica guards, VTK attribution

Review fixes for the fast-draw mapper:

- Shader-property changes after the first draw: VTK's low-memory mapper
  (9.5 through 9.7) never compares the actor's shader property against its
  built program, so a GeometryNode::add*ShaderReplacement (or a custom
  uniform declared) after the first draw never reached the GPU on wasm. The
  GetShaderMTime() > ShaderBuildTimeStamp check that only
  StreamingLowMemoryPolyDataMapper had moves into
  LowMemoryPolyDataMapper::RenderPieceStart (every VTK version), and the
  streaming subclass's copy is removed. A rebuild bumps the stamp, so the
  next draw also takes the full shader-cache lookup.

- The shader-cache skip's invariant is stated in the header: this->Shaders
  changes only in UpdateShaders, always followed by
  ShaderBuildTimeStamp.Modified(). vtkDrawTexturedElements::GetShader() is
  public, so the class hides it with a version that drops the cached
  program first (as upstream f59c2d0b297 does), and invalidateProgramCache()
  is there for any other route.

- The 9.5.0 replica no longer trusts the version number alone. Contract for
  the libcvc-deps VTK recipe: any patch to the low-memory / DTE / GLSL-mod
  code defines VTK_CVC_LOWMEM_PATCHLEVEL in the installed
  vtkOpenGLLowMemoryPolyDataMapper.h, and the replica turns itself off when
  it is defined. cvcGL's configure step also hashes the three installed
  headers the replica builds on against pristine v9.5.0 (line endings
  normalised) and builds the replica off (CVC_GL_LOWMEM_REPLICA_DISABLED)
  on any difference.

- Records in the header that VTK v9.7.0 (2026-08-14) contains f59c2d0b297
  (DTE program re-bind) and 15ed6edc478 (per-program uniform location and
  value caches) but still runs all four agents, so a VTK upgrade must
  re-port the empty-agent skip; and that its d1cde2af1a8 adds two
  glCheckFramebufferStatus per vtkglBlitFramebuffer (4 synchronous calls a
  frame on WebGL) and should be reverted or avoided for the wasm build.

- License: the regions of LowMemoryPolyDataMapper.cpp derived from VTK
  9.5.0 (GLSLModCoincidentTopology::GetCoincidentParameters,
  RenderPieceDraw's agent loop, CellTypeAgent PreDraw/Draw and the
  Vertices/Lines/Polygons PreDrawInternal) are marked BEGIN/END VTK
  9.5.0-DERIVED with VTK's BSD-3-Clause notice (Copyright.txt wording) at
  the top of the file. THIRD_PARTY_NOTICES.md records it (with the existing
  XmlRpc++ note), the README's License section points to it, and both the
  top-level and the standalone cvcGL installs ship it in the doc dir.

- A stage observer (null by default; one relaxed load per stage) lets the
  test interleave the replica's draw stages with a GL trace.
…e bench without wrappers

gl_call_counter now wraps every GL entry point VTK 9.5.0's Rendering/OpenGL2,
Rendering/VolumeOpenGL2 and Rendering/Core and cvcGL call (a grep of their
sources against glad: glVertexAttribDivisor, glBindBufferBase, queries,
renderbuffers, glMapBuffer*, glUniform2iv ...) plus the GLES3/GL4 core
families they may grow into (glUniform*ui*, the full glUniform*v /
glUniformMatrix* set, glClearBuffer*, glTexStorage*, glColorMaski,
glBindBufferRange ...), so frame totals are complete. Its "sync" column is
now "desktop round-trips (upper bound)": a desktop classification that
Firefox partly answers client-side, not a WebGL census. While a trace is
open it records every call in order with the bound program (uniforms,
draws) or active texture unit (texture state), its argument bytes and the
uniform values it sends. glcallsUninstall() restores VTK's pointers.

cvcgl_lowmem_fastdraw:
- asserts the AllCellTypes frame's trace equals Stock's exactly (bake and
  steady frames, shadows on and off, the hardware selector's cell and point
  picking), and that Fast's equals it with the removed work DERIVED from
  the AllCellTypes trace: the replica marks each draw's stages, and every
  cell-type block that issued no draw call goes, in every draw where the
  depth-offset rules allow skipping. A skipped shader-cache lookup removes
  no GL call (its GL calls are the bind tail the re-bind issues too); the
  test checks the lookups are accounted for and the stats show the CPU
  saving. Two desktop-only refinements, both documented in the test:
  glPointSize/glLineWidth go through vtkOpenGLState's cache, so they are
  compared as the state each points/lines draw sees; and float values are
  recorded with -0.0 as +0.0 (VTK's camera key-matrix cache hands out
  either sign for a zero normal-matrix entry depending on earlier renders --
  two Stock selections in a row differ that way).
- a Stock-twice control proves the trace is reproducible.
- the scene now covers VTK_WIREFRAME and VTK_POINTS representations of a
  triangle mesh, edge visibility (one that must fall back to all four cell
  types, one that skips), cell colours on quads (cell map), point colours,
  texture coordinates, translucency, flat / lit / wide lines, lines as
  points, a QUADS GeometryNode, the shadow bake and the hardware selector.
- GetShader() and invalidateProgramCache() force exactly one full lookup.
- a plain GeometryNode's shader replacement added after its first draw,
  and cleared again, reaches the GPU on every draw path (forced low-memory
  mapper on desktop; fails without the RenderPieceStart fix).
- CVCGL_LOWMEM_DUMP=<dir> writes the traces of a failing comparison.

cvcgl_lowmem_bench: every configuration runs a timed pass with the
wrappers uninstalled (each mapper draw only reads the clock) and a counted
pass for the call counts, so the uniform-redundancy bookkeeping no longer
inflates Stock's CPU time. Steady frame, shadows on: draw stage 1.65 ->
0.41 ms, frame 2.33 -> 1.05 ms (was reported 1.93 -> 0.49 / 2.67 -> 1.19
with the wrappers in the timed loop); GL calls unchanged at 2532 -> 831.
…, not coincidentSkipSafe

cvcgl_lowmem_fastdraw derived "Fast trace == AllCellTypes minus its empty
cell-type blocks" only in the draws whose DrawBegin marker said skipping was
allowed -- and that marker was coincidentSkipSafe()'s own answer. With
coincidentSkipSafe forced to return true, every ordinary-frame trace check
still passed.

The oracle no longer asks the mapper:

- gl_call_counter decodes a traced glUniform* call into the per-location
  writes it makes (glcallsUniformWrites; an n-element array upload is n
  writes, and each value carries its type, so Uniform1i and Uniform1iv agree).
- The test replays a trace's uniform writes into a (program, location) ->
  value map and records, at every draw call, the bound program's state.
  Two traces see the same uniforms when each draw call sees the same written
  value at every location (or both see an unwritten one), and -- for the
  locations a draw reads before the frame writes them -- the frame leaves the
  same value for the next frame. End values nobody reads before writing are
  not compared: a lines-only draw on AllCellTypes leaves the empty strips
  block's cellType/primitiveSize behind, and every block writes both first.
- expectFast() removes a draw's empty blocks iff removing them (that draw's
  alone, its program state on entry treated as unknown) changes no uniform any
  draw call sees. Fast's GL trace must equal that, Fast's own trace replayed
  must show every draw call AllCellTypes' uniform state, and the mapper's
  markers and stats are only counted against the derivation (blocks skipped,
  fallbacks).

Mutation check: coincidentSkipSafe forced true now fails 30 checks, among
them the derived-trace and uniform-state checks of every ordinary frame
(shadows on/off, first and steady) and the POLYGON_OFFSET frame; forced false
fails 82 (the derivation expects the skips).

The new oracle also showed the guards scene's POLYGON_OFFSET case was
vacuous: the mode was switched on after the shaders were built, and
vtkGLSLModCoincidentTopology declares cOffset/cFactor only when a shader is
built under an offset, so no offset ever reached the GPU and the expected
fallback was for nothing. The scene now sets the mode before its first
render and checks that the programs carry the offset uniforms.

LowMemoryPolyDataMapper: coincidentSkipSafe() is asked on Fast only (it was
also asked on AllCellTypes when observed, for the old derivation); the
DrawBegin arg is the draw's decision, for counting. Its comment notes it is
conservative where a program lacks the offset uniforms. No change to Fast's
GL output.

cvcgl_lowmem_fastdraw: 264/264 on the GTX 1650 and on llvmpipe
(CVC_REQUIRE_RENDER=1). cvcGL ctest suite 42/42 under default,
CVCGL_LOWMEM_MAPPER=force + CVCGL_LOWMEM_DRAW=stock, and force + fast.
… headers change

The configure-time check hashes the installed vtkOpenGLLowMemoryPolyDataMapper.h,
vtkDrawTexturedElements.h and vtkGLSLModCoincidentTopology.h, but only when
CMake runs: a patched VTK re-installed into the same prefix under an existing
build directory left the replica on until someone re-ran the configure. The
headers that exist are now CMAKE_CONFIGURE_DEPENDS, so the next build re-runs
the configure and the check (build.ninja lists all three as RERUN_CMAKE
inputs).
The file said the notices for code under other licenses "are recorded here"
but listed only XmlRpc++ and the VTK-derived mapper code. An audit of the tree
(third_party dirs, license and copyright headers, "derived from / ported from"
notes) and of what the libraries compile in adds, each with its notice exactly
as the source carries it (checked mechanically against the files):

- stb_image 2.30 / stb_image_write 1.16 (src/cvc/image/third_party/; MIT or
  public domain, Copyright (c) 2017 Sean Barrett) -- always compiled into
  libcvc by stb_io.cpp.
- PocketFFT (BSD-3-Clause; not vendored -- the cvcpkg pocketfft package, commit
  c90e55b3d5...) -- its header-only templates are compiled into libcvc's
  volume_ops with CVC_FFT_PROVIDER=pocketfft, the default.
- readvtk/writevtk in src/cvc/volume/vtk_io.cpp (Technical University of
  Denmark; no license terms stated).
- The contour library (Dan Schikore, Emilio Camahort; no license terms stated)
  and kazlib's dict.c/dict.h (Kaz Kylheku's free software license), compiled
  with CVC_ENABLE_MESHER (default ON).
- sdf (Xiaoyu Zhang, LGPL-2.1) and mtxlib (Dante Treglia II and Mark A.
  DeLoura, 2000), compiled with CVC_ENABLE_SDF (default ON).
- VTK's marching-cubes triCases table in inc/cvc/volren/detail/mc_tables.h,
  under the same VTK BSD-3 notice as the mapper code.
- ABYSSAL (MIT, Copyright (c) 2026 Davi (Token-Gremlin)), ported in
  src/cvcGL/examples/OceanFFT.{h,cpp} for lsystem_coast and
  cvcgl_ocean_fft; the notice is upstream's LICENSE verbatim.
- The vendored CMake modules (Kitware BSD, FindPETSc, NVIDIA/SCI FindCUDA MIT),
  marked as build scripts that nothing compiles or installs.

The XmlRpc++ entry now quotes its header notice and says which files carry it
(the headers; the .cpp files carry none, base64.h only its author line). The
README's license paragraph points at the full list.
…with cvcgl-examples

License metadata for the packages that carry the VTK-derived code in
libcvcGL (LowMemoryPolyDataMapper.cpp), as SPDX expressions (the recipe
schema's license field is an SPDX expression; jack, libjpeg-turbo and others
already use AND):

- cvcgl, cvcgl-cuda: "LGPL-2.1-only AND BSD-3-Clause".
- cvcgl-examples: "LGPL-2.1-only AND BSD-3-Clause AND MIT" -- the demos link
  cvcGL statically, and lsystem_coast's OceanFFT is a port of ABYSSAL (MIT).

libcvc keeps LGPL-2.1-only here: its package does not contain cvcGL (the
recipe leaves CVC_BUILD_CVCGL at its OFF default).

cvcgl-examples did not ship THIRD_PARTY_NOTICES.md: build.sh installs only
the cvcgl-examples component, and the wasm build installs nothing but the
gallery. Now:

- native (build.sh / build.ps1): src/cvcGL/examples installs
  THIRD_PARTY_NOTICES.md and the LICENSE it links to into
  share/cvcgl-examples/ as part of the cvcgl-examples component;
- wasm (build-wasm.sh): the same two files go beside the gallery in
  share/cvcgl-examples/web/, so the served payload carries them.

cvc_revision bumps (this commit): cvcgl and libcvc 15 -> 16 (cvc.15 is already
published without the fast-draw mapper), cvcgl-cuda 1 -> 2, cvcgl-examples 3 -> 4.
… not zero

On macOS CI (GL 4.1, Apple's software renderer, a core profile) the
litLines probe failed "draws issue no desktop round-trip" on every run,
shadows on and off. Its draws issue one glGetString per draw on every
path, Stock included: VTK 9.5.0's vtkDrawTexturedElements::PreDraw sends
a line width above vtkOpenGLRenderWindow::GetMaximumHardwareLineWidth()
(1 on a core profile without wide lines) to ReportUnsupportedLineWidth,
which logs the warning with glGetString(GL_VERSION). litLines draws at
width 3; NVIDIA (10) and llvmpipe (255) draw it, so Stock has none there.

The per-probe check now asks that Fast and AllCellTypes add no round-trip
to Stock's draw, and that Stock's own be none or exactly that warning
(only glGetString, on a probe whose width the driver lacks) -- still zero
on every probe on NVIDIA and llvmpipe. The test also prints the GL
renderer, version and maximum hardware line width it ran on.
…easured self-variance

On macOS CI (Apple's software renderer, GL 4.1) the pixel-exact checks
failed in some runs and not in others, shadows on only: AllCellTypes vs
Stock 1-4 differing bytes although the two GL traces were identical call
for call, Fast vs Stock up to 5, and Stock before a shader-cache release
vs Fast after it, or vs Fast in a new window, up to 9. Three of the four
ctest reruns across the two jobs had no pixel difference at all. That
renderer is not bit-deterministic for one GL stream; NVIDIA and llvmpipe
are.

The ordered GL-trace equalities stay the exact oracle everywhere. Pixels
are now judged against a bound measured in-test: pairs of Stock frames
whose traces are identical -- three Stock runs interleaved with the
AllCellTypes and Fast runs in each scene, Stock twice across the lifecycle
in one window, the old window's frame against the new window's, the guards
scene twice -- give the most bytes and the largest per-byte step the
renderer differed by on one stream. Every pixel check (Fast/AllCellTypes vs
Stock, across a recompile or a new window, a shader replacement cleared) is
recorded and judged at the end against the whole run's bound: no more
differing bytes, no larger a step. A noisy pair is printed with where it
differs; a pair whose traces differ is not a sample. A last check keeps the
bound itself small (<= 0.1% of a frame).

On NVIDIA and llvmpipe all 18 pairs are byte-identical, so the bound is 0
and every pixel check is still byte-exact (run 3x each with
CVC_REQUIRE_RENDER=1). Injected noise in a third of the Stock runs passes
within the bound; the same noise in Fast alone, with clean Stock runs,
fails every Fast check.
…use what they judge

Review of 63b77d8: the old window's Stock frame against the new window's
fed the noise bound, while "new window: same frame as the old window"
judged nearly that pair, so the check excused itself on every renderer.
A cell colour changed by +40 on re-open (scratch build, every path alike)
read as 18 bytes / step 24 of self-variance on NVIDIA and passed, and that
bound then relaxed every Fast-vs-Stock pixel check too.

Which pairs may be samples:
- two Stock frames of identical GL traces in ONE window; nothing across the
  re-open. The new window samples its own Stock twice, and its Stock and
  Fast frames are judged against the old window's Stock frame;
- neither right after a Fast run: whatever a Fast frame leaves that the
  trace cannot see would otherwise count as noise. The settles end on
  Stock, and the Stock run right after Fast is judged ("a Stock frame
  right after Fast == Stock"), in each scene and in the lifecycle;
- no compared pair is a sampled pair, and every reference frame of a check
  is a member of a sampled pair.

How a check is judged: against one observed pair in the same scene
(shadows on, shadows off, guards) that differed by at least as many bytes
AND at least as large a step -- not the most bytes of one pair with the
largest step of another, nor another scene's noise. A sampled pair beyond
64 bytes or a step of 8 (was 0.1% of a frame, 307 bytes, no step limit)
excuses nothing and fails the run. The old check logged byte counts only
(1-9 on macOS); counts that are no multiple of 3 are pixels changed in some
channels and not others, i.e. rounding, so the step cap is inferred, not
measured -- the next macOS run prints each pair's bytes and step.

testLateShaderReplacement's "cleared again" check is exact again (48x48,
no shadows, never seen to vary, nothing sampled there).

litLines round-trips: without a vtkOpenGLRenderWindow the line-width limit
is +inf (fail closed), and VTK's warning explains Stock's round-trips only
on a probe that draws lines, as exactly one glGetString per draw.

NVIDIA and llvmpipe, CVC_REQUIRE_RENDER=1, 5 runs each: 271 checks pass,
all 18 sampled pairs byte-identical, so every pixel check is byte-exact.
The +40 mutation now fails on both ("BEYOND every pair of Stock frames
sampled in shadows on", 18 bytes, step 24). Injected noise (scratch build):
in the reference Stock frames -- the macOS pattern -- passes; the same
noise only in Fast, only in the Stock frame after Fast, or in another scene
fails; 30 bytes / step 5 fails against pairs of (60, 1) and (2, 8); a
sampled pair of 100 bytes, or of step 30, fails the cap.
…ext)

Every cvcGL test that renders through a GL context now takes the CTest
resource lock cvcgl_gl_context, so ctest never runs two of them at once;
the 23 headless tests (listed by name) keep running in parallel. The lock
is assigned by exclusion -- every test in the directory minus the headless
list -- so a GL test added later is locked unless someone lists it as
headless.

Why: on the macOS runners every failed attempt of cvcgl_shadow_casters
(4 of the 12 attempts across the 8 package-macos jobs of #521-#523) ran
alongside cvcgl_shadow_caster_growth and other GL tests under ctest
--parallel; no passing attempt overlapped growth, and each retry passed.
That is a correlation, not a reproduction -- no Mac reproduced the flake.

On every platform, on purpose: the cost is wall time only. On macOS Debug
the GL tests' own times summed to ~370 s (cvcgl_volren_node 230 s of it)
where they overlapped into ~250 s, so that ctest step can grow by up to
~2 minutes of ~9.
…enderer only

Review of a802747: the sample rules narrowed what a self-variance pair
could excuse but did not close it. A sampled pair can still share a frame
with the checks it judges: comparePaths samples (stock, stock2) and
(stock, stock3) while all six of its pixel checks compare against stock;
the lifecycle samples (a, a2) while judging (before, a) and (a, b). If that
one frame alone is wrong and the other member is right, the pair measures
exactly the check's difference and excuses it. Shown on NVIDIA and llvmpipe,
both byte-exact renderers:
- cellColours cell 0 +4 after the window re-open, undone before a2:
  "new window: ... within new window: Stock twice", run passed;
- cellColours cell 0 +4 during only the first counted Stock run of the
  shadows-on scene: all six shadows-on pixel checks passed "within shadows
  on: Stock twice".

The relaxation is now Apple-only, since a shared frame could still
self-excuse. renderAvailable() keeps the GL renderer and version strings
from ReportCapabilities(); Apple-tolerance mode is on when either names
Apple (macOS CI: GL "4.1 APPLE-23.1.1"). Everywhere else (exact mode):
- a sampled pair of Stock frames of one GL trace is itself a check,
  "... Stock twice, <frame> is byte-identical on a deterministic renderer",
  and fails the run if it differs;
- every pixel check is byte for byte.
On Apple the per-pair, same-scope, 64-byte / step-8 logic is unchanged and
each differing pair still prints its bytes and step. The mode is printed
once at startup with the renderer and version:
  pixel checks: exact|apple-tolerance (OpenGL renderer "...", version "...")

Both mutations above now fail on NVIDIA and llvmpipe (4 and 10 checks), as
does the earlier +40 on re-open (2). With the mode forced to
apple-tolerance in a scratch build, every check line matches the previous
commit's, mutations included.
libcvc and cvcgl 3.4.0+cvc.16 went out as a native publish of master without the
fast-draw mapper, and cvcgl-examples is at +cvc.15 natively. The wasm publish packs
the recipe revision verbatim, so take the next number above every platform's.
…hes and windows

On macOS CI ("Apple Software Renderer", GL 4.1 APPLE-23.1.1) the first
attempt of every package-macos job on the last three heads failed the same
three pixel checks: Fast after the shader cache's programs were released
and recompiled, the new window's Stock frame against the old window's Stock
frame, and the new window's Fast frame against it -- 7 to 10 differing
bytes, every one at max delta 1, more than any pair of Stock frames sampled
in one window that run. The Stock-against-Stock failure never touches the
fast path: Apple's software renderer rounds differently across a program
recompile or a new context than within one window.

Apple-tolerance mode only: a check whose two frames come from different
program caches or windows is now judged by a fixed envelope, pure LSB
rounding -- max per-byte delta <= 1 and <= 64 differing bytes -- that no
in-window sample widens, so no frame can excuse itself. In-window checks
keep the sampled-pair rule. Every pixel check prints the rule it was judged
by.

Point picking on Apple's renderer may differ by 1 point per prop: Stock
picked 1118 points of one prop where Fast (once also AllCellTypes, whose
GL trace equals Stock's) picked 1117. Cell picking, and every other
renderer, stay exact.

Exact mode (NVIDIA, Mesa llvmpipe) is unchanged: every pair and every check
byte for byte, 288 checks.
transfix added a commit that referenced this pull request Oct 2, 2026
…e wasm app contract

Next free revisions above master, the low-memory draw PR (#522: cvcgl 17,
cvcgl-cuda 2, cvcgl-examples 16) and the highest published on any platform
(cvcgl +cvc.16, cvcgl-examples +cvc.15, cvcgl-cuda none), so the publish
jobs ship share/cvcGL/, cvcGLWasm.cmake and the examples' devtools instead of
skipping an already-published name+version.
…y pixel check

macOS CI ("Apple Software Renderer", GL 4.1 APPLE-23.1.1) still failed first
attempts on pixel checks in one window, which the sampled-pair rule judged:
four shadows-on checks (Fast against Stock, and a Stock frame right after Fast
against Stock, first and steady frames) 1 to 2 bytes apart at max delta 1, in
a run where all ten pairs of Stock frames sampled in that window were
byte-identical; on earlier heads, Fast after a mapper's resources were
released, 1 to 8 bytes at max delta 1. Apple's renderer rounds by 1 LSB
nondeterministically within a window too, so no pair a run samples can bound
what its other frames do.

Apple-tolerance mode only: every pixel check -- in one window, across a
program recompile, across a new window, and the late shader replacement's
cleared frame -- now passes iff its two frames are byte-identical, or differ
by LSB rounding alone: max per-byte delta <= 1 (kAppleLsbMaxDelta) and at
most 64 differing bytes (kSelfVarianceMaxBytes). The envelope is fixed, so no
frame can excuse itself. Every difference measured on macOS so far fits it:
1 to 8 bytes in one window, 1 to 10 across a recompile or a new window, all at
max delta 1. The pairs of Stock frames are still sampled and printed, as
diagnostics only; they excuse nothing (the sampled-pair excuse, its cap check
and FramePair are gone). The startup "pixel checks:" line states the rule.

Point picking on Apple's renderer may still differ by 1 point per prop; cell
picking stays exact.

Exact mode (NVIDIA, Mesa llvmpipe) is unchanged: every pair and every check
byte for byte, 288 checks, the same output line for line.
Conflicts:
- cvcpkg/recipes/{cvcgl,cvcgl-cuda,cvcgl-examples}/recipe.yaml: #537 took
  cvcgl 18, cvcgl-cuda 3 and cvcgl-examples 17 for the wasm app contract, above
  the revisions this branch had set aside (17, 2, 16). Master's bump comments
  are kept; this branch's revisions move to the next number above master's and
  above the highest published on any platform (cvcgl 16, cvcgl-examples 15,
  cvcgl-cuda none): cvcgl 19, cvcgl-cuda 4, cvcgl-examples 18. libcvc stays at
  17 (master 15, published 16).
- src/cvcGL/CMakeLists.txt: both installs kept -- master's wasm state shim and
  this branch's THIRD_PARTY_NOTICES.md. The test region merged clean: master's
  JS-test registration and its REMOVE_ITEM additions, one GL-test lock block.
… the LSB envelope too

In Apple-tolerance mode the sampled pairs of Stock frames were only
printed. The Stock frame drawn right after AllCellTypes in comparePaths()
is compared only through such pairs, so on Apple's renderer it could be
off by any amount and pass. Each pair is now a check there as well,
against the same fixed envelope as every other pixel check (no byte off
by more than 1, at most 64 bytes); a pair that differs at all is still
printed as the renderer's self-variance. Exact mode is unchanged, and
both modes now run the same 288 checks.

Also: kSelfVarianceMaxBytes is renamed kAppleLsbMaxBytes (it bounds the
Apple envelope only), and the comments and startup/summary lines that
still said the pairs are printed only or judge other frames are updated.
@transfix
transfix merged commit 2ebc8f1 into master Oct 2, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant