cvcGL: fast-draw low-memory mapper (~2/3 fewer GL calls per frame on WebGL2) + reship cvc.16 - #522
Merged
Merged
Conversation
… per-draw shader lookup On GLES3/WebGL2 every cvcGL mesh draws through VTK 9.5.0's vtkOpenGLLowMemoryPolyDataMapper, whose RenderPieceDraw runs four cell-type agents (verts, lines, polys, strips) whether or not they have anything to draw. Each empty one still re-binds every array texture, re-sends every camera/material/shadow/agent uniform and binds+unbinds the VAO: ~110 of the ~160 GL calls a lit mesh costs per draw. Every draw also re-resolves its program through the shader cache (5 source copies, re-substitution, an MD5 of ~16 KB) to find the program it already has. cvc::gl::LowMemoryPolyDataMapper replicates the drawn agent's PreDraw/Draw/PostDraw from VTK 9.5.0 (the agents are private, hidden VTK classes), skips the cell types that cannot render, and re-binds the cached program while ShaderBuildTimeStamp and the window's shader cache are unchanged. The drawn type gets the same uniforms, textures and draw call in the same order, so frames are byte-identical. The stock draw is kept for picking, vertex visibility, and coincident-offset layouts where skipping a type would change the depth offset a drawn type (or the next draw of the same program) sees. The replica compiles only against exactly 9.5.0; elsewhere the class forwards to VTK. It also zero-initialises CoordinateShiftAndScaleInUse/ShiftValues/ ScaleValues, which 9.5.0 leaves uninitialised. newPolyDataMapper() picks the mapper for every cvcGL node (GeometryNode, BBoxNode, GridNode, GeometryShape): CVCGL_LOWMEM_MAPPER=auto (default: LowMemoryPolyDataMapper exactly where VTK's factory hands out the low-memory mapper, i.e. wasm), force (also on desktop GL, for native tests and measurement) or off. DrawPath Stock / AllCellTypes / Fast (default; initial value from CVCGL_LOWMEM_DRAW) is a process-wide switch for A/B.
…t; add a bench cvcgl_lowmem_fastdraw forces the low-memory mapper natively and runs cvcGL's real pipeline (shadow baker -> shadow map -> translucent -> volumetric -> overlay, and shadows off) over every GeometryNode kind plus raw actors whose mapper tallies the GL calls of its own draws. test/gl_call_counter swaps VTK's glad pointers for counting wrappers (with uniform-redundancy tracking). It checks the mapper policy; that DrawPath::AllCellTypes issues exactly VTK's GL calls per entry point on bake and steady frames; that Stock, AllCellTypes and Fast frames are byte-identical RGBA; the per-draw saving (one VAO bind/unbind, no sync query, no shader-cache lookup on a steady frame); lifecycle (shadow-bake rebuilds, a mapper's released resources, programs released in the shader cache, the scene re-opened in a new window); and the guards (offset layouts that need all four agents, picking, vertex visibility). Frames render without MSAA: VTK's default 8x is not frame-to-frame deterministic on NVIDIA (1 LSB at a few edge pixels for an identical GL stream). cvcgl_lowmem_bench (not a test) prints GL calls and draw-stage CPU per frame and per draw, Stock vs Fast, for a demo-style scene with shadows on and off.
…aMapper #512 and the fast-draw mapper landed independently; re-parent the streaming low-memory mapper so streaming overlays (tracks, ribbons, draped links) get the per-draw cuts too. Its draw-range RenderPieceDraw narrows cell group 0 and then calls the base, which draws exactly that group. Its constructor's shift/scale initialisation moves to the base (now applied on every VTK 9.x: 9.5.1+ initialise only the bool).
…ica guards, VTK attribution Review fixes for the fast-draw mapper: - Shader-property changes after the first draw: VTK's low-memory mapper (9.5 through 9.7) never compares the actor's shader property against its built program, so a GeometryNode::add*ShaderReplacement (or a custom uniform declared) after the first draw never reached the GPU on wasm. The GetShaderMTime() > ShaderBuildTimeStamp check that only StreamingLowMemoryPolyDataMapper had moves into LowMemoryPolyDataMapper::RenderPieceStart (every VTK version), and the streaming subclass's copy is removed. A rebuild bumps the stamp, so the next draw also takes the full shader-cache lookup. - The shader-cache skip's invariant is stated in the header: this->Shaders changes only in UpdateShaders, always followed by ShaderBuildTimeStamp.Modified(). vtkDrawTexturedElements::GetShader() is public, so the class hides it with a version that drops the cached program first (as upstream f59c2d0b297 does), and invalidateProgramCache() is there for any other route. - The 9.5.0 replica no longer trusts the version number alone. Contract for the libcvc-deps VTK recipe: any patch to the low-memory / DTE / GLSL-mod code defines VTK_CVC_LOWMEM_PATCHLEVEL in the installed vtkOpenGLLowMemoryPolyDataMapper.h, and the replica turns itself off when it is defined. cvcGL's configure step also hashes the three installed headers the replica builds on against pristine v9.5.0 (line endings normalised) and builds the replica off (CVC_GL_LOWMEM_REPLICA_DISABLED) on any difference. - Records in the header that VTK v9.7.0 (2026-08-14) contains f59c2d0b297 (DTE program re-bind) and 15ed6edc478 (per-program uniform location and value caches) but still runs all four agents, so a VTK upgrade must re-port the empty-agent skip; and that its d1cde2af1a8 adds two glCheckFramebufferStatus per vtkglBlitFramebuffer (4 synchronous calls a frame on WebGL) and should be reverted or avoided for the wasm build. - License: the regions of LowMemoryPolyDataMapper.cpp derived from VTK 9.5.0 (GLSLModCoincidentTopology::GetCoincidentParameters, RenderPieceDraw's agent loop, CellTypeAgent PreDraw/Draw and the Vertices/Lines/Polygons PreDrawInternal) are marked BEGIN/END VTK 9.5.0-DERIVED with VTK's BSD-3-Clause notice (Copyright.txt wording) at the top of the file. THIRD_PARTY_NOTICES.md records it (with the existing XmlRpc++ note), the README's License section points to it, and both the top-level and the standalone cvcGL installs ship it in the doc dir. - A stage observer (null by default; one relaxed load per stage) lets the test interleave the replica's draw stages with a GL trace.
…e bench without wrappers gl_call_counter now wraps every GL entry point VTK 9.5.0's Rendering/OpenGL2, Rendering/VolumeOpenGL2 and Rendering/Core and cvcGL call (a grep of their sources against glad: glVertexAttribDivisor, glBindBufferBase, queries, renderbuffers, glMapBuffer*, glUniform2iv ...) plus the GLES3/GL4 core families they may grow into (glUniform*ui*, the full glUniform*v / glUniformMatrix* set, glClearBuffer*, glTexStorage*, glColorMaski, glBindBufferRange ...), so frame totals are complete. Its "sync" column is now "desktop round-trips (upper bound)": a desktop classification that Firefox partly answers client-side, not a WebGL census. While a trace is open it records every call in order with the bound program (uniforms, draws) or active texture unit (texture state), its argument bytes and the uniform values it sends. glcallsUninstall() restores VTK's pointers. cvcgl_lowmem_fastdraw: - asserts the AllCellTypes frame's trace equals Stock's exactly (bake and steady frames, shadows on and off, the hardware selector's cell and point picking), and that Fast's equals it with the removed work DERIVED from the AllCellTypes trace: the replica marks each draw's stages, and every cell-type block that issued no draw call goes, in every draw where the depth-offset rules allow skipping. A skipped shader-cache lookup removes no GL call (its GL calls are the bind tail the re-bind issues too); the test checks the lookups are accounted for and the stats show the CPU saving. Two desktop-only refinements, both documented in the test: glPointSize/glLineWidth go through vtkOpenGLState's cache, so they are compared as the state each points/lines draw sees; and float values are recorded with -0.0 as +0.0 (VTK's camera key-matrix cache hands out either sign for a zero normal-matrix entry depending on earlier renders -- two Stock selections in a row differ that way). - a Stock-twice control proves the trace is reproducible. - the scene now covers VTK_WIREFRAME and VTK_POINTS representations of a triangle mesh, edge visibility (one that must fall back to all four cell types, one that skips), cell colours on quads (cell map), point colours, texture coordinates, translucency, flat / lit / wide lines, lines as points, a QUADS GeometryNode, the shadow bake and the hardware selector. - GetShader() and invalidateProgramCache() force exactly one full lookup. - a plain GeometryNode's shader replacement added after its first draw, and cleared again, reaches the GPU on every draw path (forced low-memory mapper on desktop; fails without the RenderPieceStart fix). - CVCGL_LOWMEM_DUMP=<dir> writes the traces of a failing comparison. cvcgl_lowmem_bench: every configuration runs a timed pass with the wrappers uninstalled (each mapper draw only reads the clock) and a counted pass for the call counts, so the uniform-redundancy bookkeeping no longer inflates Stock's CPU time. Steady frame, shadows on: draw stage 1.65 -> 0.41 ms, frame 2.33 -> 1.05 ms (was reported 1.93 -> 0.49 / 2.67 -> 1.19 with the wrappers in the timed loop); GL calls unchanged at 2532 -> 831.
…, not coincidentSkipSafe cvcgl_lowmem_fastdraw derived "Fast trace == AllCellTypes minus its empty cell-type blocks" only in the draws whose DrawBegin marker said skipping was allowed -- and that marker was coincidentSkipSafe()'s own answer. With coincidentSkipSafe forced to return true, every ordinary-frame trace check still passed. The oracle no longer asks the mapper: - gl_call_counter decodes a traced glUniform* call into the per-location writes it makes (glcallsUniformWrites; an n-element array upload is n writes, and each value carries its type, so Uniform1i and Uniform1iv agree). - The test replays a trace's uniform writes into a (program, location) -> value map and records, at every draw call, the bound program's state. Two traces see the same uniforms when each draw call sees the same written value at every location (or both see an unwritten one), and -- for the locations a draw reads before the frame writes them -- the frame leaves the same value for the next frame. End values nobody reads before writing are not compared: a lines-only draw on AllCellTypes leaves the empty strips block's cellType/primitiveSize behind, and every block writes both first. - expectFast() removes a draw's empty blocks iff removing them (that draw's alone, its program state on entry treated as unknown) changes no uniform any draw call sees. Fast's GL trace must equal that, Fast's own trace replayed must show every draw call AllCellTypes' uniform state, and the mapper's markers and stats are only counted against the derivation (blocks skipped, fallbacks). Mutation check: coincidentSkipSafe forced true now fails 30 checks, among them the derived-trace and uniform-state checks of every ordinary frame (shadows on/off, first and steady) and the POLYGON_OFFSET frame; forced false fails 82 (the derivation expects the skips). The new oracle also showed the guards scene's POLYGON_OFFSET case was vacuous: the mode was switched on after the shaders were built, and vtkGLSLModCoincidentTopology declares cOffset/cFactor only when a shader is built under an offset, so no offset ever reached the GPU and the expected fallback was for nothing. The scene now sets the mode before its first render and checks that the programs carry the offset uniforms. LowMemoryPolyDataMapper: coincidentSkipSafe() is asked on Fast only (it was also asked on AllCellTypes when observed, for the old derivation); the DrawBegin arg is the draw's decision, for counting. Its comment notes it is conservative where a program lacks the offset uniforms. No change to Fast's GL output. cvcgl_lowmem_fastdraw: 264/264 on the GTX 1650 and on llvmpipe (CVC_REQUIRE_RENDER=1). cvcGL ctest suite 42/42 under default, CVCGL_LOWMEM_MAPPER=force + CVCGL_LOWMEM_DRAW=stock, and force + fast.
… headers change The configure-time check hashes the installed vtkOpenGLLowMemoryPolyDataMapper.h, vtkDrawTexturedElements.h and vtkGLSLModCoincidentTopology.h, but only when CMake runs: a patched VTK re-installed into the same prefix under an existing build directory left the replica on until someone re-ran the configure. The headers that exist are now CMAKE_CONFIGURE_DEPENDS, so the next build re-runs the configure and the check (build.ninja lists all three as RERUN_CMAKE inputs).
The file said the notices for code under other licenses "are recorded here"
but listed only XmlRpc++ and the VTK-derived mapper code. An audit of the tree
(third_party dirs, license and copyright headers, "derived from / ported from"
notes) and of what the libraries compile in adds, each with its notice exactly
as the source carries it (checked mechanically against the files):
- stb_image 2.30 / stb_image_write 1.16 (src/cvc/image/third_party/; MIT or
public domain, Copyright (c) 2017 Sean Barrett) -- always compiled into
libcvc by stb_io.cpp.
- PocketFFT (BSD-3-Clause; not vendored -- the cvcpkg pocketfft package, commit
c90e55b3d5...) -- its header-only templates are compiled into libcvc's
volume_ops with CVC_FFT_PROVIDER=pocketfft, the default.
- readvtk/writevtk in src/cvc/volume/vtk_io.cpp (Technical University of
Denmark; no license terms stated).
- The contour library (Dan Schikore, Emilio Camahort; no license terms stated)
and kazlib's dict.c/dict.h (Kaz Kylheku's free software license), compiled
with CVC_ENABLE_MESHER (default ON).
- sdf (Xiaoyu Zhang, LGPL-2.1) and mtxlib (Dante Treglia II and Mark A.
DeLoura, 2000), compiled with CVC_ENABLE_SDF (default ON).
- VTK's marching-cubes triCases table in inc/cvc/volren/detail/mc_tables.h,
under the same VTK BSD-3 notice as the mapper code.
- ABYSSAL (MIT, Copyright (c) 2026 Davi (Token-Gremlin)), ported in
src/cvcGL/examples/OceanFFT.{h,cpp} for lsystem_coast and
cvcgl_ocean_fft; the notice is upstream's LICENSE verbatim.
- The vendored CMake modules (Kitware BSD, FindPETSc, NVIDIA/SCI FindCUDA MIT),
marked as build scripts that nothing compiles or installs.
The XmlRpc++ entry now quotes its header notice and says which files carry it
(the headers; the .cpp files carry none, base64.h only its author line). The
README's license paragraph points at the full list.
…with cvcgl-examples License metadata for the packages that carry the VTK-derived code in libcvcGL (LowMemoryPolyDataMapper.cpp), as SPDX expressions (the recipe schema's license field is an SPDX expression; jack, libjpeg-turbo and others already use AND): - cvcgl, cvcgl-cuda: "LGPL-2.1-only AND BSD-3-Clause". - cvcgl-examples: "LGPL-2.1-only AND BSD-3-Clause AND MIT" -- the demos link cvcGL statically, and lsystem_coast's OceanFFT is a port of ABYSSAL (MIT). libcvc keeps LGPL-2.1-only here: its package does not contain cvcGL (the recipe leaves CVC_BUILD_CVCGL at its OFF default). cvcgl-examples did not ship THIRD_PARTY_NOTICES.md: build.sh installs only the cvcgl-examples component, and the wasm build installs nothing but the gallery. Now: - native (build.sh / build.ps1): src/cvcGL/examples installs THIRD_PARTY_NOTICES.md and the LICENSE it links to into share/cvcgl-examples/ as part of the cvcgl-examples component; - wasm (build-wasm.sh): the same two files go beside the gallery in share/cvcgl-examples/web/, so the served payload carries them. cvc_revision bumps (this commit): cvcgl and libcvc 15 -> 16 (cvc.15 is already published without the fast-draw mapper), cvcgl-cuda 1 -> 2, cvcgl-examples 3 -> 4.
… not zero On macOS CI (GL 4.1, Apple's software renderer, a core profile) the litLines probe failed "draws issue no desktop round-trip" on every run, shadows on and off. Its draws issue one glGetString per draw on every path, Stock included: VTK 9.5.0's vtkDrawTexturedElements::PreDraw sends a line width above vtkOpenGLRenderWindow::GetMaximumHardwareLineWidth() (1 on a core profile without wide lines) to ReportUnsupportedLineWidth, which logs the warning with glGetString(GL_VERSION). litLines draws at width 3; NVIDIA (10) and llvmpipe (255) draw it, so Stock has none there. The per-probe check now asks that Fast and AllCellTypes add no round-trip to Stock's draw, and that Stock's own be none or exactly that warning (only glGetString, on a probe whose width the driver lacks) -- still zero on every probe on NVIDIA and llvmpipe. The test also prints the GL renderer, version and maximum hardware line width it ran on.
…easured self-variance On macOS CI (Apple's software renderer, GL 4.1) the pixel-exact checks failed in some runs and not in others, shadows on only: AllCellTypes vs Stock 1-4 differing bytes although the two GL traces were identical call for call, Fast vs Stock up to 5, and Stock before a shader-cache release vs Fast after it, or vs Fast in a new window, up to 9. Three of the four ctest reruns across the two jobs had no pixel difference at all. That renderer is not bit-deterministic for one GL stream; NVIDIA and llvmpipe are. The ordered GL-trace equalities stay the exact oracle everywhere. Pixels are now judged against a bound measured in-test: pairs of Stock frames whose traces are identical -- three Stock runs interleaved with the AllCellTypes and Fast runs in each scene, Stock twice across the lifecycle in one window, the old window's frame against the new window's, the guards scene twice -- give the most bytes and the largest per-byte step the renderer differed by on one stream. Every pixel check (Fast/AllCellTypes vs Stock, across a recompile or a new window, a shader replacement cleared) is recorded and judged at the end against the whole run's bound: no more differing bytes, no larger a step. A noisy pair is printed with where it differs; a pair whose traces differ is not a sample. A last check keeps the bound itself small (<= 0.1% of a frame). On NVIDIA and llvmpipe all 18 pairs are byte-identical, so the bound is 0 and every pixel check is still byte-exact (run 3x each with CVC_REQUIRE_RENDER=1). Injected noise in a third of the Stock runs passes within the bound; the same noise in Fast alone, with clean Stock runs, fails every Fast check.
…use what they judge Review of 63b77d8: the old window's Stock frame against the new window's fed the noise bound, while "new window: same frame as the old window" judged nearly that pair, so the check excused itself on every renderer. A cell colour changed by +40 on re-open (scratch build, every path alike) read as 18 bytes / step 24 of self-variance on NVIDIA and passed, and that bound then relaxed every Fast-vs-Stock pixel check too. Which pairs may be samples: - two Stock frames of identical GL traces in ONE window; nothing across the re-open. The new window samples its own Stock twice, and its Stock and Fast frames are judged against the old window's Stock frame; - neither right after a Fast run: whatever a Fast frame leaves that the trace cannot see would otherwise count as noise. The settles end on Stock, and the Stock run right after Fast is judged ("a Stock frame right after Fast == Stock"), in each scene and in the lifecycle; - no compared pair is a sampled pair, and every reference frame of a check is a member of a sampled pair. How a check is judged: against one observed pair in the same scene (shadows on, shadows off, guards) that differed by at least as many bytes AND at least as large a step -- not the most bytes of one pair with the largest step of another, nor another scene's noise. A sampled pair beyond 64 bytes or a step of 8 (was 0.1% of a frame, 307 bytes, no step limit) excuses nothing and fails the run. The old check logged byte counts only (1-9 on macOS); counts that are no multiple of 3 are pixels changed in some channels and not others, i.e. rounding, so the step cap is inferred, not measured -- the next macOS run prints each pair's bytes and step. testLateShaderReplacement's "cleared again" check is exact again (48x48, no shadows, never seen to vary, nothing sampled there). litLines round-trips: without a vtkOpenGLRenderWindow the line-width limit is +inf (fail closed), and VTK's warning explains Stock's round-trips only on a probe that draws lines, as exactly one glGetString per draw. NVIDIA and llvmpipe, CVC_REQUIRE_RENDER=1, 5 runs each: 271 checks pass, all 18 sampled pairs byte-identical, so every pixel check is byte-exact. The +40 mutation now fails on both ("BEYOND every pair of Stock frames sampled in shadows on", 18 bytes, step 24). Injected noise (scratch build): in the reference Stock frames -- the macOS pattern -- passes; the same noise only in Fast, only in the Stock frame after Fast, or in another scene fails; 30 bytes / step 5 fails against pairs of (60, 1) and (2, 8); a sampled pair of 100 bytes, or of step 30, fails the cap.
…ext) Every cvcGL test that renders through a GL context now takes the CTest resource lock cvcgl_gl_context, so ctest never runs two of them at once; the 23 headless tests (listed by name) keep running in parallel. The lock is assigned by exclusion -- every test in the directory minus the headless list -- so a GL test added later is locked unless someone lists it as headless. Why: on the macOS runners every failed attempt of cvcgl_shadow_casters (4 of the 12 attempts across the 8 package-macos jobs of #521-#523) ran alongside cvcgl_shadow_caster_growth and other GL tests under ctest --parallel; no passing attempt overlapped growth, and each retry passed. That is a correlation, not a reproduction -- no Mac reproduced the flake. On every platform, on purpose: the cost is wall time only. On macOS Debug the GL tests' own times summed to ~370 s (cvcgl_volren_node 230 s of it) where they overlapped into ~250 s, so that ctest step can grow by up to ~2 minutes of ~9.
…enderer only Review of a802747: the sample rules narrowed what a self-variance pair could excuse but did not close it. A sampled pair can still share a frame with the checks it judges: comparePaths samples (stock, stock2) and (stock, stock3) while all six of its pixel checks compare against stock; the lifecycle samples (a, a2) while judging (before, a) and (a, b). If that one frame alone is wrong and the other member is right, the pair measures exactly the check's difference and excuses it. Shown on NVIDIA and llvmpipe, both byte-exact renderers: - cellColours cell 0 +4 after the window re-open, undone before a2: "new window: ... within new window: Stock twice", run passed; - cellColours cell 0 +4 during only the first counted Stock run of the shadows-on scene: all six shadows-on pixel checks passed "within shadows on: Stock twice". The relaxation is now Apple-only, since a shared frame could still self-excuse. renderAvailable() keeps the GL renderer and version strings from ReportCapabilities(); Apple-tolerance mode is on when either names Apple (macOS CI: GL "4.1 APPLE-23.1.1"). Everywhere else (exact mode): - a sampled pair of Stock frames of one GL trace is itself a check, "... Stock twice, <frame> is byte-identical on a deterministic renderer", and fails the run if it differs; - every pixel check is byte for byte. On Apple the per-pair, same-scope, 64-byte / step-8 logic is unchanged and each differing pair still prints its bytes and step. The mode is printed once at startup with the renderer and version: pixel checks: exact|apple-tolerance (OpenGL renderer "...", version "...") Both mutations above now fail on NVIDIA and llvmpipe (4 and 10 checks), as does the earlier +40 on re-open (2). With the mode forced to apple-tolerance in a scratch build, every check line matches the previous commit's, mutations included.
transfix
force-pushed
the
perf/cvcgl-lowmem-draw
branch
from
October 2, 2026 05:09
4899052 to
130f8c9
Compare
libcvc and cvcgl 3.4.0+cvc.16 went out as a native publish of master without the fast-draw mapper, and cvcgl-examples is at +cvc.15 natively. The wasm publish packs the recipe revision verbatim, so take the next number above every platform's.
…hes and windows
On macOS CI ("Apple Software Renderer", GL 4.1 APPLE-23.1.1) the first
attempt of every package-macos job on the last three heads failed the same
three pixel checks: Fast after the shader cache's programs were released
and recompiled, the new window's Stock frame against the old window's Stock
frame, and the new window's Fast frame against it -- 7 to 10 differing
bytes, every one at max delta 1, more than any pair of Stock frames sampled
in one window that run. The Stock-against-Stock failure never touches the
fast path: Apple's software renderer rounds differently across a program
recompile or a new context than within one window.
Apple-tolerance mode only: a check whose two frames come from different
program caches or windows is now judged by a fixed envelope, pure LSB
rounding -- max per-byte delta <= 1 and <= 64 differing bytes -- that no
in-window sample widens, so no frame can excuse itself. In-window checks
keep the sampled-pair rule. Every pixel check prints the rule it was judged
by.
Point picking on Apple's renderer may differ by 1 point per prop: Stock
picked 1118 points of one prop where Fast (once also AllCellTypes, whose
GL trace equals Stock's) picked 1117. Cell picking, and every other
renderer, stay exact.
Exact mode (NVIDIA, Mesa llvmpipe) is unchanged: every pair and every check
byte for byte, 288 checks.
transfix
added a commit
that referenced
this pull request
Oct 2, 2026
…e wasm app contract Next free revisions above master, the low-memory draw PR (#522: cvcgl 17, cvcgl-cuda 2, cvcgl-examples 16) and the highest published on any platform (cvcgl +cvc.16, cvcgl-examples +cvc.15, cvcgl-cuda none), so the publish jobs ship share/cvcGL/, cvcGLWasm.cmake and the examples' devtools instead of skipping an already-published name+version.
…y pixel check
macOS CI ("Apple Software Renderer", GL 4.1 APPLE-23.1.1) still failed first
attempts on pixel checks in one window, which the sampled-pair rule judged:
four shadows-on checks (Fast against Stock, and a Stock frame right after Fast
against Stock, first and steady frames) 1 to 2 bytes apart at max delta 1, in
a run where all ten pairs of Stock frames sampled in that window were
byte-identical; on earlier heads, Fast after a mapper's resources were
released, 1 to 8 bytes at max delta 1. Apple's renderer rounds by 1 LSB
nondeterministically within a window too, so no pair a run samples can bound
what its other frames do.
Apple-tolerance mode only: every pixel check -- in one window, across a
program recompile, across a new window, and the late shader replacement's
cleared frame -- now passes iff its two frames are byte-identical, or differ
by LSB rounding alone: max per-byte delta <= 1 (kAppleLsbMaxDelta) and at
most 64 differing bytes (kSelfVarianceMaxBytes). The envelope is fixed, so no
frame can excuse itself. Every difference measured on macOS so far fits it:
1 to 8 bytes in one window, 1 to 10 across a recompile or a new window, all at
max delta 1. The pairs of Stock frames are still sampled and printed, as
diagnostics only; they excuse nothing (the sampled-pair excuse, its cap check
and FramePair are gone). The startup "pixel checks:" line states the rule.
Point picking on Apple's renderer may still differ by 1 point per prop; cell
picking stays exact.
Exact mode (NVIDIA, Mesa llvmpipe) is unchanged: every pair and every check
byte for byte, 288 checks, the same output line for line.
Conflicts:
- cvcpkg/recipes/{cvcgl,cvcgl-cuda,cvcgl-examples}/recipe.yaml: #537 took
cvcgl 18, cvcgl-cuda 3 and cvcgl-examples 17 for the wasm app contract, above
the revisions this branch had set aside (17, 2, 16). Master's bump comments
are kept; this branch's revisions move to the next number above master's and
above the highest published on any platform (cvcgl 16, cvcgl-examples 15,
cvcgl-cuda none): cvcgl 19, cvcgl-cuda 4, cvcgl-examples 18. libcvc stays at
17 (master 15, published 16).
- src/cvcGL/CMakeLists.txt: both installs kept -- master's wasm state shim and
this branch's THIRD_PARTY_NOTICES.md. The test region merged clean: master's
JS-test registration and its REMOVE_ITEM additions, one GL-test lock block.
… the LSB envelope too In Apple-tolerance mode the sampled pairs of Stock frames were only printed. The Stock frame drawn right after AllCellTypes in comparePaths() is compared only through such pairs, so on Apple's renderer it could be off by any amount and pass. Each pair is now a check there as well, against the same fixed envelope as every other pixel check (no byte off by more than 1, at most 64 bytes); a pair that differs at all is still printed as the renderer's self-variance. Exact mode is unchanged, and both modes now run the same 288 checks. Also: kSelfVarianceMaxBytes is renamed kAppleLsbMaxBytes (it bounds the Apple envelope only), and the comments and startup/summary lines that still said the pairs are printed only or judge other frames are updated.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
LowMemoryPolyDataMapper— a cvcGL subclass of VTK 9.5'svtkOpenGLLowMemoryPolyDataMapper(the mapper VTK uses on WebGL2/wasm) that removes per-draw GL work which cannot change a pixel:newPolyDataMapper()for GeometryNode, BBoxNode, GridNode and GeometryShape wherever VTK would hand out the low-memory mapper (wasm);CVCGL_LOWMEM_MAPPER=forceuses it on desktop;LowMemoryPolyDataMapper::setDrawPath()/CVCGL_LOWMEM_DRAW=stock|all|fastfor A/B.StreamingLowMemoryPolyDataMapper(cvcGL: streaming overlay geometry with partial uploads (StreamingGeometryNode, RibbonNode, terrain-hugging links) #512) now derives from it.Measured (native, low-memory mapper forced on desktop; bench scene: textured ground, 20k-box merged city, 6 vehicles, 4 translucent ribbons, 4 line trails, 960×540):
Pixels byte-identical across stock / all-cell-types / fast. In Firefox the GL-call reduction also cuts the out-of-process command volume that triggers WebGL flushes (async present is lost past 10 flushes between presents).
Correctness evidence (
cvcgl_lowmem_fastdraw, 288 checks, NVIDIA + llvmpipe): an ordered GL trace (entry point, bound program / texture unit, argument bytes, uniform values) of the all-cell-types path equals stock's exactly; the fast path's trace equals that minus the empty blocks, with the expected removals derived independently by replaying uniform state (a wrong skip-safety decision fails 30+ checks — mutation-tested); scenes cover polys/lines/points/strips, wireframe and points representations, edge visibility, cell and point colours, texture coordinates, translucency, shadow bake, the hardware selector, and the depth-offset guard. Full cvcGL suite (48 binaries) passes under default, forced-stock and forced-fast.VTK version guard: the replicated agent draw code compiles only against exactly VTK 9.5.0 and switches itself off when a patched VTK defines
VTK_CVC_LOWMEM_PATCHLEVELor when the installed headers differ from pristine 9.5.0 (configure-time hash, re-run on header change). Header notes record that VTK 9.7.0 has upstream equivalents of the program re-bind and adds 4 synchronousglCheckFramebufferStatusper frame (d1cde2af1a8) — relevant before any VTK upgrade for WebGL.Licensing: ~60 lines are derived from VTK 9.5.0 (BSD-3-Clause), marked with BEGIN/END blocks and the full notice. New
THIRD_PARTY_NOTICES.mdlists every third-party component found in the tree (VTK, stb, PocketFFT, XmlRpc++, the contour library, kazlib, sdf, mtxlib, ABYSSAL for the ocean-FFT example, DTU vtk_io);cvcgl/cvcgl-cudadeclareLGPL-2.1-only AND BSD-3-Clause,cvcgl-examplesLGPL-2.1-only AND BSD-3-Clause AND MIT, and the examples packages (native + wasm web payload) now ship the notices.Reship:
cvc_revisionlibcvc 15 → 17 (+cvc.16 went out as a native publish of master without this), cvcgl 18 → 19, cvcgl-cuda 3 → 4, cvcgl-examples 17 → 18. Each is one above master's wasm app contract bump (#537: cvcgl 18, cvcgl-cuda 3, cvcgl-examples 17) and above every 3.4.0 revision already published on any platform (cvcgl +cvc.16, libcvc +cvc.16, cvcgl-examples +cvc.15; no cvcgl-cuda yet); the wasm publish packs the recipe revision verbatim, so it must clear every platform.Apple's renderer: macOS CI draws on Apple's software renderer, which rounds a few bytes 1 LSB apart between frames of an identical GL stream (measured 1–10 bytes, max delta 1, in one window and across windows). There, and only there (detected from the GL renderer/version string), every pixel check in
cvcgl_lowmem_fastdrawis judged by one fixed LSB envelope: byte-identical, or no byte off by more than 1 and at most 64 bytes. That covers AllCellTypes and Fast against Stock, Stock right after Fast, frames across a recompile and across a new window, and the sampled pairs of Stock frames. The envelope belongs to the renderer, not the run, so no frame can widen it or excuse itself, and a shading change or a changed region still fails. Every other renderer (NVIDIA, llvmpipe) stays byte for byte, the GL traces stay the exact oracle everywhere, point picking on Apple allows at most 1 point per prop, and cell picking stays exact.Independently reviewed twice (adversarial); all findings fixed.