Skip to content

Store the entire figpack bundle as the object, not just data.zarr #7

Description

@dimitri-yatsenko

Decision

Store the whole figpack bundle as the object, not just data.zarr. Settled with Jeremy Magland (Flatiron Institute) on 2026-09-04.

The codec today calls value.save(), which emits index.html, assets/, data.zarr, and extension_manifest.json, then uploads only data.zarr (codec.py encode()). FigpackRef.serve_under() reassembles a servable bundle at render time by copying the viewer dist out of the installed figpack package (figpack.__file__ / figpack-figure-dist). The stored object is therefore pure Zarr — byte-identical in kind to what the <zarr> codec stores — and the figpack marker exists only so the platform knows to lay a viewer over it.

Jeremy's position: a figpack figure is the folder. The viewer is not rendering code layered onto data; it is part of the object. Storing the bundle is what makes the figure independent of the figpack package at render time.

Why this is the better design

The render path loses its figpack dependency. Today the dashboard container must install figpack at a version whose viewer dist is compatible with every stored figure. Store the bundle and serving is static file serving. Jeremy on five-year durability: as long as browsers stay backward-compatible the figure keeps working, whereas a Python dependency graph three years out is not a safe bet. The dependency moves entirely to write time, inside the codec.

Extension views become supportable. validate() currently rejects ExtensionView because extension JavaScript has nowhere to live. Jeremy maintains extension libraries built for specific groups, and this is the capability he cared most about: an extension is installed on the machine that populates the table and would never be available in the serving container. Storing the bundle is what makes those views work. Closes the design half of #4.

One object instead of two linked things. The figure's HTML and its data stop being separate artifacts joined by a hash. Sharing is a link to one object.

It stops assuming a layout we do not control. serve_under() currently hardcodes data.zarr and asserts .zmetadata is present. Jeremy flagged this directly: figpack's flexibility means a custom view can name its Zarr folder differently, or hold more than one. Treating the bundle as opaque removes that class of breakage, along with our reach into the private figpack-figure-dist path.

Scope

  • encode(): upload the full bundle directory value.save() produces, not bundle_path / "data.zarr".
  • validate(): drop the ExtensionView rejection and its message. Verify an extension-backed view round-trips.
  • serve_under(): serve the stored bundle directly. No viewer overlay, no import figpack, no manifest synthesis, no data.zarr layout assertion.
  • load() / show() (FigpackRef.load()/show() broken on DataJoint 2.3 (remote get_folder AttributeError; file branch caches None) #3): reconsider. Both are blocked on figpack exposing no data.zarr → FigpackView function. With the bundle stored, browser display needs no FigpackView at all, so show() can be implemented against the bundle. Decide whether load() still has a purpose or should be removed rather than left raising.
  • Column JSON: {path, store, title, description} stays as-is. Title and description are still lifted at encode time so metadata reads without I/O — the lazy-reference property is unaffected by this change.
  • README "Storage Structure", the codec.py class docstring ("File format: Zarr folder (figpack native)"), and the encode() comment stating the store never duplicates viewer code.

No migration. Only demo pipelines hold figpack data today, confirmed in the meeting. Reseed rather than convert.

Deliberately deferred: content-addressed viewer sharing

The bundle adds ~2.2 MB per figure, essentially all one JS bundle (assets/index-*.js; 4 files total in figpack 0.3's dist). Storing that per figure duplicates it across every row.

DataJoint supports content-addressed storage alongside schema-addressed, so the viewer could be stored once per distinct version and referenced. Jeremy's read: in the regime he works in the data is 20 MB to 1 GB against a few MB of viewer, so separating them is not worth it; it only pays when figures are ~1 MB and there are a million of them. His framing, which we should adopt: the object is the folder, and dedup is an optimization layer on top — for instance marking files cacheable — never a change to what the object is.

So: implement the simple case first. Open a separate issue for cacheable-file dedup if a pipeline shows the small-figure, high-count profile. Do not build it now.

Follow-on, not in this issue

Jeremy's observation that the design generalizes: once the codec stores an opaque static bundle, nothing in it is figpack-specific. A <htmlbundle@store> codec accepting any self-contained static HTML + data folder would serve tools beyond figpack. Worth its own issue after this lands, and worth naming in the blog post as the general property.

Also from the meeting, for the record and not for this repo: figpack embeds the generating script in the figure's info panel, and Jeremy suggested structured fields could be added there. Under DataJoint there is no single generating script — the table definition and its make() are the equivalent — so what we would want written into figure info is a provenance reference. Separate conversation with him.

Blocks

The blog post (dj-thought essay visualization-codecs-in-the-schema.md, PRs #92/#93) asserts the current split as a design choice. Target publication is end of September, so this lands first and the essay describes the shipped design.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions