feat(extract): export tables and lists as JSON/CSV with provenance - #382
Open
NianJiuZst wants to merge 1 commit into
Open
NianJiuZst wants to merge 1 commit into
NianJiuZst wants to merge 1 commit into
Conversation
NianJiuZst
force-pushed
the
codex/structured-extraction
branch
from
October 1, 2026 04:51
4cffd5e to
b93bef1
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Agents can already read page text and HTML, but exporting orders, reports or search
results requires site-specific scripts and a separate reconstruction step. This
adds a passive extraction primitive that returns columns, rows and their source
together. For example, an order ID such as "000123" remains a string, merged cells
retain their anchors, and a virtual grid declaring 101 rows but loading only two
returns those two rows with incomplete coverage.
Implementation
tool.extractandbsk extract discover|table|list. Discovery returns boundedsession/document handles for semantic containers in the main page, open shadow
roots and available frames, including OOPIFs.
explicit header IDs, spans, footer rows, row/column indices, and schema-driven
list/card extraction with text, link or attribute fields.
row locations, spans, coverage and warnings. Bound collection by rows, columns,
bytes, traversal work and time; only return complete rows.
null positions and record the CSV SHA-256. Optional
--csv-safeescapesformula-like values while retaining originals. Require
--overwriteto replaceexisting output.
browser_inspectaction, using itssession registry and normal runner/queue. Add protocol schemas, focused tests,
CLI/plugin references and focused unit tests.
Document/attachment checks fence collection. Navigation, detached nodes, expired
handles and foreign sessions cannot reuse a discovered target. Cancellation
releases remote object groups. Extraction does not scroll, paginate or change
focus; CSV encoding and file paths stay on the CLI host.
Real-browser validation: 13 scenario groups
I ran this feature end to end in a real Chrome for Testing 153.0.8010.12
browser on macOS / Apple Silicon, using a purpose-built HTML page with realistic
order, report and search-result layouts. The browser had an owned profile, the
built CLI, an isolated foreground daemon and the unpacked extension.
All 13 scenario groups passed through the actual CLI → daemon → extension →
Chrome chain. The HTML test page and the 13-scenario runner were executed
locally and are retained outside this PR. The screenshots and results below
document that run. These are real-browser tests against controlled fixtures,
not claims of testing production customer websites.
Viewport and full-page screenshots were captured through
bsk screenshot. Aseparate Playwright layout check confirmed a 760 px viewport had no document
horizontal overflow and no console warnings/errors. That check is supplementary;
the extraction proof comes from the real CLI → daemon → extension chain.
Screenshots from the test run
Order and report fixtures captured through
bsk screenshot:Full fixture: tables, search results, shadow DOM and frames
Supplementary layout check at 760 px (Playwright)
Automated validation and CI
includes 28 extraction normalization and lifecycle tests.
checks and Cargo/npm skill package checks passed.
unit tests, extension type check and Biome check passed again. Unit tests now
use minimal case-specific DOM inputs instead of importing the HTML test page.
checkouts. CLI and DSH entry points stay below their respective 7,000/4,500-byte
budgets; Cargo/npm package-content checks passed after the documentation update.
The strict all-target Clippy command reports the existing
clippy::redundant_iter_cloneddiagnostic incrates/bsk-cli/src/cli/update.rs:1428. The same diagnostic was reproduced on theunmodified base commit using a fresh target directory. This change does not edit
that file. One full-suite run also hit the existing
idle_exit.rstest whilereading
daemon.jsonafter the two-second idle period; its focused retry passed.The original failed run is retained alongside the final full-suite result.
Hosted CI results are available in the PR Checks tab.
Scope and validation boundaries
Extraction reads the loaded DOM. It does not claim the entire server-side dataset
is complete, follow pagination/infinite scrolling, OCR canvas content, or discover
closed shadow roots. CSV/metadata writes are individually atomic, not one
filesystem transaction; consumers can verify the recorded CSV hash.
This validates desktop Chromium on macOS. DSH integration was tested at the
adapter/unit level; no live DSH-host session or Edge run is claimed.
See the extraction contract for the API and CSV behavior.
The code diff contains implementation, focused unit tests and usage documentation.
The HTML test page, 13-scenario browser runner, screenshots, logs and one-off
JSON/CSV exports are kept outside the source commits. Screenshots are attached
above; the remaining materials are preserved in the local evidence bundle.