Skip to content

feat(extract): export tables and lists as JSON/CSV with provenance - #382

Open
NianJiuZst wants to merge 1 commit into
Tencent:mainfrom
NianJiuZst:codex/structured-extraction
Open

NianJiuZst wants to merge 1 commit into
Tencent:mainfrom
NianJiuZst:codex/structured-extraction

Conversation

@NianJiuZst

@NianJiuZst NianJiuZst commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Agents can already read page text and HTML, but exporting orders, reports or search
results requires site-specific scripts and a separate reconstruction step. This
adds a passive extraction primitive that returns columns, rows and their source
together. For example, an order ID such as "000123" remains a string, merged cells
retain their anchors, and a virtual grid declaring 101 rows but loading only two
returns those two rows with incomplete coverage.

Implementation

  • Add tool.extract and bsk extract discover|table|list. Discovery returns bounded
    session/document handles for semantic containers in the main page, open shadow
    roots and available frames, including OOPIFs.
  • Support HTML/ARIA tables, hierarchical or missing headers, duplicate labels,
    explicit header IDs, spans, footer rows, row/column indices, and schema-driven
    list/card extraction with text, link or attribute fields.
  • Return stable column keys, original string/null values, page/frame provenance,
    row locations, spans, coverage and warnings. Bound collection by rows, columns,
    bytes, traversal work and time; only return complete rows.
  • Add UTF-8 CSV export with provenance columns and a metadata sidecar. Preserve
    null positions and record the CSV SHA-256. Optional --csv-safe escapes
    formula-like values while retaining originals. Require --overwrite to replace
    existing output.
  • Expose extraction through the existing DSH browser_inspect action, using its
    session registry and normal runner/queue. Add protocol schemas, focused tests,
    CLI/plugin references and focused unit tests.

Document/attachment checks fence collection. Navigation, detached nodes, expired
handles and foreign sessions cannot reuse a discovered target. Cancellation
releases remote object groups. Extraction does not scroll, paginate or change
focus; CSV encoding and file paths stay on the CLI host.

Real-browser validation: 13 scenario groups

I ran this feature end to end in a real Chrome for Testing 153.0.8010.12
browser on macOS / Apple Silicon, using a purpose-built HTML page with realistic
order, report and search-result layouts. The browser had an owned profile, the
built CLI, an isolated foreground daemon and the unpacked extension.

All 13 scenario groups passed through the actual CLI → daemon → extension →
Chrome chain.
The HTML test page and the 13-scenario runner were executed
locally and are retained outside this PR. The screenshots and results below
document that run. These are real-browser tests against controlled fixtures,
not claims of testing production customer websites.

# Scenario group Verified result
1 Container discovery Native tables, open shadow roots, same-origin frames and cross-site frames are discovered; an actual OOPIF target is asserted; hidden data is excluded.
2 Orders and provenance Three records; leading zeroes, Chinese text, commas/quotes/newlines and empty cells preserved; page/frame URL and original row positions retained; selector and target extraction agree.
3 Hierarchical headers, merged cells and footers Header paths include 营收 → 本月 / 上月; row labels stay data; merged positions are null with explicit span anchors; footer provenance is separate.
4 Nested, duplicate, headerless and empty tables Parent rows exclude nested cells; duplicate labels keep different keys; missing headers are generated; the first headerless data row is retained.
5 Partial ARIA grid Loaded source rows 51–52 returned with declared row count 101 and incomplete coverage; missing column 3 remains null.
6 Search cards and semantic lists Field schemas return titles, resolved links, summaries and null missing attributes; default semantic-list extraction also succeeds.
7 Shadow/frame target extraction Discovery handles read open-shadow, same-origin and cross-site tables with the correct page and frame provenance.
8 Row limits and invalid targets A one-row cap returns a complete record and marks truncation; ambiguous, invalid, hidden and missing targets fail explicitly.
9 CSV export and metadata CSV quoting, provenance, matching SHA-256, formula-text escaping and refusal to overwrite existing output are verified.
10 Passive behavior Page data, scroll position and focus remain unchanged after extraction.
11 Byte limits A 2048-byte budget returns three complete rows plus metadata, fits the compact JSON budget and marks truncation/incompleteness.
12 Removed nodes and navigation Handles are rejected as stale after their node is removed or the page reloads.
13 Session isolation Another session cannot reuse a discovery handle.

Viewport and full-page screenshots were captured through bsk screenshot. A
separate Playwright layout check confirmed a 760 px viewport had no document
horizontal overflow and no console warnings/errors. That check is supplementary;
the extraction proof comes from the real CLI → daemon → extension chain.

Screenshots from the test run

Order and report fixtures captured through bsk screenshot:

Orders and report fixture in real Chrome

Full fixture: tables, search results, shadow DOM and frames

Full fixture captured through BrowserSkill

Supplementary layout check at 760 px (Playwright)

760 px fixture with no document horizontal overflow

Automated validation and CI

  • Extension suite: 2,449 passed, 122 skipped by the existing opt-in configuration;
    includes 28 extraction normalization and lifecycle tests.
  • DSH plugin suite: 451 passed.
  • CSV-focused Rust tests: 3 passed.
  • Final full Rust workspace suite: 833 passed, 1 ignored.
  • Extension and plugin type checks/builds passed.
  • Biome, Rust formatting, generated protocol schemas, skill reference/budget
    checks and Cargo/npm skill package checks passed.
  • Production-target Clippy passed with warnings denied.
  • After removing the local browser-test materials from the PR, the 28 extraction
    unit tests, extension type check and Biome check passed again. Unit tests now
    use minimal case-specific DOM inputs instead of importing the HTML test page.
  • Node script suite: 13 passed, including both authored skills in LF and CRLF
    checkouts. CLI and DSH entry points stay below their respective 7,000/4,500-byte
    budgets; Cargo/npm package-content checks passed after the documentation update.

The strict all-target Clippy command reports the existing
clippy::redundant_iter_cloned diagnostic in
crates/bsk-cli/src/cli/update.rs:1428. The same diagnostic was reproduced on the
unmodified base commit using a fresh target directory. This change does not edit
that file. One full-suite run also hit the existing idle_exit.rs test while
reading daemon.json after the two-second idle period; its focused retry passed.
The original failed run is retained alongside the final full-suite result.

Hosted CI results are available in the PR Checks tab.

Scope and validation boundaries

Extraction reads the loaded DOM. It does not claim the entire server-side dataset
is complete, follow pagination/infinite scrolling, OCR canvas content, or discover
closed shadow roots. CSV/metadata writes are individually atomic, not one
filesystem transaction; consumers can verify the recorded CSV hash.

This validates desktop Chromium on macOS. DSH integration was tested at the
adapter/unit level; no live DSH-host session or Edge run is claimed.

See the extraction contract for the API and CSV behavior.
The code diff contains implementation, focused unit tests and usage documentation.
The HTML test page, 13-scenario browser runner, screenshots, logs and one-off
JSON/CSV exports are kept outside the source commits. Screenshots are attached
above; the remaining materials are preserved in the local evidence bundle.

@NianJiuZst
NianJiuZst force-pushed the codex/structured-extraction branch from 4cffd5e to b93bef1 Compare October 1, 2026 04:51

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant