From bc32e8c38a6a9c9385b4b0e902107b8bbef714a9 Mon Sep 17 00:00:00 2001 From: "Ian T." Date: Fri, 25 Sep 2026 10:08:01 -0700 Subject: [PATCH 1/2] docs(paper): bring TECHNICAL.tex up to version 2.7.3 The Zenodo record (v1.1.0, 23 Sep) predates the scrape path and the v2.6.0 error handling work, so the note no longer described what the code does. Diffed the paper against HEAD and closed the gaps. Added a subsection, "The pipeline without the LLM" (sec:nollm), under System overview. `dataforge scrape` had only been mentioned in passing, in the agent guide list and as an MCP table row. It now covers: the same politeness gate as a run (robots.txt, per-domain rate limit, Crawl-delay); no key, session or database; page title, Markdown and text; every HTML table extracted as a header row plus data rows, with the three header forms real pages use, colspan and rowspan expanded so the result is rectangular, nested tables kept separate, layout tables skipped and spans capped; the --check rule-based checks (no text, under fifty words, duplicate text); per-table CSV and JSON plus tables.jsonl carrying source_url; and the scrape folder feeding back in as a run's input so collection and generation need not be one decision. Cost control now records the v2.6.0 change: a budget bounds what a working run spends but not what a broken one wastes, so credential and connection failures are exempt from retry, propagate out of the per-chunk handler, and stop dispatch on the first one. Two corrections in Limitations. The claim that tabular data is "flattened or lost" stopped being true in v2.4.5, but only halfway: the extractor serves the no-LLM path, while the dataset pipeline still flattens a table through the Markdown conversion. Routing those rows into chunking is now named as the reachable half of that gap. The lineage paragraph gained the v2.6.0 recurrence, where a try written for one failure class had quietly become the handler for every class. Smaller edits: Processing says what conversion does to a table and points at both paths; the MCP section describes scrape_page as a library call beside explore_site rather than leaving it a bare table row; Discovery cross-references the new subsection; the abstract gains a sentence on the no-LLM path; System overview states that the note describes version 2.7.3. Verified structurally only: refs, cites, braces and environments all balance, but there is no LaTeX toolchain on this machine and no CI job for the .tex, so it has not been through pdflatex. The DOI in the preamble still points at the v1.1.0 record. A new Zenodo version mints a new DOI, so that is set at publish time; the concept DOI in CITATION.cff stays as it is. --- docs/TECHNICAL.tex | 120 +++++++++++++++++++++++++++++++++++++++------ 1 file changed, 104 insertions(+), 16 deletions(-) diff --git a/docs/TECHNICAL.tex b/docs/TECHNICAL.tex index 6d6f960..0da3b11 100644 --- a/docs/TECHNICAL.tex +++ b/docs/TECHNICAL.tex @@ -70,9 +70,13 @@ \texttt{robots.txt} alone would have permitted at least six of them. The LLM judge approved 97.6\% of samples, but the same model generated and judged them, so we treat that figure as unvalidated until a human audit is -complete. We also describe agent access through the Model Context -Protocol, situate the resulting synthetic, fine-tuning-oriented -dataset relative to retrieval-augmented generation as a complementary +complete. We describe the same collection machinery run without the LLM, +as a scrape that extracts a page's text and its tables as structured rows +and whose output can later be fed back in as a run's input, so gathering +a corpus and paying to generate from it need not be the same decision. We +also describe agent access through the Model Context Protocol, situate +the resulting synthetic, fine-tuning-oriented dataset relative to +retrieval-augmented generation as a complementary route to reducing hallucination, discuss the responsible-use considerations inherent to an unattended web-scraping and synthetic-data tool, and situate the system relative to prior work in web-scale data @@ -172,7 +176,8 @@ \section{System overview} returns a distinct exit code for each class of outcome (invalid recipe, zero URLs after filtering, paused, zero approved samples), so it composes with a scheduler or CI pipeline without a human watching the -log. +log. This note describes version~2.7.3; where behaviour changed, the +version that changed it is named. \subsection{Discovery} \label{sec:discovery} @@ -211,9 +216,9 @@ \subsection{Discovery} \texttt{source.language} / \texttt{source.include} / \texttt{source.exclude} / \texttt{source.max\_urls}. The wizard also accepts the folder written by a no-AI scrape -(\texttt{dataforge scrape}) in place of URLs: its pages are stored as the -session's collection, without being fetched again, and the run starts at -processing. +(\texttt{dataforge scrape}, Section~\ref{sec:nollm}) in place of URLs: +its pages are stored as the session's collection, without being fetched +again, and the run starts at processing. \subsection{Collection} \label{sec:collection} @@ -259,7 +264,11 @@ \subsection{Processing} model, so chunk sizes are approximate for models with a different tokenizer. Chunk size and overlap are recipe-configurable; overlap exists so that a fact spanning a chunk boundary is not lost to the generation -stage entirely. +stage entirely. The conversion keeps a table as Markdown, which reads +acceptably to the generation model but is not structured data; the +no-LLM path of Section~\ref{sec:nollm} extracts tables as rows instead, +and Section~\ref{sec:limitations} treats structured input as future +work. \subsection{Generation} \label{sec:generation} @@ -292,6 +301,18 @@ \subsubsection{Cost control} chunks are skipped cleanly and the run reports how many calls were skipped, rather than crashing or continuing to spend. +A budget bounds what a working run may spend; it does not bound what a +broken one wastes in time. Until version~2.6.0 a missing or invalid API +key, or an unreachable provider, was caught per chunk and retried, so a +run with a bad key made three calls for every chunk and printed the same +error once per chunk before ending with nothing. Credential and +connection failures are now classified as fatal: they are exempt from the +retry policy, they propagate out of the per-chunk handler to the +generator and to the streaming agent, and the first one stops dispatch +and cancels the work already queued. A run against forty chunks with a +bad key now ends in about three seconds with a single message naming the +cause. The judge stage does the same, and reports how far it got. + \subsection{Quality} \label{sec:quality} Samples pass through filters in the following order, so that the LLM @@ -386,6 +407,55 @@ \subsubsection{Dataset introspection} also holds a \texttt{run\_summary.json} with timing per stage, LLM calls and cost, budget use, and errors. +\subsection{The pipeline without the LLM} +\label{sec:nollm} +The six stages above answer one question: what does it take to turn a +site into a fine-tuning dataset. A recurring second question is narrower +and much more common: what does this page say, and what is in the table +on it. Answering it through the full pipeline means a session, a +database, an API key and a bill, for work no model needs to do. Since +version~2.5.0 that path exists on its own, as \texttt{dataforge scrape} +(and, for an agent, the \texttt{scrape\_page} tool of +Section~\ref{sec:mcp}). + +A scrape fetches each URL once, in order, through the same client as +every other request DataForge makes, so \texttt{robots.txt}, the +per-domain rate limit and a declared \texttt{Crawl-delay} bind it exactly +as they bind a run (Section~\ref{sec:collection}). No LLM is called, no +API key is read, and nothing is written to the session database; the +output is a folder. Each page yields its title, its Markdown and its +plain text, and every \texttt{} on it is extracted as a header row +plus data rows rather than being flattened into prose. Header detection +accepts the three forms real pages use: a row of \texttt{}, or, on hand-made pages, a first row whose +non-empty cells are entirely bold while the second row's are not; +duplicate header names are made unique. \texttt{colspan} repeats a cell +across the columns it spans and \texttt{rowspan} carries it down into the +rows below, so every row has every column and the result is rectangular. +Nested tables are extracted in their own right and do not fold their text +into the cell containing them, tables used for layout (one row, or one +column with no header) are skipped, and spans are capped so that a +hostile page cannot force a large allocation. Tables are written per page +as CSV and JSON, and together as \texttt{tables.jsonl} with each row's +\texttt{source\_url}, so lineage survives this path too. + +An optional \texttt{--check} pass applies the rule-based checks that need +no model: a page with no extracted text at all (the usual cause being +that it is built by JavaScript), a page under fifty words, and a page +whose normalized text is identical to one already fetched, which is +reported against the URL that produced it first. These are the same +failure modes the quality stage of Section~\ref{sec:quality} would spend +LLM calls to notice. + +The two paths meet again at the folder. A scrape folder can be given to +the wizard in place of URLs (Section~\ref{sec:discovery}): its pages +become the session's collection without being fetched a second time, and +the run starts at processing. Together with the stage exports of +Section~\ref{sec:export}, this means the collection work and the +generation work can be done at different times, by different people, or +on different terms: a corpus can be gathered, inspected and kept while +the decision to spend anything on generation is still open. + \section{Concurrency: streaming vs.\ batch} \label{sec:streaming} @@ -877,9 +947,13 @@ \subsection{A local MCP server} (\texttt{list\_sessions}, \texttt{session\_stats}, \texttt{view\_samples}) invoke the CLI's own \texttt{--json} output rather than reimplementing its queries, and \texttt{validate\_recipe} runs the CLI's dry run, so the -CLI stays the source of truth for those; \texttt{explore\_site} calls the -discovery library directly. Third, and most important, the safety -constraints that matter most are \emph{enforced by the server} rather +CLI stays the source of truth for those; \texttt{explore\_site} and +\texttt{scrape\_page} call the discovery and scrape libraries directly. +\texttt{scrape\_page} in particular gives the no-LLM path of +Section~\ref{sec:nollm} to an agent: asked what a page says, or what is +in the table on it, the agent answers from one read instead of starting +a run that crawls and spends to find out. Third, and most important, +the safety constraints that matter most are \emph{enforced by the server} rather than requested of the model. \texttt{start\_run} refuses a recipe that sets no spending cap (\texttt{generation.max\_cost\_usd} or \texttt{generation.max\_llm\_calls}, Section~\ref{sec:budget}) unless the @@ -1144,17 +1218,31 @@ \section{Limitations and future work} split's output file, now shuffled (Section~\ref{sec:row-shuffle}). Both are recorded here as a reminder that lineage and ordering guarantees need to be verified per output format and per pipeline stage, not -assumed to hold uniformly once established for one. +assumed to hold uniformly once established for one. The same pattern +recurred in version~2.6.0 with error handling rather than lineage: the +per-chunk \texttt{try} that made the generation stage tolerant of one bad +chunk also swallowed the failures that were not per-chunk at all +(Section~\ref{sec:budget}). A handler written for one class of failure +had silently become the handler for every class. The most significant near-term extension, and the one we consider the project's actual next step rather than a background item, is support for \textbf{non-text source material}. The pipeline today assumes its input is HTML prose reachable by an HTTP client: PDFs and images are filtered out of crawl candidates before any request is made, rather than processed, -tabular and structured data embedded in a page is flattened or lost by -the Markdown-conversion step (Section~\ref{sec:processing}), and there is no handling for -audio or video sources at all. Domains where the authoritative content is -a table (regulatory filings, statistical releases), a scanned or +and there is no handling for audio or video sources at all. Structured +data embedded in a page is half-handled. Version~2.4.5 added the table +extractor of Section~\ref{sec:nollm}, so an HTML table can leave the +system as rows, with its headers, its spans expanded and its source URL +attached; but that extractor serves the no-LLM path only. The dataset +pipeline still passes a table through the Markdown-conversion step +(Section~\ref{sec:processing}), which flattens it, so the rows a scrape +can hand to a spreadsheet are not the rows the generation stage sees. +Routing the extractor's output into chunking, so that a chunk can be a +group of rows with their headers rather than a paragraph of run-together +cells, is the smaller half of this gap and the part already within +reach. Domains where the authoritative content is a table (regulatory +filings, statistical releases), a scanned or image-based PDF (many government and legal documents), or a recorded briefing are currently out of scope, not because the downstream pipeline (chunking, generation, quality, leak-aware export) is From 1be7fc3d9d7e0e2cc3fef50fc17691ac82d1d32c Mon Sep 17 00:00:00 2001 From: "Ian T." Date: Fri, 25 Sep 2026 10:46:59 -0700 Subject: [PATCH 2/2] docs(paper): add five figures to TECHNICAL.tex The note had four tables and no figures, and three of its arguments are easier to see than to read: the stage structure, why overlapping the stages matters, and why a random split leaks. All five figures are TikZ or pgfplots in the document itself, so there are no binary assets and they diff in git. - Figure 1 (System overview): the six stages, what passes between them, and where the guarantees attach. Bands mark the three stages that overlap as concurrent pools and the two that cannot. Complements Table 1 by showing structure rather than repeating its columns. - Figure 2 (Concurrency): batch vs streaming as a timeline, with the LLM idle span marked on the batch row. Drawn schematically and the caption says so, because the benchmark ran in streaming mode only and the saving is a design rationale, not a measured result. - Figure 3 (Leak-aware splitting): page to chunks to samples, then the same samples under a random split and under a group-aware split. The random panel marks one page landing on both sides of the boundary. This is the paper's central claim and the one that most wanted a picture. - Figure 4 (Empirical evaluation): per-site stage timings as stacked bars, from evals/results/pipeline_metrics_20260922T202407Z.json. Shows USCIS bound by its Crawl-delay and the Python tutorial bound by the LLM, which the surrounding prose argues. - Figure 5 (Empirical evaluation): the quality funnel, 831 generated to 811 approved, with each gate's rejections broken out. Carries the ordering claim that the free filters run before the judge that costs money. Every figure is referenced from the prose. Colours were checked for colour-vision-deficiency separation against a light surface, and every coloured mark also carries a direct label or a letter, so the figures survive greyscale printing. Verified: compiles clean under tectonic (XeTeX), 24 pages, no errors, and each figure was rendered and inspected. Not yet compiled under pdfLaTeX, which is what Overleaf and arXiv run; pgfplots compat is pinned to 1.16 rather than a newer value so it builds on older TeX Live as well. --- docs/TECHNICAL.tex | 326 ++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 320 insertions(+), 6 deletions(-) diff --git a/docs/TECHNICAL.tex b/docs/TECHNICAL.tex index 0da3b11..6de66d5 100644 --- a/docs/TECHNICAL.tex +++ b/docs/TECHNICAL.tex @@ -10,6 +10,32 @@ \usepackage{xcolor} \usepackage{enumitem} \usepackage{amsmath} +\usepackage{tikz} +\usetikzlibrary{positioning, arrows.meta, fit, backgrounds, calc} +\usepackage{pgfplots} +% 1.16 rather than a newer value: nothing here depends on later behaviour, and +% this compiles on older TeX Live installations as well as current Overleaf. +\pgfplotsset{compat=1.16} + +% Figure palette. Validated for colour-vision deficiency separation against a +% light surface; every coloured mark also carries a direct label, so the +% figures survive greyscale printing. +\definecolor{figblue}{HTML}{2A78D6} +\definecolor{figorange}{HTML}{EB6834} +\definecolor{figaqua}{HTML}{1BAF7A} +\definecolor{figink}{HTML}{0B0B0B} +\definecolor{figmuted}{HTML}{52514E} +\definecolor{figrule}{HTML}{C9C8C3} + +\tikzset{ + stage/.style = {draw=figink, fill=white, rounded corners=2pt, + minimum height=7mm, inner xsep=3pt, font=\small}, + artifact/.style= {font=\footnotesize\itshape, text=figmuted, inner sep=1.5pt, + fill=white}, + note/.style = {font=\footnotesize, text=figmuted, align=left}, + flow/.style = {-{Stealth[length=4pt]}, draw=figink}, + band/.style = {rounded corners=3pt, draw=figrule, line width=0.6pt}, +} \hypersetup{ colorlinks=true, @@ -165,7 +191,57 @@ \section{System overview} \end{tabular} \end{table} -DataForge is organized as six stages (Table~\ref{tab:architecture}): +\begin{figure}[h] + \centering + \begin{tikzpicture}[node distance=6mm and 0mm] + \node[stage] (disc) {Discovery}; + \node[stage, below=of disc] (coll) {Collection}; + \node[stage, below=of coll] (proc) {Processing}; + \node[stage, below=of proc] (gen) {Generation}; + \node[stage, below=of gen] (qual) {Quality}; + \node[stage, below=of qual] (exp) {Export}; + + \draw[flow] (disc) -- node[artifact, right=6mm] {candidate URLs} (coll); + \draw[flow] (coll) -- node[artifact, right=6mm] {pages} (proc); + \draw[flow] (proc) -- node[artifact, right=6mm] {chunks} (gen); + \draw[flow] (gen) -- node[artifact, right=6mm] {candidate samples} (qual); + \draw[flow] (qual) -- node[artifact, right=6mm] {approved samples} (exp); + + \node[note, right=26mm of disc, text width=38mm] + {sitemap first, then a bounded best-first crawl}; + \node[note, right=26mm of coll, text width=38mm] + {\texttt{robots.txt}, per-domain rate limit, \texttt{Crawl-delay}}; + \node[note, right=26mm of gen, text width=38mm] + {$n$ samples per chunk, under one shared cost and call budget}; + \node[note, right=26mm of qual, text width=38mm] + {free filters first; only the judge costs money}; + \node[note, right=26mm of exp, text width=38mm] + {group-aware split; lineage on every record}; + + \begin{scope}[on background layer] + \node[band, fill=figblue!7, fit=(coll)(proc)(gen), + inner xsep=4mm, inner ysep=2mm] (streamband) {}; + \node[band, fill=figorange!7, fit=(qual)(exp), + inner xsep=4mm, inner ysep=2mm] (batchband) {}; + \end{scope} + + \node[note, left=1mm of streamband, text width=24mm, align=right] + {concurrent pools joined by bounded queues}; + \node[note, left=1mm of batchband, text width=24mm, align=right] + {batch only: needs the whole sample set}; + \end{tikzpicture} + \caption{The six stages, what passes between them, and where the + guarantees attach. Collection, Processing and Generation overlap as + concurrent worker pools (Section~\ref{sec:streaming}); Quality and + Export cannot, because deduplication and group-aware splitting are + properties of the whole sample set. Every URL, page, chunk and sample + is checkpointed as it completes, so a resumed run replays only the + unfinished work.} + \label{fig:pipeline} +\end{figure} + +DataForge is organized as six stages (Table~\ref{tab:architecture} and +Figure~\ref{fig:pipeline}): Discovery, Collection, Processing, Generation, Quality, and Export. A run is specified either by a YAML \emph{recipe} --- a declarative, version-controllable description of every stage's parameters --- or by @@ -464,12 +540,63 @@ \section{Concurrency: streaming vs.\ batch} for the entire duration of the crawl. DataForge's streaming mode (\texttt{stream: true}, the recipe default) instead runs collection, chunking, and generation as three concurrent worker pools connected by -bounded queues: +bounded queues (Figure~\ref{fig:streaming}): \[ \text{urls} \;\longrightarrow\; [\text{scrape pool}] \;\longrightarrow\; \text{pages} \;\longrightarrow\; [\text{chunk pool}] \;\longrightarrow\; \text{chunks} \;\longrightarrow\; [\text{LLM pool}] \;\longrightarrow\; \text{samples} \] +\begin{figure}[h] + \centering + \begin{tikzpicture}[ + bar/.style={draw=figink, rounded corners=1pt, minimum height=5mm, + font=\scriptsize, inner sep=1pt, anchor=west}, + rowlbl/.style={font=\small, anchor=east}, + tick/.style={font=\scriptsize, text=figmuted, anchor=north}, + span/.style={{Bar[width=4pt]}-{Bar[width=4pt]}, draw=figorange, + line width=0.6pt}, + ] + % batch + \node[rowlbl] at (-0.2,0) {batch}; + \node[bar, fill=figblue!25, minimum width=36mm] at (0,0) {collection}; + \node[bar, fill=figrule!50, minimum width=7mm] at (3.6,0) {}; + \node[bar, fill=figorange!30, minimum width=36mm] at (4.3,0) {generation}; + + \draw[span] (0,0.55) -- (4.3,0.55); + \node[font=\scriptsize, text=figorange, anchor=south] at (2.15,0.6) + {LLM idle for the whole crawl}; + + % streaming + \node[rowlbl] at (-0.2,-1.5) {streaming}; + \node[bar, fill=figblue!25, minimum width=36mm] at (0,-1.2) {collection}; + \node[bar, fill=figrule!50, minimum width=36mm] at (0.4,-1.8) {processing}; + \node[bar, fill=figorange!30, minimum width=38mm] at (0.8,-2.4) {generation}; + + \draw[figmuted, line width=0.4pt, dashed] (4.6,-3.1) -- (4.6,0.35); + \draw[figmuted, line width=0.4pt, dashed] (7.9,-3.1) -- (7.9,0.35); + + % axis + \draw[draw=figrule, line width=0.6pt, -{Stealth[length=4pt]}] + (0,-3.1) -- (8.4,-3.1); + \node[tick, anchor=west] at (8.4,-3.05) {time}; + \node[tick] at (4.6,-3.15) {streaming done}; + \node[tick] at (7.9,-3.15) {batch done}; + + \node[font=\scriptsize, text=figmuted, anchor=west, align=left] + at (5.0,-1.8) {pool sizes tuned per stage;\\bounded queues apply\\backpressure between them}; + \end{tikzpicture} + \caption{Why the stages are overlapped, drawn schematically. Run + sequentially, the LLM sits idle for the entire crawl, and the crawl is + rate-limit bound rather than compute bound. Streaming runs collection, + chunking and generation as concurrent pools joined by bounded queues, + so generation starts on the first chunks while later pages are still + being fetched. The proportions here are illustrative: the benchmark of + Section~\ref{sec:eval} ran in streaming mode only, so the saving is a + design rationale and not a measured result + (Section~\ref{sec:eval-limits}).} + \label{fig:streaming} +\end{figure} + Pool sizes are tuned independently, since each stage is bound by a different resource: scraping is rate-limit bound (typically one or a few requests per second per domain), while generation is bound by LLM @@ -498,8 +625,8 @@ \section{Leak-aware dataset splitting} \subsection{The problem} A synthetic dataset produced this way has a clustered generative -structure: $n$ samples are drawn per chunk, and multiple chunks are -drawn per page. Two samples generated from the same chunk --- or from +structure (Figure~\ref{fig:leakage}): $n$ samples are drawn per chunk, +and multiple chunks are drawn per page. Two samples generated from the same chunk --- or from adjacent, overlapping chunks of the same page --- draw on the same underlying facts and the same grounding text, even when they are worded differently. This is precisely the condition under @@ -513,6 +640,101 @@ \subsection{The problem} the near-duplicate contamination concerns raised for pretraining corpora \cite{lee2022deduplicating}. +\begin{figure}[h] + \centering + \begin{tikzpicture}[ + every node/.style={font=\footnotesize}, + src/.style={draw=figink, fill=white, rounded corners=2pt, + minimum width=11mm, minimum height=5mm, inner sep=1pt}, + chunk/.style={draw=figmuted, fill=white, rounded corners=1.5pt, + minimum width=7mm, minimum height=4.5mm, inner sep=1pt, + font=\scriptsize}, + pchip/.style={draw=figblue, fill=figblue!12, rounded corners=1.5pt, + minimum width=6mm, minimum height=4.5mm, inner sep=1pt, + font=\scriptsize, text=figink}, + qchip/.style={draw=figaqua, fill=figaqua!12, rounded corners=1.5pt, + minimum width=6mm, minimum height=4.5mm, inner sep=1pt, + font=\scriptsize, text=figink}, + bin/.style={draw=figrule, rounded corners=2pt, inner xsep=2mm, + inner ysep=1.5mm}, + lbl/.style={font=\scriptsize, text=figmuted}, + lnk/.style={draw=figmuted, line width=0.4pt}, + ] + + % ---- (a) how the data is generated ------------------------------------- + \node[src] (P) at (0,0) {page $P$}; + \node[src] (Q) at (4.2,0) {page $Q$}; + + \node[chunk] (c1) at (-0.9,-0.95) {$c_1$}; + \node[chunk] (c2) at ( 0.9,-0.95) {$c_2$}; + \node[chunk] (c3) at ( 4.2,-0.95) {$c_3$}; + + \foreach \a/\b in {P/c1, P/c2, Q/c3} \draw[lnk] (\a) -- (\b); + + \node[pchip] (p1) at (-1.5,-1.95) {$P_1$}; + \node[pchip] (p2) at (-0.3,-1.95) {$P_2$}; + \node[pchip] (p3) at ( 0.3,-1.95) {$P_3$}; + \node[pchip] (p4) at ( 1.5,-1.95) {$P_4$}; + \node[qchip] (q1) at ( 3.6,-1.95) {$Q_1$}; + \node[qchip] (q2) at ( 4.8,-1.95) {$Q_2$}; + + \foreach \a/\b in {c1/p1, c1/p2, c2/p3, c2/p4, c3/q1, c3/q2} + \draw[lnk] (\a) -- (\b); + + \node[lbl, anchor=west] at (5.6,0) {source pages}; + \node[lbl, anchor=west] at (5.6,-0.95) {chunks}; + \node[lbl, anchor=west, align=left] at (5.6,-1.95) + {samples: not independent\\draws, they cluster by page}; + + % ---- (b) the two splits ------------------------------------------------ + \node[pchip] (rt1) at (-0.6,-3.6) {$P_1$}; + \node[pchip] (rt2) at ( 0.1,-3.6) {$P_3$}; + \node[qchip] (rt3) at ( 0.8,-3.6) {$Q_1$}; + \node[bin, fit=(rt1)(rt2)(rt3)] (rtrain) {}; + \node[lbl, left=1mm of rtrain] {train}; + + \node[pchip] (rs1) at (-0.6,-4.5) {$P_2$}; + \node[pchip] (rs2) at ( 0.1,-4.5) {$P_4$}; + \node[qchip] (rs3) at ( 0.8,-4.5) {$Q_2$}; + \node[bin, fit=(rs1)(rs2)(rs3)] (rtest) {}; + \node[lbl, left=1mm of rtest] {test}; + + \draw[figorange, dashed, line width=0.7pt] (rt1) -- (rs1); + \node[lbl, text=figorange, anchor=north, align=center] at (0.1,-5.1) + {$P$ on both sides: a test item is\\answerable from a training paraphrase}; + + \node[font=\small] at (0.1,-2.95) {(a) uniformly random}; + + \node[pchip] (gt1) at (4.0,-3.6) {$P_1$}; + \node[pchip] (gt2) at (4.7,-3.6) {$P_2$}; + \node[pchip] (gt3) at (5.4,-3.6) {$P_3$}; + \node[pchip] (gt4) at (6.1,-3.6) {$P_4$}; + \node[bin, fit=(gt1)(gt2)(gt3)(gt4)] (gtrain) {}; + \node[lbl, left=1mm of gtrain] {train}; + + \node[qchip] (gs1) at (4.0,-4.5) {$Q_1$}; + \node[qchip] (gs2) at (4.7,-4.5) {$Q_2$}; + \node[bin, fit=(gs1)(gs2)(gtrain.east |- gs1)] (gtest) {}; + \node[lbl, left=1mm of gtest] {test}; + + \node[lbl, anchor=north, align=center] at (5.05,-5.1) + {page-disjoint: whole pages move\\as atomic units}; + + \node[font=\small] at (5.05,-2.95) {(b) group-aware}; + + \draw[figrule, line width=0.6pt] (2.6,-2.7) -- (2.6,-5.6); + \end{tikzpicture} + \caption{Why a uniformly random split leaks. Several samples are drawn + per chunk and several chunks per page, so samples cluster by source + page rather than arriving as independent draws. Splitting rows at + random (a) scatters one page's samples across the boundary, and the + held-out score then partly measures memorization of a near-duplicate + seen in training. Group-aware splitting (b) assigns whole pages as + atomic units, and the assignment is verified page-disjoint before any + file is written.} + \label{fig:leakage} +\end{figure} + \subsection{The mitigation} DataForge's export stage implements \emph{group-aware} splitting: the grouping key (\texttt{export.split.group\_by}, default \texttt{page}) @@ -749,7 +971,8 @@ \subsection{Pipeline results} the number of chunks because the judge makes one call per chunk: for example, 128 calls for Ready.gov's 64 chunks. -The stage timings suggest which resource limited each run, although we +The stage timings (Figure~\ref{fig:timings}) suggest which resource +limited each run, although we did not measure LLM idle time directly. USCIS spent 212.7 of its 296 seconds in the streaming stage, close to the 200-second floor that 20 pages at one request per 10 seconds imposes, and a further 54.2 seconds @@ -759,7 +982,98 @@ \subsection{Pipeline results} many chunks as pages, spent 142.5 seconds streaming and 83.8 seconds in the judge, consistent with a run limited by the LLM. -The quality stage rejected 20 of 831 samples (2.4\%). The free filters +\begin{figure}[h] + \centering + \begin{tikzpicture} + \begin{axis}[ + xbar stacked, + bar width=4.5mm, + y=10mm, + width=11cm, + height=5.4cm, + xmin=0, xmax=330, + symbolic y coords={iantoo.space, Python tutorial, USCIS, Ready.gov}, + ytick=data, + xlabel={seconds}, + xlabel style={font=\small, text=figmuted}, + tick label style={font=\small}, + axis lines=left, + axis line style={draw=figrule}, + xmajorgrids, grid style={draw=figrule, line width=0.3pt}, + tickwidth=0pt, + legend style={at={(0.5,1.14)}, anchor=north, legend columns=3, + draw=none, font=\small, column sep=1.5ex}, + point meta=explicit symbolic, + every node near coord/.append style={font=\tiny, text=figink, + anchor=center}, + enlarge y limits=0.22, + ] + \addplot+[fill=figblue!30, draw=figink, nodes near coords] + coordinates {(0.3,Ready.gov) [] (54.2,USCIS) [54] + (31.1,Python tutorial) [31] (0.2,iantoo.space) []}; + \addplot+[fill=figorange!30, draw=figink, nodes near coords] + coordinates {(53.9,Ready.gov) [54] (212.7,USCIS) [213] + (142.5,Python tutorial) [142] (9.0,iantoo.space) [9]}; + \addplot+[fill=figaqua!30, draw=figink, nodes near coords] + coordinates {(36.3,Ready.gov) [36] (27.6,USCIS) [28] + (83.8,Python tutorial) [84] (4.6,iantoo.space) []}; + \legend{discovery, streaming, judge} + \end{axis} + \end{tikzpicture} + \caption{Where each benchmark run spent its time. USCIS is dominated + by its declared \texttt{Crawl-delay}: 212.7 seconds streaming against + the 200-second floor that 20 pages at one request per 10 seconds + imposes, plus 54.2 in discovery fetching a sitemap index and five + sub-sitemaps at the same spacing. The Python tutorial, with no crawl + delay and nine times as many chunks as pages, spends its time in the + LLM instead. Export takes under a second everywhere and is omitted. + These are stage timings, not a direct measurement of LLM idle time, + which the benchmark did not record.} + \label{fig:timings} +\end{figure} + +\begin{figure}[h] + \centering + \begin{tikzpicture}[ + gate/.style={draw=figink, fill=white, rounded corners=2pt, + minimum height=9mm, minimum width=22mm, font=\small, + align=center, inner sep=2pt}, + drop/.style={-{Stealth[length=4pt]}, draw=figorange, line width=0.6pt}, + droplbl/.style={font=\scriptsize, text=figmuted, align=left, anchor=north west}, + cost/.style={font=\scriptsize\itshape, text=figmuted, anchor=south}, + ] + \node[gate] (gen) at (0,0) {831\\generated}; + \node[gate] (free) at (3.4,0) {free\\filters}; + \node[gate] (judge) at (6.8,0) {LLM\\judge}; + \node[gate, fill=figaqua!10, draw=figaqua] (ok) at (10.2,0) {811 approved\\(97.6\%)}; + + \draw[flow] (gen) -- (free); + \draw[flow] (free) -- node[artifact, above=0pt] {820} (judge); + \draw[flow] (judge) -- (ok); + + \node[cost] at (3.4,0.55) {costs nothing}; + \node[cost] at (6.8,0.55) {one call per chunk}; + + \draw[drop] (free) -- (3.4,-1.1); + \node[droplbl] at (3.0,-1.15) + {11 rejected\\9 refer to the source\\2 exact duplicates\\0 near-duplicates}; + + \draw[drop] (judge) -- (6.8,-1.1); + \node[droplbl] at (6.4,-1.15) + {9 rejected\\6 not standalone\\3 scored below 4\\0 ungrounded}; + \end{tikzpicture} + \caption{How the 831 generated samples were filtered. The order is + deliberate: the filters that cost nothing run first, so the judge, the + only stage that spends money, never runs on a sample a free filter + would have rejected. Counts are pooled across the four sites. The + near-duplicate and ungrounded checks never fired, which + Section~\ref{sec:eval-limits} discusses rather than treats as + confirmation that they work.} + \label{fig:funnel} +\end{figure} + +The quality stage rejected 20 of 831 samples (2.4\%), in the order +Figure~\ref{fig:funnel} shows. The free filters rejected 11: 9 caught by the source-reference regular expressions and 2 exact duplicates. The judge rejected 9: 6 as not standalone and 3 with a score of 3, below the minimum of 4. (In these runs both kinds of
} cells, a +row inside \texttt{