Skip to content

Commit a190b7b

Browse files
donislawdevclaude
andauthored
format: a CSV quotes the fields that need it, and you choose which (#46)
* format: a CSV quotes the fields that need it, and you choose which Adds quote_style to the csv format, taking minimal, all or none. The default is minimal, which changes the bytes of every table this tool writes - the description column used to be quoted on every row and is quoted now only when it carries the separator. This is a breaking change under D11 and it is deliberate. Measured on a 4 kB table at seed 7 before the change: 9 of 44 rows carry a description with no separator in it, because the phrase is three to seven words and drops a separator every third one. Those nine lose their quotes. Sizes are unchanged and every reader that took these files still takes them. The values are the RFC 4180 vocabulary and nothing outside it. A fourth name that preserved today's bytes was rejected: the release this belongs to closes with a major bump either way, so a clean vocabulary costs nothing now and a fourth name would have been carried forever. Two things are not obvious from the list of values. none changes the CONTENT, not only the punctuation. An unquoted field cannot hold a separator without ending early, so the description stops carrying one - in the phrase and in the padding both. A value that only removed the quotes would produce a ragged row at exactly the right size. And the closing row is built to the byte, so under minimal the decision to quote changes the length that the decision depends on. Measured over 59 sizes: the padding first carries a separator at 30 B of description, so the ambiguous band is two sizes per dialect. It is resolved by MEASURING - the quoted length is built first, and if it carries the separator the quotes are earned, otherwise the description is rebuilt to the full room with the separator withheld. A threshold constant was rejected as arithmetic that has to keep agreeing with the bytes beside it. The floor moves with the setting, the way the dialect already does: 115 B rather than 117 under minimal and none, 139 B under all. Twelve combinations, eight distinct floors, all measured with the binary. all quotes the header too, because a header is a row of fields and a writer told to quote everything quotes those as well. The alternative left the structural checker with an "except the first row" exception, which is where a defect hides. The checker is now told the style and judges it, with negative controls in both directions. That matters more here than for the other axes: every one of the three styles produces a well formed file, so nothing about the table gives the style away and a rubber stamp would have been silent. Verified: 1152 files swept over every size from 115 to 260 B across three styles and two delimiters - exact size, six columns under Python's csv module, quoting matching the style, and no needless quote under minimal. All 48 dialect combinations through the structural checker. LibreOffice Calc headless reads all three variants as six columns on every row, which closes D4 for the new values. 16 of 16 mutations caught. Two of those mutations were NOT CAUGHT at first, and the fault was in the guard: both only shift the ANNOUNCED floor, which is invisible elsewhere because Shortest is the worst draw and a real row sits some forty bytes under it. The only handle is the count of distinct floors, and it was written as six where the axes make eight. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci: state the coverage timeout, and guard the README against the registry Two things, and the first is why CI was red. The coverage gate died on a limit nobody had chosen. Go allows ten minutes per package by default while that job allows twenty for all of it, so internal/guard hit the first without coming near the second - a stack trace out of whichever test was running when the alarm went off, instead of a failure naming something. Exactly what the race job above it met on 2026-08-25 and fixed the same way. Measured on the runner rather than guessed: the step took 375 s and 457 s on two consecutive main runs, which is 22 percent of variance on code that barely moved, and the default cuts in at 600 s. This branch added eight seconds of coverage instrumented work - measured, both new CSV guards together - and tipped it. Eight seconds is not what went wrong. 457 against 600 was never a margin. The second thing is O176: the README settings table had no guard and disagreed with the registry on seven rows. log said "none" while carrying eight settings, zip and targz listed three of eight, and avif and jxl had no row at all. The site has had this guard since it was built. The one page a visitor reads first did not. Two guards rather than one, because the table turned out to be the second half of the problem. The list at the top of the README was missing jxl outright - it arrived on 2026-08-31 as the twenty fourth format and never reached that list, so the page offered twenty three while the binary shipped twenty four. The prose said "twenty two" in three places and "24" in two: one file, three different numbers about one thing. Settings are compared as SETS. Whether a row reads "width, height" or the other way round is a question about English, and a guard answering it would be refusing prose rather than catching a lie. The count spelled out in words is deliberately not guarded. That was measured and rejected on 2026-08-05 in copiednumbers_test.go: the general form raised 43 findings and most were false, because "24 formats" and "25 formats" and "150 formats" answer three different questions here. Both guards are proven by mutation - a setting renamed in the registry, and a format registered under a name the list does not carry. That matters because the neighbouring TestTheFormatDocumentAgreesWithTheRegistry sits on notProvenByMutation, and that list is only allowed to shrink. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci: every whole tree test run states its timeout, and a guard keeps it that way Follow through on the commit before this one, which fixed the coverage gate and left three jobs sitting in exactly the same place. Measured on the green run of 2026-09-03, rather than assumed from the one job that went red: the test step takes 399 s on ubuntu, 476 s on windows and 491 s on macOS, all against Go's unstated ten minutes a package. macOS had 109 s of room. The same fleet was measured swinging from 479 s to past 600 s between two runs of one branch, so the margin was smaller than the variance on every one of them. The release workflow had the same gap, and there a timeout would read as a red tree and stop a release that was fine. Six whole tree runs now state a timeout. The two that already did are unchanged. The guard is the point of this commit rather than the flags. This is the SECOND time the project has lost a run to Go's default - the race detector met it on 2026-08-25 and the answer was a long comment beside that one job, which is why the coverage gate met it again eight days later. A diagnosis recorded at one step does not protect the next one, so the reasoning has moved out of the comments and into TestEveryWholeTreeTestRunStatesItsOwnTimeout. Proven by mutation. It reads the workflows through the YAML parser rather than as text, because a run block can be folded and the flag would then sit on a different line from the command. Only whole tree runs are asked: a targeted -run walks a handful of tests, and the fuzz step carries -fuzztime, which is its own budget. One stale claim fixed on the way. The matrix job's own comment said "the matrix runs in about a minute", which had not been true for a long time - it is eight, and the numbers are written down now instead of a word. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
1 parent e8cd383 commit a190b7b

16 files changed

Lines changed: 993 additions & 195 deletions

File tree

‎.github/workflows/ci.yml‎

Lines changed: 41 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -32,9 +32,12 @@ jobs:
3232
test:
3333
name: test on ${{ matrix.os }}
3434
# A hung job otherwise holds a runner until the GitHub default of six
35-
# hours. Every number here is well above what the job takes today: the
36-
# matrix runs in about a minute, the race detector took 148 s when it was
37-
# measured, and fuzzing is given 5 minutes a target by its own loop.
35+
# hours. Remeasured 2026-09-03, because the sentence here said "the matrix
36+
# runs in about a minute" and had not been true for a long time: the test
37+
# step alone takes 399 s on ubuntu, 476 s on windows and 491 s on macOS.
38+
# The race detector took 148 s when it was measured and has its own job
39+
# and its own numbers now, and fuzzing is given 5 minutes a target by its
40+
# own loop.
3841
timeout-minutes: 20
3942
strategy:
4043
fail-fast: false
@@ -244,7 +247,19 @@ jobs:
244247
shell: bash
245248

246249
- name: test
247-
run: go test -tags "$(cat .github/build-tags)" ./... -count=1
250+
# The timeout is stated for the reason the race job and the coverage
251+
# gate both state theirs: Go allows ten minutes PER PACKAGE by default,
252+
# this job allows twenty for all of it, and internal/guard is one
253+
# package holding almost every test there is. A run that went past the
254+
# first without approaching the second would die as a stack trace out
255+
# of whichever test happened to be running, which is what the coverage
256+
# gate did on 2026-09-03.
257+
#
258+
# Measured that day, on the run that caught it: 399 s on ubuntu, 476 s
259+
# on windows, 491 s on macOS. macOS therefore had 109 s of room under a
260+
# limit nobody had chosen, and the same fleet was measured swinging by
261+
# more than 25 percent between two runs of one branch.
262+
run: go test -tags "$(cat .github/build-tags)" ./... -count=1 -timeout 18m
248263

249264
- name: build the command line binary
250265
run: go build -tags "$(cat .github/build-tags)" ./cmd/tfg
@@ -727,10 +742,31 @@ jobs:
727742
# package, and by default Go credits coverage only to the package
728743
# under test - which reports 0.0% and makes the gate meaningless.
729744
# Measured, not assumed.
745+
#
746+
# The timeout is stated rather than left to Go, since 2026-09-03, and
747+
# for the same reason the race job above states its own. Go allows ten
748+
# minutes PER PACKAGE by default while this job allows twenty for all of
749+
# it, so internal/guard died on a limit nobody had chosen - a stack
750+
# trace out of whichever test was running when the alarm went off,
751+
# instead of a failure naming something.
752+
#
753+
# Measured on the runner rather than guessed. This step took 375 s and
754+
# 457 s on two consecutive main runs of 2026-09-02 and 2026-09-03, which
755+
# is 22 percent of variance on code that barely moved between them, and
756+
# the default cuts in at 600 s. A branch adding eight seconds of
757+
# coverage instrumented work then timed out. Eight seconds is not what
758+
# went wrong: 457 against 600 was never a margin, and a limit that
759+
# decides on how busy the runner is tells you nothing about the code.
760+
#
761+
# Atomic counters are the cost. Every statement in every internal
762+
# package pays one, and this package renders twenty five screens and
763+
# generates files for twenty four formats. Eighteen minutes sits under
764+
# the job's own ceiling on purpose, so a genuinely stuck run still fails
765+
# as a test with output rather than as a killed job without any.
730766
run: >
731767
go test -tags "$(cat .github/build-tags)" ./... -count=1 -covermode=atomic
732768
-coverpkg=./internal/...,./cmd/...
733-
-coverprofile=coverage.out
769+
-coverprofile=coverage.out -timeout 18m
734770
735771
- name: gate
736772
# The threshold lives in exactly one place, .github/coverage-threshold.

‎.github/workflows/release.yml‎

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -51,7 +51,13 @@ jobs:
5151
# A release built from a red tree is the one kind of release that
5252
# cannot be taken back, because the binaries are already on somebody's
5353
# disk. This is the same command CI runs.
54-
run: go test -tags "$(cat .github/build-tags)" ./... -count=1
54+
#
55+
# Including the timeout, and that is the point of saying so. Go allows
56+
# ten minutes PER PACKAGE by default while this job allows thirty for
57+
# all of it, and the suite measured 399 s to 491 s across the three
58+
# systems on 2026-09-03. A release run dying on the default would look
59+
# like a red tree and stop a release that was fine.
60+
run: go test -tags "$(cat .github/build-tags)" ./... -count=1 -timeout 18m
5561

5662
- name: what version the code says
5763
id: version

‎CHANGELOG.md‎

Lines changed: 30 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -16,6 +16,36 @@ because it turns other people's test suites red.
1616

1717
### Breaking
1818

19+
- **A generated `.csv` quotes only the fields that need it, so its bytes are
20+
different.** Sizes are unchanged. Every size that worked before still works,
21+
every reader that took these files still takes them, and the file is still
22+
RFC 4180.
23+
24+
The description column used to be quoted on every row. It is quoted now only
25+
when it carries the separator, which is the one thing that makes a quote
26+
necessary. The description is three to seven words and drops a separator
27+
every third one, so a short one carries none - measured on a 4 kB table,
28+
**9 of its 44 rows** lost their quotes.
29+
30+
This arrives as a new setting, `quote_style`, which takes `minimal`, `all` or
31+
`none`:
32+
33+
- `minimal` is the new default and is what a spreadsheet writes.
34+
- `all` wraps every field on every row, the header included.
35+
- `none` wraps nothing. It also stops the description carrying the separator,
36+
because an unquoted field cannot hold one without ending early - so this
37+
value changes what the file says and not only how it is punctuated.
38+
39+
**The smallest `.csv` is 115 B rather than 117**, because the shortest row
40+
has an empty description and an empty field needs no quotes. With
41+
`quote_style=all` the smallest is 139 B. `tfg formats` prints the current
42+
numbers.
43+
44+
**A suite pinning `.csv` hashes will go red once and then stay green.** There
45+
is no switch back to the old bytes: they were not any of the three styles RFC
46+
4180 describes, and carrying a fourth name for them forever costs more than
47+
the one red run.
48+
1949
- **Six formats have different bytes, because the tool is built with Go 1.27
2050
now.** Sizes are unchanged. Every size that worked before still works, every
2151
reader that took these files still takes them, and the same sizes are

‎README.md‎

Lines changed: 10 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -76,13 +76,13 @@ reference is below it.
7676

7777
## 📁 Formats it generates
7878

79-
Twenty two, and every one is a **real file of that format** - it opens in the
79+
Twenty four, and every one is a **real file of that format** - it opens in the
8080
software that owns it, at the exact size you asked for:
8181

8282
| group | formats |
8383
|---|---|
8484
| 📄 **Documents** | `pdf`, `docx` (Word), `xlsx` (Excel), `pptx` (PowerPoint) |
85-
| 🖼️ **Images** | `png`, `jpg`, `bmp`, `gif`, `ico`, `svg`, `tiff`, `webp`, `avif` |
85+
| 🖼️ **Images** | `png`, `jpg`, `bmp`, `gif`, `ico`, `svg`, `tiff`, `webp`, `avif`, `jxl` |
8686
| 📝 **Text and markup** | `txt`, `md`, `csv`, `json`, `xml`, `html`, `log` |
8787
| 🗜️ **Archives** | `zip`, `targz` (`.tar.gz`) |
8888
| 🔊 **Audio** | `wav` |
@@ -440,7 +440,7 @@ ignored quietly: `extends`, `with`, `policy`, `engine`, `defaults.fill`,
440440

441441
## 📁 Formats in detail
442442

443-
The twenty two formats are listed near the top of this file. Each is produced at an
443+
The twenty four formats are listed near the top of this file. Each is produced at an
444444
exact size and checked against independent readers before it ships - a PNG is
445445
opened and its pixels compared, a DOCX is read back by three separate
446446
libraries, an archive is extracted.
@@ -460,14 +460,17 @@ recipe. `tfg formats <id>` prints the allowed range or list for each:
460460
| `pdf` | `pages`, `page_size` |
461461
| `png`, `bmp`, `tiff`, `webp` | `width`, `height` |
462462
| `gif` | `width`, `height`, `frames` |
463-
| `jpg` | `width`, `height`, `quality` |
463+
| `avif`, `jpg`, `jxl` | `width`, `height`, `quality` |
464464
| `ico` | `width`, `height`, `embed` |
465465
| `wav` | `sample_rate`, `bit_depth`, `channels`, `content` |
466-
| `zip`, `targz` | `entries`, `entry_format`, `entry_size` |
466+
| `zip` | `entries`, `entry_format`, `entry_size`, `compression`, `depth`, `directory_entries`, `password`, `encryption` |
467+
| `targz` | `entries`, `entry_format`, `entry_size`, `compression`, `depth`, `directory_entries`, `entry_mode`, `entry_owner` |
467468
| `docx` | `paragraphs` |
468469
| `xlsx` | `rows`, `columns` |
469470
| `pptx` | `slides` |
470-
| `csv`, `json`, `xml`, `html`, `md`, `log`, `txt`, `svg` | none |
471+
| `csv` | `delimiter`, `line_ending`, `header`, `quote_style` |
472+
| `log` | `entry_format`, `timestamps`, `rate`, `methods`, `status_mix`, `level_mix`, `ip_version`, `line_ending` |
473+
| `json`, `xml`, `html`, `md`, `txt`, `svg` | none |
471474
472475
```
473476
tfg generate --format jpg --size 500kb --set width=1920 --set height=1080 --set quality=85
@@ -684,7 +687,7 @@ a valid one of its format.
684687

685688
Honest scope, because a tool that oversells itself wastes your afternoon.
686689

687-
**Working end to end:** twenty two formats, recipes, presets, the desktop window,
690+
**Working end to end:** twenty four formats, recipes, presets, the desktop window,
688691
`generate`, `validate`, `verify`, `cleanup`, boundary sets, archive contents,
689692
size ranges, per format settings, manifests and every exit code above.
690693

‎internal/format/csvfile/csv.go‎

Lines changed: 108 additions & 24 deletions
Original file line numberDiff line numberDiff line change
@@ -42,9 +42,10 @@ const (
4242
// drawn from below guarantees it.
4343
amountWidth = 9
4444

45-
// closingQuote ends the description. What follows it is the row ending,
46-
// which the dialect decides, so the two are no longer one constant.
47-
closingQuote = `"`
45+
// quoteMark wraps a field. It is the one character RFC 4180 gives a value
46+
// for holding a separator inside it, and how many of them a row carries is
47+
// what quote_style decides.
48+
quoteMark = '"'
4849

4950
// maxRowDigits bounds the width of the row number. A row is at least one
5051
// byte, so a file can never hold more rows than it has bytes, and a size is
@@ -54,21 +55,23 @@ const (
5455
maxRowDigits = 19
5556

5657
// fixedBeforeEnding is every byte of a row except the row number, the name
57-
// (which also forms the address), the description and the row ending. A
58-
// constant expression, so it cannot drift away from the template above.
58+
// (which also forms the address), the description, the quotes and the row
59+
// ending. A constant expression, so it cannot drift away from the template
60+
// above.
5961
//
6062
// The five separators count one byte each, which is a fact about the
6163
// separators offered rather than an assumption: every one of them is a
6264
// single byte, and dialect.go says so where they are declared.
6365
fixedBeforeEnding = 5 /* separators */ + len(emailDomain) + amountWidth +
64-
len(createdDate) + 1 /* the opening quote */ + len(closingQuote)
66+
len(createdDate)
6567
)
6668

67-
// fixedWidth is fixedBeforeEnding plus the row ending, which the dialect
68-
// decides. A CRLF row costs one byte more than an LF one, on every row, which
69-
// is why the minimum moves with this setting.
69+
// fixedWidth is fixedBeforeEnding plus the quotes and the row ending, both of
70+
// which the dialect decides. A CRLF row costs one byte more than an LF one on
71+
// every row, and quoting every field costs two bytes per column, which is why
72+
// the minimum moves with either setting.
7073
func fixedWidth(d dialect) int64 {
71-
return int64(fixedBeforeEnding + len(d.eol))
74+
return int64(fixedBeforeEnding + d.quotes.quoteBytes() + len(d.eol))
7275
}
7376

7477
func init() {
@@ -100,8 +103,8 @@ func init() {
100103
// name and the manifest carry it instead.
101104
Label: format.LabelExternalOnly,
102105
Oracle: "python-csv",
103-
// Quoting, column count and column types come later. Declaring none of
104-
// them now makes a recipe asking for one fail loudly.
106+
// Column count and column types come later. Declaring neither of them
107+
// now makes a recipe asking for one fail loudly.
105108
Properties: properties(),
106109
GeneratorVersion: generatorVersion,
107110
Generator: generator{},
@@ -146,6 +149,7 @@ func (generator) Plan(r format.Request) (format.Plan, error) {
146149
"line_ending": d.lineEndingID,
147150
"separator": string(d.sep),
148151
"header": d.header,
152+
"quote_style": d.quotes.id,
149153
"columns": len(columnNames),
150154
// Stated even though it is always false here, so a test can assert
151155
// on it without knowing which formats carry a label internally.
@@ -186,6 +190,15 @@ func (generator) Write(ctx context.Context, w io.Writer, p format.Plan) error {
186190
type rows struct {
187191
next int64
188192
dia dialect
193+
194+
// scratch holds the description while it is being asked whether it needs
195+
// quotes. It cannot be written straight into the row, because the answer
196+
// decides whether a quote goes in FRONT of it.
197+
//
198+
// Reused rather than allocated per row. A table of any size is millions of
199+
// rows and the resource guard measures exactly that. It stays small: the
200+
// closing row is the longest and is bounded by twice the shortest row.
201+
scratch []byte
189202
}
190203

191204
// Shortest is the smallest row this builder can close a file with: the widest
@@ -229,33 +242,100 @@ func (r *rows) append(dst []byte, rng *rand.Rand, want int64) []byte {
229242
cents := rng.IntN(100)
230243

231244
sep := r.dia.sep
245+
q := r.dia.quotes
232246

247+
dst = q.mark(dst)
233248
dst = strconv.AppendInt(dst, r.next, 10)
249+
dst = q.mark(dst)
234250
dst = append(dst, sep)
251+
dst = q.mark(dst)
235252
dst = append(dst, name...)
253+
dst = q.mark(dst)
236254
dst = append(dst, sep)
255+
dst = q.mark(dst)
237256
dst = append(dst, name...)
238257
dst = append(dst, emailDomain...)
258+
dst = q.mark(dst)
239259
dst = append(dst, sep)
260+
dst = q.mark(dst)
240261
dst = strconv.AppendInt(dst, int64(whole), 10)
241262
dst = append(dst, '.')
242263
if cents < 10 {
243264
dst = append(dst, '0')
244265
}
245266
dst = strconv.AppendInt(dst, int64(cents), 10)
267+
dst = q.mark(dst)
246268
dst = append(dst, sep)
269+
dst = q.mark(dst)
247270
dst = append(dst, createdDate...)
248-
dst = append(dst, sep, '"')
271+
dst = q.mark(dst)
272+
dst = append(dst, sep)
273+
274+
return r.appendDescription(dst, rng, want, int64(len(dst)-start))
275+
}
249276

277+
// appendDescription writes the last field and ends the row.
278+
//
279+
// want below zero means a natural row, any other value is the exact length the
280+
// whole row must have. used is what the row has spent already, measured rather
281+
// than worked out beside the bytes.
282+
func (r *rows) appendDescription(dst []byte, rng *rand.Rand, want, used int64) []byte {
250283
if want < 0 {
251-
dst = appendPhrase(dst, rng, 3+rng.IntN(5), sep)
252-
} else {
253-
// Everything written so far, plus what still has to follow.
254-
used := int64(len(dst)-start) + int64(len(closingQuote)) + int64(len(r.dia.eol))
255-
dst = appendFiller(dst, want-used, sep)
284+
r.scratch = appendPhrase(r.scratch[:0], rng, 3+rng.IntN(5),
285+
r.dia.sep, r.dia.quotes.separatorsInDescription)
286+
return r.closeRow(dst, r.scratch)
287+
}
288+
289+
// What is left for the description AND its quotes together. Which of the
290+
// two it is comes out of fill below.
291+
r.scratch = r.fill(r.scratch[:0], want-used-int64(len(r.dia.eol)))
292+
return r.closeRow(dst, r.scratch)
293+
}
294+
295+
// fill builds the description of the closing row so the row lands on exactly
296+
// the length it was asked for.
297+
//
298+
// room is the description and its quotes together, and how it divides between
299+
// them is the whole of this function. With "all" the quotes are certain. With
300+
// "none" there are none. With "minimal" it depends on the description itself,
301+
// so the quoted length is built first and MEASURED: if it carries the
302+
// separator the quotes are earned and that is the answer, and if it does not,
303+
// the description is built to the full room with the separator withheld, which
304+
// leaves nothing for a quote to be needed for.
305+
//
306+
// Measured on 2026-09-03: the filler first carries a separator at 30 B of
307+
// description, so the second branch is reached by the two lengths either side
308+
// of that. It is a narrow band and it is the only place the two halves of this
309+
// setting could have disagreed.
310+
func (r *rows) fill(dst []byte, room int64) []byte {
311+
q, sep := r.dia.quotes, r.dia.sep
312+
313+
if q.everyField {
314+
return appendFiller(dst, room-2, sep, true)
315+
}
316+
if q.separatorsInDescription && room >= 2 {
317+
dst = appendFiller(dst, room-2, sep, true)
318+
if q.wraps(dst, sep) {
319+
return dst
320+
}
321+
dst = dst[:0]
256322
}
323+
return appendFiller(dst, room, sep, false)
324+
}
257325

258-
dst = append(dst, closingQuote...)
326+
// closeRow puts the description into the row and ends it.
327+
//
328+
// It asks the setting whether these bytes carry quotes, and fill above asked
329+
// the same question to work the length out - one question, one answer, so the
330+
// arithmetic and the bytes cannot part company.
331+
func (r *rows) closeRow(dst, description []byte) []byte {
332+
if r.dia.quotes.wraps(description, r.dia.sep) {
333+
dst = append(dst, quoteMark)
334+
dst = append(dst, description...)
335+
dst = append(dst, quoteMark)
336+
} else {
337+
dst = append(dst, description...)
338+
}
259339
return append(dst, r.dia.eol...)
260340
}
261341

@@ -268,10 +348,13 @@ func (r *rows) append(dst []byte, rng *rand.Rand, want int64) []byte {
268348
// needs no quoting, so a description that kept dropping commas would leave a
269349
// semicolon file never exercising the quoted path at all - the file would be
270350
// the right size, parse everywhere, and quietly test less than the comma one.
271-
func appendPhrase(dst []byte, rng *rand.Rand, n int, sep byte) []byte {
351+
//
352+
// carries is false only under quote_style none, where an unquoted field cannot
353+
// hold a separator without ending early.
354+
func appendPhrase(dst []byte, rng *rand.Rand, n int, sep byte, carries bool) []byte {
272355
for i := 0; i < n; i++ {
273356
if i > 0 {
274-
if i%3 == 0 {
357+
if carries && i%3 == 0 {
275358
dst = append(dst, sep)
276359
}
277360
dst = append(dst, ' ')
@@ -291,11 +374,12 @@ func appendPhrase(dst []byte, rng *rand.Rand, n int, sep byte) []byte {
291374
// A separator every fourth word, unlike every other format here, and on
292375
// purpose: the description is a quoted field, so the padding is what makes a
293376
// long file keep exercising the quoting rather than turning into plain words.
294-
// It follows the dialect for the reason appendPhrase gives.
295-
func appendFiller(dst []byte, n int64, sep byte) []byte {
377+
// It follows the dialect for the reason appendPhrase gives, and it withholds
378+
// the separator for the reason appendPhrase gives too.
379+
func appendFiller(dst []byte, n int64, sep byte, carries bool) []byte {
296380
both := string(sep) + " "
297381
return core.AppendFiller(dst, words, n, func(i int) string {
298-
if i%4 == 0 {
382+
if carries && i%4 == 0 {
299383
return both
300384
}
301385
return " "

0 commit comments

Comments
 (0)