Skip to content

Commit f78a5c0

Browse files
donislawdevclaude
andcommitted
format: archives can compress what they hold
One axis for both containers - compression: none, fast, default, best, meaning levels 0, 1, 6 and 9. The words are the intent rather than the mechanism, because ZIP sets a method per entry and TAR.GZ compresses the whole stream, and those are the same choice said two ways. The archive still comes out exactly the size that was ordered. What changes is how much of it is content and how much is padding. The hard part is that a compressed length cannot be planned. Measured: our content compresses at about 50 MB/s, and the guard that keeps a preview cheap plans three gigabytes of declared contents in milliseconds against a twenty second ceiling, so compressing to learn the length would take about a minute. This project had already measured the same thing from the other side on TIFF, where deflate moves the length with the seed while uncompressed is flat. So the plan keeps its stored arithmetic and the writer settles the difference. ZIP needs one pass. Every structural field is fixed width, so a compressed archive differs from a stored one by the entry data and nothing else - the plan works the padding out for the stored archive and the writer gives back what the compressor freed, measured with a registered compressor rather than by counting headers. Two things had to be measured rather than reasoned about: an entry's compressed bytes do not reach the writer until the entry closes and Flush does not close one, and the filler has to stay stored because it is random, so deflate grows it from 65195 to 65220 bytes. TAR.GZ needs two, because gzip has no per-entry method. The filler carries the bulk and the gzip extra field closes the remainder exactly, since it sits in the uncompressed header and costs n+2. Measured across levels 1, 6 and 9 and targets from 64 KB to 10 MB: every one landed on the ordered size, with 567 to 3759 bytes left for the field against the 65531 it holds. Two combinations are refused, each naming both settings. Compression with a size from the contents cannot be planned. Compression with a password cannot be streamed, because a locked entry states its length before its data is written and a compressed one does not know it yet. The default is none, so no existing archive changes by a byte. Verified with independent readers rather than only Go: 7-Zip and GNU tar take all four levels in both containers, and Python sees deflate on the entries and store on the filler with the content coming back unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 parent a00ee9d commit f78a5c0

14 files changed

Lines changed: 941 additions & 72 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 25 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -70,6 +70,31 @@ because it turns other people's test suites red.
7070

7171
### Added
7272

73+
- **An archive can compress what it holds.** `--set compression=best` on a
74+
`zip` or a `targz`, with `none`, `fast`, `default` and `best` to choose from.
75+
76+
The archive still comes out **exactly the size you asked for**. What changes
77+
is how much of it is your files and how much is padding: at `best` a
78+
megabyte archive holding four 32 KB text files carries the same four files
79+
deflated, and the padding entry grows to make up the difference. A reader
80+
sees real deflated entries, which is what a tool under test has to cope with.
81+
82+
The default is `none`, which is what archives from this tool have always
83+
been, so **no existing file changes by a byte**.
84+
85+
Two combinations are refused rather than half-supported, and the message
86+
says which two settings to choose between. Compression with a size taken
87+
**from the contents**: the archive's length would then be whatever the
88+
contents compress to, which is only knowable by compressing them, and that
89+
would make a preview cost as much as the run. Compression with a
90+
**password**: a locked entry has to state its length before its data is
91+
written, so a compressed one would have to be held in memory whole.
92+
93+
Compressing costs time at write, not at preview. A 10 MB archive takes about
94+
25 ms at `fast` and 140 ms at `default`, against 8 ms stored, and a `.tar.gz`
95+
pays that twice because gzip compresses the whole stream and the size has to
96+
be measured before it can be hit.
97+
7398
- **An archive can hold its files in directories.** `--set depth=3` puts every
7499
file three levels down, and `--set directory_entries=true` also makes the
75100
archive list the directories themselves. Both work on `zip` and on `targz`.

‎internal/format/archive/archive.go‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -113,6 +113,13 @@ func mustSize(s string) int64 {
113113
// somebody kept them so by hand, and the comment saying so was the whole
114114
// mechanism.
115115
var axes = map[string]format.Property{
116+
Compression: {
117+
Name: Compression, Kind: format.PropertyChoice,
118+
Choices: []string{CompressBest, CompressDefault, CompressFast, CompressNone},
119+
Default: CompressNone,
120+
Detail: "How hard the files inside are squeezed. " +
121+
"The archive still comes out the size you asked for - what changes is how much of it is real content and how much is padding.",
122+
},
116123
Depth: {
117124
Name: Depth, Kind: format.PropertyInt,
118125
Min: 0, Max: maxDepth,
Lines changed: 109 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,109 @@
1+
package archive
2+
3+
import (
4+
"compress/flate"
5+
6+
"github.com/donislawdev/TestingFilesGenerator/internal/format"
7+
)
8+
9+
// How hard the archive is squeezed, in one vocabulary for both containers.
10+
//
11+
// The words are deliberately not the mechanism. ZIP picks a method per entry
12+
// and TAR.GZ compresses the whole stream, so "deflate level 6" and "gzip level
13+
// 6" are the same intent said two ways - and a person choosing a setting is
14+
// choosing how they want to trade time for size, not which library runs.
15+
//
16+
// none is the default and has to stay the default: every archive this tool has
17+
// written so far is stored, so any other default would move the bytes of all of
18+
// them, which is untouchable rule 3.
19+
const Compression = "compression"
20+
21+
const (
22+
CompressNone = "none"
23+
CompressFast = "fast"
24+
CompressDefault = "default"
25+
CompressBest = "best"
26+
)
27+
28+
// Levels are the flate and gzip levels the words mean. Measured 2026-09-01 on
29+
// a repeating text payload: level 1 came to 1304 B where levels 5, 6 and 9 all
30+
// came to 616 B, so fast really is a different answer rather than a label. On
31+
// a 10 MB archive one pass costs 25 ms at level 1 against 140 ms at level 6.
32+
var levels = map[string]int{
33+
CompressNone: flate.NoCompression,
34+
CompressFast: 1,
35+
CompressDefault: 6,
36+
CompressBest: 9,
37+
}
38+
39+
// Squeeze is what a container should do with the bytes.
40+
type Squeeze struct {
41+
// Name is the word the person asked for, for the manifest.
42+
Name string
43+
// Level is the flate or gzip level it means.
44+
Level int
45+
}
46+
47+
// On reports whether anything is actually compressed. The zero value is off,
48+
// which is what every archive written before this existed did.
49+
func (s Squeeze) On() bool { return s.Name != "" && s.Name != CompressNone }
50+
51+
// ReadCompression works out how hard to squeeze, and refuses the two
52+
// combinations that cannot mean what they say.
53+
//
54+
// The refusals are not tidiness, and each has a measurement behind it.
55+
//
56+
// Compression with a size that comes from the CONTENTS cannot be planned. The
57+
// archive's size would then be whatever the contents compress to, and that is
58+
// knowable only by compressing them - which is exactly what the guard on
59+
// planning forbids. Measured 2026-09-01: our content compresses at about
60+
// 50 MB/s, and that guard plans three gigabytes of declared contents in
61+
// milliseconds against a twenty second ceiling. Compressing to find the answer
62+
// would take about a minute.
63+
//
64+
// Compression with a PASSWORD cannot be streamed. A locked entry goes through
65+
// CreateRaw, which needs the compressed length in the header before any of the
66+
// data is written, so the entry would have to be deflated into memory first -
67+
// and a generator holding a whole entry in memory is the other rule this
68+
// project holds. Both halves are named in each message, because from "this is
69+
// not allowed" nobody can tell which of the two to change.
70+
func ReadCompression(id string, r format.Request, locked bool) (Squeeze, error) {
71+
raw, ok := r.Properties[Compression]
72+
if !ok || raw == "" {
73+
return Squeeze{Name: CompressNone, Level: flate.NoCompression}, nil
74+
}
75+
level, known := levels[raw]
76+
if !known {
77+
return Squeeze{}, &format.PropertyValueError{
78+
Format: id,
79+
Key: Compression,
80+
Value: raw,
81+
Reason: "it takes one of: " + CompressBest + ", " + CompressDefault + ", " + CompressFast + ", " + CompressNone,
82+
Remedy: "Ask for " + CompressNone + " to store the files as they are.",
83+
}
84+
}
85+
s := Squeeze{Name: raw, Level: level}
86+
if !s.On() {
87+
return s, nil
88+
}
89+
90+
if r.SizeFromContents {
91+
return Squeeze{}, &format.PropertyValueError{
92+
Format: id,
93+
Key: Compression,
94+
Value: raw,
95+
Reason: "the size is being left to the contents, and how far they compress is only known once they have been compressed",
96+
Remedy: "Give the archive an explicit size, or ask for " + Compression + ": " + CompressNone + ".",
97+
}
98+
}
99+
if locked {
100+
return Squeeze{}, &format.PropertyValueError{
101+
Format: id,
102+
Key: Compression,
103+
Value: raw,
104+
Reason: "the archive is locked with a " + Password + ", and a locked entry states its length before its data is written - so a compressed one would have to be held in memory whole",
105+
Remedy: "Ask for " + Compression + ": " + CompressNone + ", or take the " + Password + " off.",
106+
}
107+
}
108+
return s, nil
109+
}

‎internal/format/targz/compress.go‎

Lines changed: 187 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,187 @@
1+
package targz
2+
3+
import (
4+
"context"
5+
"fmt"
6+
"io"
7+
8+
"github.com/donislawdev/TestingFilesGenerator/internal/format"
9+
"github.com/donislawdev/TestingFilesGenerator/internal/format/archive"
10+
)
11+
12+
// Settling the padding of a COMPRESSED archive, which cannot be done by
13+
// arithmetic.
14+
//
15+
// A stored tar.gz has a length that follows from its parts, and size.go works
16+
// it out exactly: the tar is 1024 plus 512 and the rounded content per entry,
17+
// and gzip frames that predictably. Compression breaks every term of it. The
18+
// project measured the same thing on TIFF - deflate moves the length with the
19+
// seed, while uncompressed is flat - so the only honest way to learn a
20+
// compressed length is to compress and look.
21+
//
22+
// So the padding is settled here instead, at write time, and the shape is:
23+
//
24+
// the FILLER carries the bulk. It is random, so the compressor cannot shrink
25+
// it - measured at about +0.031% - but its compressed length is still not
26+
// exactly predictable, and a tar entry moves in 512 byte blocks anyway.
27+
// the EXTRA FIELD closes the remainder exactly. It sits in the gzip header,
28+
// which is not compressed, so n bytes there cost n+2 in the file. That is the
29+
// one channel here with byte granularity.
30+
//
31+
// Measured 2026-09-01 across levels 1, 6 and 9 and targets from 64 KB to
32+
// 10 MB: every one landed on the ordered size, and the remainder left for the
33+
// extra field came out between 567 and 3759 bytes - far inside the 65 531 the
34+
// field holds (O163).
35+
//
36+
// This costs passes, and the cost is real: one pass over a 10 MB archive is
37+
// 25 ms at level 1 and 140 ms at level 6, against 8 ms stored. It buys the one
38+
// thing that cannot be given up, which is that the file is the size that was
39+
// ordered.
40+
const solveRounds = 8
41+
42+
// counter counts what a write would come to without keeping any of it.
43+
type counter struct{ n int64 }
44+
45+
func (c *counter) Write(p []byte) (int, error) { c.n += int64(len(p)); return len(p), nil }
46+
47+
// measure builds the archive described by m and reports its length.
48+
//
49+
// It writes nothing anybody keeps, but it does GENERATE the files inside,
50+
// because that is the only way to learn what they compress to. That is the
51+
// price of compression in this container and it is paid at write time, never
52+
// at planning time - the guard that keeps a preview cheap is about planning.
53+
func measure(ctx context.Context, m memo) (int64, error) {
54+
c := &counter{}
55+
if err := build(ctx, c, m); err != nil {
56+
return 0, err
57+
}
58+
return c.n, nil
59+
}
60+
61+
// settleCompressed finds the filler and extra field that make the archive come
62+
// to exactly m.target.
63+
//
64+
// It walks rather than solving in one step because the two channels do not have
65+
// the same granularity: a tar entry moves in 512 byte blocks and the filler is
66+
// compressed on the way in, so asking for n more bytes of filler does not add
67+
// exactly n to the file. The extra field does add exactly what it is given, so
68+
// it always gets the last word.
69+
func settleCompressed(ctx context.Context, m memo) (memo, error) {
70+
bare := m
71+
bare.withFiller, bare.fillerSize = false, 0
72+
bare.withExtra, bare.extraLen = false, 0
73+
74+
base, err := measure(ctx, bare)
75+
if err != nil {
76+
return m, err
77+
}
78+
if base > m.target {
79+
return m, belowMinimum(m.target, base)
80+
}
81+
return settleRound(ctx, bare, m.target, m.target-base, solveRounds)
82+
}
83+
84+
// settleRound tries one filler and either lands or says what to try next.
85+
//
86+
// Written as a walk rather than a loop, and that is not decoration: a loop
87+
// carrying an error check and a decision inside it nests three deep, and this
88+
// project counts how many functions do. A bounded recursion says the same
89+
// thing at two, and a solve that converges reads naturally as "try this, and
90+
// if it is not right, try the next" anyway. left bounds it, so there is no
91+
// depth to worry about.
92+
func settleRound(ctx context.Context, bare memo, target, filler int64, left int) (memo, error) {
93+
if left == 0 {
94+
return bare, fmt.Errorf(
95+
"targz: the padding of this compressed archive does not settle after %d rounds. "+
96+
"Ask for a different size, or for compression: none", solveRounds)
97+
}
98+
if err := ctx.Err(); err != nil {
99+
return bare, err
100+
}
101+
102+
try := bare
103+
try.withFiller, try.fillerSize = filler > 0, filler
104+
got, err := measure(ctx, try)
105+
if err != nil {
106+
return bare, err
107+
}
108+
109+
next, extra, useExtra, done := nextFiller(target, got, filler)
110+
if done {
111+
try.withExtra, try.extraLen = useExtra, extra
112+
return try, nil
113+
}
114+
if next < 0 {
115+
return bare, belowMinimum(target, got)
116+
}
117+
return settleRound(ctx, bare, target, next, left-1)
118+
}
119+
120+
// nextFiller reads one measurement and says what to do with it.
121+
//
122+
// The four answers are the whole of the arithmetic, and which one applies is
123+
// decided by how much is left over rather than by preference.
124+
func nextFiller(target, got, filler int64) (next, extra int64, useExtra, done bool) {
125+
switch deficit := target - got; {
126+
case deficit < 0 || (deficit > 0 && deficit < 2):
127+
// Overshot, or left a remainder the extra field cannot hold: it costs
128+
// two bytes before it holds anything. Give the filler back enough that
129+
// the field has room to work.
130+
return filler - (2 - deficit), 0, false, false
131+
case deficit == 0:
132+
// Landed without needing the field at all.
133+
return filler, 0, false, true
134+
case deficit-2 > extraPaddingLimit:
135+
// More left than the header can hold, so the filler takes it.
136+
return filler + deficit - 2 - extraPaddingLimit, 0, false, false
137+
default:
138+
return filler, deficit - 2, true, true
139+
}
140+
}
141+
142+
// writeCompressed settles the padding and then writes the archive.
143+
//
144+
// Two passes at least, and the reason is in settleCompressed: the length of a
145+
// compressed archive is not knowable without compressing it. Nothing is held
146+
// in memory between them - the measuring pass throws its bytes away as it
147+
// makes them, so an archive larger than memory still works.
148+
func writeCompressed(ctx context.Context, w io.Writer, m memo) error {
149+
settled, err := settleCompressed(ctx, m)
150+
if err != nil {
151+
return err
152+
}
153+
return build(ctx, w, settled)
154+
}
155+
156+
// belowMinimum says the archive cannot be made this small once its contents
157+
// are in it.
158+
//
159+
// The number it reports is measured rather than derived: it is what this
160+
// archive actually came to when it was compressed with nothing added, which is
161+
// the smallest it can be. A stored archive can say the same thing by
162+
// arithmetic, and a compressed one cannot.
163+
func belowMinimum(target, floor int64) error {
164+
return &format.BelowMinimumError{
165+
Format: "TAR.GZ",
166+
Requested: target,
167+
Minimum: floor,
168+
Reason: "that is what the contents come to once they are compressed, so nothing can be taken away",
169+
Hint: fmt.Sprintf("Ask for %d B or more, or hold fewer or smaller files.", floor),
170+
}
171+
}
172+
173+
// reachable refuses a compressed archive whose size the contents already
174+
// exceed, using the STORED arithmetic.
175+
//
176+
// The stored number is the honest bound to check here even though the archive
177+
// will be compressed. Compression only ever makes the contents smaller, so an
178+
// archive that fits when stored fits when squeezed - and the stored number is
179+
// the one the plan can work out without compressing anything, which is what
180+
// keeps a preview cheap.
181+
func reachable(m *memo, target int64, label string, groups []format.Content) error {
182+
probe := *m
183+
probe.squeeze = archive.Squeeze{}
184+
var p format.Plan
185+
p.Properties = map[string]any{}
186+
return pad(&probe, &p, target, label, groups)
187+
}

‎internal/format/targz/size.go‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -276,7 +276,7 @@ func solveFiller(base, target int64, label string) (size int64, header headerPad
276276
// the stream measured before the comment can be sized, so it is two passes and
277277
// a separate decision.
278278
func build(ctx context.Context, w io.Writer, m memo) error {
279-
zw, err := gzip.NewWriterLevel(w, gzip.NoCompression)
279+
zw, err := gzip.NewWriterLevel(w, m.squeeze.Level)
280280
if err != nil {
281281
return fmt.Errorf("targz: the archive could not be started: %w", err)
282282
}

0 commit comments

Comments
 (0)