Audit EPUB body text and repair broken EPUBs with safe, deterministic fixes, validated by epubcheck, with optional in-place replacement in a Calibre library.
Bindery makes accidentally broken markup well-formed again. It does not rewrite or reflow content; it only applies a small set of deterministic, semantics-preserving fixes that real-world EPUBs (especially Calibre conversions) trip over:
- Unclosed void elements (
<link>,<br>,<img>, ...) get self-closed. - Undeclared named entities (
,°,é, ...) become numeric character references that every XML parser understands. - Bare
&(common intoc.ncx) is escaped to&. - Junk before the XML prolog (BOM, stray bytes) is stripped.
- Duplicate
xmlnson the root<html>is collapsed to one. - NCX-001:
toc.ncxdtb:uidis synced to the OPF unique identifier. - mimetype is rewritten first and stored, fixing the common ordering defect; a missing entry is added and wrong or whitespace-padded content is normalized to the OCF constant.
Five opt-in fixes go further:
--fix-ids: rewrite ids that are not valid XML names (start with a digit, contain a colon) in the OPF manifest, updating every reference to them (spine, fallback, media-overlay, the EPUB 2 cover meta), and in the NCX (where old conversions stamp navPoint ids from UUIDs, one error per id). Touches the OPF, so it is off by default; the dc: metadata is never altered.--add-img-alt: addalt=""to<img>elements missing the required attribute. Renders identically, but it is the one fix that adds markup the author never wrote, and an empty alt tells a screen reader the image is decorative; hence opt-in.--reserialize: rebuild content documents that are still malformed by re-parsing them with html5lib and re-emitting XHTML, closing unclosed<p>/<div>/<span>/<blockquote>that the regex transforms cannot. Runs only on documents that are not already well-formed, so good files are untouched.--strip-bad-attrs: drop attributes that are invalid XML (a name starting with a digit, or a namespaced name whose prefix is never declared, like Office VMLv:shapes). Surgical and a no-op on well-formed files.--escape-unknown-entities: escape entity names that are not in the HTML5 table (&foo;becomes&foo;), which renders exactly as browsers already render an unknown entity: the literal text. Documents whose DOCTYPE carries an internal subset are skipped wholesale, since a subset can declare custom entities.--fix-ncx-playorder: safely rewrites duplicatedplayOrderattributes sequentially in the NCX table of contents.--fix-id-colons: translates illegal colons inid="X:Y"and their matching#X:Yfragment references to valid underscores.--fix-empty-body: appends a non-breaking space to strictly empty body tags to satisfy parser requirements.--fix-missing-title: injects a<title>Unknown</title>fallback in the<head>if missing, handling both empty<title/>self-closing tags and entirely absent tags.--unwrap-block-in-inline: safely unwraps<span>tags that illegally contain a block-level element (e.g.<div>or<p>), leaving the block element intact.--strip-invalid-value: systematically strips invalidvalue="..."attributes from elements like<div>,<span>,<p>, etc.--unwrap-illegal-tags: strips completely invalid or deprecated HTML tags that break EPUB3 validation (like<st>,<font>,<sentence>,<o>,<w>, and<pagebreak>) while retaining their inner text. Automatically excludes tags if they are referenced by an EPUB CSS stylesheet to guarantee 100% format preservation.
Three opt-in fixes are lossy and stand apart from the semantics-preserving rest:
--strip-pagination: remove print page numbers and running headers that a PDF/OCR conversion baked into the body text as literal paragraphs (so they reflow into the middle of a sentence: "where the hay cart 16 was taking him"). It removes only that injected furniture, never the author's prose: where a number split a sentence it rejoins the two paragraphs (closing up a word split likecompli-/mentary), and it preserves roman chapter numbers, page-list nav anchors, and years. A book is only treated as paginated when it has both a dense run of arabic numbers and several confident mid-sentence interrupts, so a merely chapter-numbered book is left alone. Three safety nets guard every edit (character conservation, tag balance, and an epubcheck no-regression check); any failure leaves the document untouched.--strip-broken-tags: remove leaked HTML closing tags missing their open brackets (e.g.</p>) that render as raw text.--strip-watermarks: remove known producer and distributor watermarks (e.g. OceanofPDF, ABC Amber LIT Converter) and stray marker files. It locates the stamp and deletes the outermost wrapper whose entire visible text is the watermark, ensuring prose that merely mentions the URL is preserved.
All lossy edits are invisible to epubcheck, so they are accepted when the result is no worse rather than measurably better.
Every repair is gated by epubcheck. The acceptance rule is two-mode, because a fatal parse error makes epubcheck stop reading a file and hides every downstream error:
- If a book had fatals, success means fewer fatals. The error count may rise as previously-hidden schema warnings become visible once the file parses; that is the book going from "won't open" to "opens with nits," not a regression.
- If a book had no fatals, an error increase is a real regression, so a strict error decrease is required.
Introducing a net-new fatal is always rejected. If epubcheck itself fails to run (crash, timeout, unparsable output), the book is reported as an error and never applied; only an explicit --no-validate skips the gate. Originals are never modified except by an explicit, atomic in-place replace (see below), and even then only after the gate accepts the result.
The lossy modes (--strip-pagination, --strip-broken-tags, and --strip-watermarks) are the exception to the "must improve" rule. Since they remove visible markup rather than correcting XML schema violations, epubcheck counts often remain unchanged. They are accepted when the result is no worse (no net-new fatals or errors), relying on strict programmatic safety nets instead.
Python 3.14+, plus epubcheck on PATH for the gate. The core is minimal-dependency (tqdm is used for output formatting);
html5lib is an optional extra, needed only for --reserialize.
uv tool install /path/to/Bindery # core tools
uv tool install "bindery[reserialize] @ /path/to/Bindery" # incl. --reserialize
# or from a checkout:
PYTHONPATH=src python3 -m bindery --help # all modes except --reserialize
PYTHONPATH=src uv run --with html5lib python3 -m bindery --help # incl. --reserializeBindery includes a comprehensive auditing tool to inspect EPUB body text for non-schema flaws that epubcheck cannot catch. It extracts and analyzes the visible text to detect content issues, producing console reports that can be used to filter your library or feed into bindery repair.
Scan a Calibre library for specific issues:
# Detect non-English content (e.g., Cyrillic or CJK in an English library)
cd ~/docs/Calibre\ Library && bindery audit content
# Find books with hardcoded print page numbers interrupting the text
bindery audit pagenumbers ~/docs/Calibre\ Library
# Find books that are empty or severely truncated
bindery audit emptytext ~/docs/Calibre\ Library
# Find books with severe OCR damage (garbage characters, excessive hyphenation)
bindery audit ocr ~/docs/Calibre\ Library
# Run all audits and generate a comprehensive CSV
cd ~/docs/Calibre\ Library && bindery audit allAudits can also be run on a directory of loose .epub files by passing the path as the second argument:
bindery audit pagenumbers /path/to/loose/epubsRepair one book to a new file (gated; writes only if it is an improvement):
bindery repair broken.epub # -> "broken (repaired).epub"
bindery repair broken.epub fixed.epub
bindery repair scanned.epub --strip-pagination # also remove baked-in page numbersScan a Calibre library and see what would be fixed, writing nothing:
bindery library ~/docs/Calibre\ Library --only fatals --audit epub_audit.csvApply accepted repairs in place, atomically, with backups:
bindery library ~/docs/Calibre\ Library --only fatals --apply --backup ~/bindery-backupsApply all safe and lossy repairs directly into the Calibre database natively:
bindery library ~/docs/Calibre\ Library --only all --apply --all --install-to-calibre--only {fatals,ncx,all}restricts the candidate set.ncxtargets NCX-001 mismatches (detected without epubcheck);fatalsneeds--audit.--audit CSV(thefatals,errors,warnings,pathformat produced by an epubcheck sweep) skips clean books so a run is fast. Paths are resolved on both sides, and a CSV that matches nothing triggers a loud warning instead of silently selecting zero books.--sweepreplaces the CSV step entirely: it runs a live epubcheck sweep for candidate selection and reuses each result as that book's before-measurement, so no book is checked twice. Bindery launches a transparent, persistent Java daemon (EpubcheckDaemon) to evaluate candidates in the background, eliminating the JVM startup penalty. This reduces validation time from ~5 seconds per book down to ~0.05 seconds, allowing--sweepto scan a 5,000-book library in just 10 minutes. Combine with--only fatalsfor a self-contained, lightning-fast "find and fix the broken books" run.--json FILEwrites a machine-readable report of the whole run (per-book status, before/after counts, applied flag, summary totals).--manual-list FILEwrites the paths of every book that was not auto-repaired, one per line, ready for manual follow-up.--applyis required to write; the default is a dry run.--backup DIRmirrors originals before replacing;--backup-inplacewrites.epub.bakbeside each file.--install-to-calibreacts as a safer alternative to atomic replacement for Calibre libraries. It usescalibredb add_format --replaceto natively swap the EPUB format without losing metadata or read progress. It falls back to atomic file replacement if the Calibre database ID cannot be parsed.--allautomatically turns on all opt-in non-fatal fixes and lossy strips (pagination, watermarks, bad attributes, unknown entities, image alt tags, etc.) in a single run.- Only the
.epubis replaced.metadata.opf,cover.jpg, andmetadata.dbare left for Calibre's Quality Check sync to reconcile. - A per-book progress line goes to stderr (stdout stays a clean report);
--quietsuppresses it. A corrupt or unreadable book is reported and skipped, never aborting the sweep. - Exit codes: 0 for a clean sweep, 1 for a usage error, 2 when any book was rejected, unreadable, or failed epubcheck (for scripts and cron).
repairrefuses to overwrite an existing output file unless--forceis given.
scripts/ holds standalone, read-only utilities that are useful for EPUB maintenance but fall outside Bindery's repair contract (fixing what they find would be a content change, which Bindery makes only via the opt-in --strip-pagination):
find_missing_images.py: scans a library tree and reports every book whose<img>tags point at files that do not exist inside the archive (a common defect in converted EPUBs). Reads the archives in place; nothing is unpacked or written. The library path is set at the bottom of the script.
./run_tests.sh # unittest suiteSee spec.md for the full contract and roadmap.md for what is planned.
MIT. See LICENSE.
If Bindery's useful to you and you'd like to chip in:
- liberapay · liberapay.com/bdkl
- bitcoin
bc1qkge6zr45tzqfwfmvma2ylumt6mg7wlwmhr05yv
