Identify and extract LittleFS filesystems (#16) - #34
Merged
Merged
Conversation
LittleFS is the default on-flash filesystem across modern MCU stacks (Zephyr,
Mbed, ESP-IDF, nRF, STM32) and neither binwalk nor unblob extract it. Add full
identification and byte-exact extraction.
- Identify: anchored on the "littlefs" magic at superblock offset 8. Parses the
superblock metadata pair, verifies the commit CRC32, reports version + geometry
(block_size / block_count), and sizes the finding to the whole filesystem.
- Extract: walk the metadata directories (the commit log of XOR-chained tags per
metadata pair, applying CREATE/DELETE id renumbering and following hardtail
split-continuations), reconstruct files from inline data and CTZ skip-lists,
and write the tree to the SafeRoot. Bounded by depth / file / dir / byte caps
and the block count; no process spawn.
The commit and CTZ skip-list algorithms are a clean reimplementation of the
littlefs on-disk format (BSD-3-Clause), written against moria's Reader and
verified byte-exact against real mklittlefs images (noted in README).
- New: signatures/littlefs.toml, src/littlefs_parse.{hpp,cpp} (shared core),
src/validators/littlefs.{hpp,cpp}, src/extract/littlefs.{hpp,cpp}. Wired into
the validator registry, extractor registry, MIME map and build.
- Tests: a minimal-superblock fixture in gen_samples (identifies verified), and
tests/test_littlefs.py (synthetic identify, corrupt-CRC / bare-string
false-positive guards, and a real mklittlefs round-trip that extracts every
file byte-exact for two block sizes, self-skipping when the tool is absent).
Wired into run.sh and CTest.
# Conflicts: # CMakeLists.txt # tests/run.sh
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
LittleFS is the default on-flash filesystem across modern MCU stacks (Zephyr, Mbed OS, ESP-IDF, Nordic nRF, STM32), and neither binwalk nor unblob extract it — config, provisioning data and keys frequently live in a LittleFS partition. This adds full identification and byte-exact extraction.
What it does
littlefsmagic at superblock offset 8. Parses the superblock metadata pair, verifies the commit CRC32, reports the version and geometry (block size / block count), and sizes the finding to the whole filesystem so interior metadata and data aren't re-detected as separate findings.SafeRoot. Bounded by depth / file / dir / byte caps and the filesystem's own block count; no process spawn.Geometry is self-contained in the superblock, so no external parameters are needed (unlike SPIFFS/yaffs2). An embedded LittleFS at a nonzero offset is handled.
Provenance
The commit-log and CTZ skip-list algorithms are a clean reimplementation of the LittleFS on-disk format (BSD-3-Clause), written against moria's
Readerand verified byte-exact against realmklittlefsimages. Noted in the README's license section.Tests
gen_samples(identifies at the verified tier; also cross-validates the parser, since the Python encoder and the C++ decoder were written independently from the spec and agree).tests/test_littlefs.py: synthetic identify, false-positive guards (a corrupt-CRC superblock and a barelittlefsstring are both rejected), and a realmklittlefsround-trip that extracts every file byte-exact across two block sizes (4096 and 512), self-skipping when the tool is absent.Closes #16