v0.4.0 : disk cache, bulk file resolutions, async behavior changes & experimental tui - #64
Open
sreedevk wants to merge 4 commits into
Open
v0.4.0 : disk cache, bulk file resolutions, async behavior changes & experimental tui#64sreedevk wants to merge 4 commits into
sreedevk wants to merge 4 commits into
Conversation
The deduplication core was a three-stage streaming producer/consumer (scan -> group-by-size -> group-by-hash) orchestrated by a Server struct that ran the stages concurrently on a threadpool. Coordination relied on AtomicBool flags polled in busy-wait loops, an Arc<Mutex<Vec>> hand-off queue, DashMap stores, and a per-file Arc<Mutex<FileState>>. Every stage had to receive its own hand-cloned Arcs with ad-hoc names, and the busy-wait branches burned a core spinning while waiting for the producer. Replace it with a staged batch pipeline over owned collections: pipeline::run(&Params) drives scan -> group_by_size -> group_by_hash in sequence, with rayon supplying parallelism. With no shared mutable state between stages, all the Arc plumbing, the AtomicBool coordination, the Mutex queue, the per-file lock, server.rs, and the threadpool/dashmap dependencies are gone. FileInfo becomes plain data and each stage is a pure, independently testable function. Behavior is preserved except for three authorized deviations: progress spinners render sequentially rather than concurrently under -p; the interactive empty-result message prints once instead of twice; and the interactive "Duplicate Set X of Y" total now counts the confirmed duplicate groups shown instead of the internal candidate-hash store size. This lands as the foundation for the cache, CLI, and TUI work that follows.
The tool could only act on duplicates through the default listing or the manual per-group interactive prompt, with no way to resolve groups in bulk by a rule. This blocked any scripted or automated use, and bulk resolution is the top item in the README's proposed operations. Add a resolver module that reduces each duplicate group to a single kept file. --keep <newest|oldest|first|last|shortest|shallowest> selects the keeper: newest/oldest by mtime, first/last and shortest/shallowest by path, each resolved to a unique keeper through total-order path tiebreaks so the outcome is deterministic and order-independent. Selection is a pure, independently tested function; deletion is a separate step. Deletion is guarded: --keep alone previews (KEEP/DELETE lines and a would-free summary) and removes nothing, so a mistaken invocation is harmless. --force performs the deletion, continues past individual failures, reports freed space, and exits non-zero if any file could not be removed. clap enforces that --keep excludes --interactive and that --force requires --keep, so misuse fails before any file is touched. Also fixes the interactive table's column width, which sized on path component count instead of path string length so the columns never aligned.
Content hashing is the pipeline's expensive stage; every run re-read and re-hashed all size-collision candidates from scratch. Persist those hashes so an unchanged file is skipped on subsequent runs. Add a path-keyed cache mapping absolute_path -> (size, mtime, strict, hash, last_seen), stored in a compact hand-rolled little-endian binary format with a magic+version header. A candidate reuses its stored hash only when size, mtime, and hash-mode all match; otherwise it is re-hashed and the entry is refreshed. On save, entries unseen for 30 days are pruned so the file cannot grow without bound. The hashing stage splits into hash_candidates (parallel, cache-consulting) and group_hashed (grouping) so lookups stay inside the parallel pass while cache mutation stays single-threaded after it. The cache is strictly a performance hint: a missing, corrupt, version- mismatched, or unwritable cache degrades to a correct full-hash run and never fails or changes the result. Caching is on by default at the platform cache directory; --no-cache disables it and --cache-file overrides the path. The gxhash seed becomes a fixed constant so stored hashes are reproducible across runs, which is what makes cache hits possible; rand is consequently no longer a production dependency (now dev-only), and dirs is added for the platform cache directory.
The tool could only act on duplicates non-interactively or through a line-oriented per-group prompt. Add a full-screen terminal UI (--tui) for browsing and resolving them visually. The UI is a two-pane master/detail view: a list of duplicate groups and, for the selected group, its files with per-file deletion marks. Files can be marked individually or in bulk by a keep-strategy (newest/oldest/first/ last/shortest/shallowest, reused from the resolver) applied to the current group or to all groups at once. A footer tracks how many files are marked and how much space would be reclaimed. Deletion is gated behind an explicit confirmation, a group always keeps at least one file, and the highlighted file can be opened in the system default application. The design keeps the interaction logic pure and independently testable: an App model translates an abstract key event into an outcome, so navigation, marking, strategy application, and post-deletion model updates are all unit tested without a terminal. Rendering (ratatui) and the destructive file I/O live in the event loop, which restores the terminal on both normal exit and panic so a crash never leaves it broken. --tui is mutually exclusive with --keep and --interactive.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.