fix(memory): make retrieval, compaction, and statistics complete at scale — prevent lossy results and quadratic work - #772
Closed
unohee wants to merge 3 commits into
Conversation
This was referenced Sep 27, 2026
unohee
pushed a commit
that referenced
this pull request
Sep 28, 2026
…past the first page Salvaged from draft PRs #772 and #776. Compaction, expiry cleanup, consolidation, and statistics all read the memory table with a single capped query, so larger stores were silently truncated: compaction refused outright at its 100k safety limit, while cleanupExpired, consolidateMemories and getMemoryStats quietly processed only the first 10k rows. They now read the complete table through memoryCore's paginated fetchAllTableRows / offset-limit pages. Deduplication is now bucket-based (stable FNV-1a hash of repo, type, derivedFrom and canonical metadata) instead of an O(n^2) pairwise scan, which matters now that the full record set is considered at once. Bucketing also guarantees duplicate pairs that straddle a page boundary meet: with the old page-by-page filtering they could not be compared. Within a bucket, records are ranked by importance then recency so the survivor is independent of page order. Also fixes the type defect in #772 where cosineSimilarity was declared : boolean while returning a similarity score, which broke the >= threshold comparison at the consolidation call site. Dropped from the drafts: getTable(tempTableName)/db.renameTable() (neither exists in the current memoryCore/lancedb API), the removals of withMemoryMutationLock/saveConversation/getRecentConversations/ runBackgroundCognition/getMemoryStats (main's API, still called by discordCore, support/chatMemory and support/web), package.json version churn, and the unconditional error-swallowing catch in compactMemoryTable that main deliberately rethrows.
unohee
pushed a commit
that referenced
this pull request
Sep 28, 2026
…past the first page Salvaged from draft PRs #772 and #776. Compaction, expiry cleanup, consolidation, and statistics all read the memory table with a single capped query, so larger stores were silently truncated: compaction refused outright at its 100k safety limit, while cleanupExpired, consolidateMemories and getMemoryStats quietly processed only the first 10k rows. They now read the complete table through memoryCore's paginated fetchAllTableRows / offset-limit pages. Deduplication is now bucket-based (stable FNV-1a hash of repo, type, derivedFrom and canonical metadata) instead of an O(n^2) pairwise scan, which matters now that the full record set is considered at once. Bucketing also guarantees duplicate pairs that straddle a page boundary meet: with the old page-by-page filtering they could not be compared. Within a bucket, records are ranked by importance then recency so the survivor is independent of page order. Also fixes the type defect in #772 where cosineSimilarity was declared : boolean while returning a similarity score, which broke the >= threshold comparison at the consolidation call site. Dropped from the drafts: getTable(tempTableName)/db.renameTable() (neither exists in the current memoryCore/lancedb API), the removals of withMemoryMutationLock/saveConversation/getRecentConversations/ runBackgroundCognition/getMemoryStats (main's API, still called by discordCore, support/chatMemory and support/web), package.json version churn, and the unconditional error-swallowing catch in compactMemoryTable that main deliberately rethrows.
unohee
pushed a commit
that referenced
this pull request
Sep 28, 2026
…past the first page Salvaged from draft PRs #772 and #776. Compaction, expiry cleanup, consolidation, and statistics all read the memory table with a single capped query, so larger stores were silently truncated: compaction refused outright at its 100k safety limit, while cleanupExpired, consolidateMemories and getMemoryStats quietly processed only the first 10k rows. They now read the complete table through memoryCore's paginated fetchAllTableRows / offset-limit pages. Deduplication is now bucket-based (stable FNV-1a hash of repo, type, derivedFrom and canonical metadata) instead of an O(n^2) pairwise scan, which matters now that the full record set is considered at once. Bucketing also guarantees duplicate pairs that straddle a page boundary meet: with the old page-by-page filtering they could not be compared. Within a bucket, records are ranked by importance then recency so the survivor is independent of page order. Also fixes the type defect in #772 where cosineSimilarity was declared : boolean while returning a similarity score, which broke the >= threshold comparison at the consolidation call site. Dropped from the drafts: getTable(tempTableName)/db.renameTable() (neither exists in the current memoryCore/lancedb API), the removals of withMemoryMutationLock/saveConversation/getRecentConversations/ runBackgroundCognition/getMemoryStats (main's API, still called by discordCore, support/chatMemory and support/web), package.json version churn, and the unconditional error-swallowing catch in compactMemoryTable that main deliberately rethrows.
unohee
added a commit
that referenced
this pull request
Sep 28, 2026
…past the first page (#785) Salvaged from draft PRs #772 and #776. Compaction, expiry cleanup, consolidation, and statistics all read the memory table with a single capped query, so larger stores were silently truncated: compaction refused outright at its 100k safety limit, while cleanupExpired, consolidateMemories and getMemoryStats quietly processed only the first 10k rows. They now read the complete table through memoryCore's paginated fetchAllTableRows / offset-limit pages. Deduplication is now bucket-based (stable FNV-1a hash of repo, type, derivedFrom and canonical metadata) instead of an O(n^2) pairwise scan, which matters now that the full record set is considered at once. Bucketing also guarantees duplicate pairs that straddle a page boundary meet: with the old page-by-page filtering they could not be compared. Within a bucket, records are ranked by importance then recency so the survivor is independent of page order. Also fixes the type defect in #772 where cosineSimilarity was declared : boolean while returning a similarity score, which broke the >= threshold comparison at the consolidation call site. Dropped from the drafts: getTable(tempTableName)/db.renameTable() (neither exists in the current memoryCore/lancedb API), the removals of withMemoryMutationLock/saveConversation/getRecentConversations/ runBackgroundCognition/getMemoryStats (main's API, still called by discordCore, support/chatMemory and support/web), package.json version churn, and the unconditional error-swallowing catch in compactMemoryTable that main deliberately rethrows. Co-authored-by: openSwarm <openswarm@local>
Collaborator
Author
|
Closing: superseded — the salvageable work in this draft was rebased onto current main, verified, and re-published as a reviewed PR: absorbed into #785 (salvage/memory-lifecycle). Closing the stale draft instead of merging it, because it was 170-250 commits behind, carried scratch files, and (in several cases) reverted main's later hardening. |
unohee
deleted the
swarm/AGT-3491-fix-memory-make-retrieval-compaction-and
branch
September 28, 2026 07:02
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Published because this run stopped and needs a human: autonomous execution failed 4 times
It has not been reviewed and is very likely incomplete — this PR is a draft on purpose. It exists so the work is reviewable instead of sitting on a branch that was never pushed.
Change shape
420 file(s): 208 source · 189 test · 3 docs · 20 other
Base freshness
Branch base is 175 commit(s) behind
mainat publication.⚠ Conflicts with
mainin:package-lock.json,package.json,src/memory/memoryOps.tsGitHub runs no
pull_requestworkflows on a PR it cannot merge, so any green checks here are not the verification gate. Opened as a draft; rebase before review.This branch changes files that other open PRs / active branches also touch. Coordinate before merging to avoid divergent parallel edits (INT-2388 #3):
package-lock.jsonpackage-lock.json,package.jsonsrc/memory/compaction.tspackage-lock.jsonpackage-lock.jsonpackage-lock.jsonpackage-lock.json,package.jsonpackage-lock.jsonpackage-lock.jsonpackage-lock.json,package.json,src/memory/memoryOps.tspackage-lock.json,package.jsonpackage-lock.jsonpackage-lock.json,package.json,src/memory/memoryOps.tspackage-lock.jsonpackage-lock.json,package.jsonpackage-lock.jsonpackage-lock.jsonpackage-lock.jsonpackage-lock.json,package.jsonpackage-lock.jsonpackage-lock.jsonpackage-lock.jsonpackage-lock.json,package.jsonpackage-lock.json,package.jsonpackage-lock.json,package.jsonpackage-lock.json,package.jsonLinear
Closes AGT-3491
🤖 Generated with OpenSwarm