Repository navigation
feat(graph): report edge counts, and narrow the NREM read in SQL - #83
Merged
Merged
Conversation
`memory_stats` and the dashboard reported `graph_nodes` and nothing else, so
the size of the graph was invisible to every tool that could read it. The
numbers behind this session's measurements were all raw SQL against the
database files, which makes them incomparable between runs.
Add `EpistemicGraph.count_edges()`, returning `{"total", "live", "dead"}`, and
surface it as `graph_edges` / `graph_edges_live` / `graph_edges_dead` in both
`memory_stats` and `dashboard.get_stats`. All three default to 0, so existing
consumers keep working.
`epi_edges` has no `layer` column, so an edge is attributed to its SOURCE
node's layer via a join — the convention `memory_graph_edges` already uses. An
edge whose source node is gone lands in no bucket rather than a guessed one;
there is a test for that.
"Dead" is `weight <= 0` rather than `weight < NREM_FLOOR`: across all three
house databases only 0.1% of edges fall strictly between zero and the 0.05
floor, so the coarser split is the accurate one. A dead edge is not garbage —
it is the state the next miner pass revives it from.
`_dream_nrem` also read every heuristic edge and decided age in Python, to
discard most of them: on the busiest database, 155 486 rows read to act on a
few thousand, and 0 rows for mimocode's over-30-day band. The two actionable
age bands are now the query's bounds. This is the same comparison the loop
made — `created_at < now − 30 d` is `age_days > 30`, and
`created_at >= now − 1 d` is `age_days <= 1`, inclusive, as before. Measured
row reduction: hermes 49.9%, mimocode 68.0%, cowagent 42.1%.
No index is added. The filter still reads 42–68% of the table and the time
fell with the row count (190 → 115 ms, 205 → 79 ms), so an index on
`created_at` would buy nothing while keeping its write cost.
A NULL `created_at` is now skipped instead of raising `float(None)`. No house
database has one; "unknown age" is not a reason to decay, boost, or prune.
Tests: four on `count_edges` (live/dead split, empty graph, source-node layer,
orphan edge) and one on NREM asserting a ten-day-old edge is left exactly as
it was, which is what proves narrowing the read kept the behaviour. `test_stats`
and `test_dashboard` assert `total == live + dead`.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this fixes
memory_statsand the dashboard reportedgraph_nodesand nothing else. The size of the graph was invisible to every tool that could read it — the numbers behind this session's measurements were all raw SQL against the database files, which makes them incomparable between runs.Changes
1. Edges become visible. New
EpistemicGraph.count_edges()returns{"total", "live", "dead"}in one query, surfaced asgraph_edges/graph_edges_live/graph_edges_deadinmemory_statsanddashboard.get_stats. All three default to0, so existing consumers are unaffected.epi_edgeshas nolayercolumn, so an edge is attributed to its source node's layer via a join — the conventionmemory_graph_edgesalready uses. An edge whose source node is gone lands in no bucket rather than a guessed one."Dead" is
weight <= 0, notweight < NREM_FLOOR: across all three house databases only 0.1% of edges fall strictly between zero and the 0.05 floor (103 of 186 523), so the coarser split is the accurate one. A dead edge is not garbage — it is the state the next miner pass revives it from, which is why nothing prunes zero eagerly.2.
_dream_nremreads only what it can act on. It previously read every heuristic edge and decided age in Python, discarding most: on the busiest database, 155 486 rows read to act on a few thousand, and 0 rows for mimocode's over-30-day band. The two actionable age bands are now the query's bounds.The comparison is the same one the loop made —
created_at < now − 30 disage_days > 30, andcreated_at >= now − 1 disage_days <= 1, inclusive, as before.A NULL
created_atis now skipped instead of raisingfloat(None). No house database has one; "unknown age" is not a reason to decay, boost, or prune.No index, deliberately
An index on
created_atwas considered and rejected on measurement: the filter still reads 42–68% of the table and the time fell with the row count (190 → 115 ms, 205 → 79 ms). An index at that selectivity would buy nothing while keeping its write cost.Tests
count_edges— live/dead split, empty graph, source-node layer attribution, orphan edge;test_statsandtest_dashboardasserttotal == live + dead.Full suite: 2078 passed.
Cost
count_edgeswith the layer join, measured on the live databases: 97 ms on hermes (186 878 edges), 30 ms on cowagent. It runs only when the tool is called.