Skip to content

fix(codebase): hub git routing + write refactors - #16

Merged
ManSio merged 1 commit into
mainfrom
fix/codebase-hub-write-actions
Aug 28, 2026
Merged

ManSio merged 1 commit into
mainfrom
fix/codebase-hub-write-actions

Conversation

@ManSio

@ManSio ManSio commented Aug 28, 2026

Copy link
Copy Markdown
Owner

Summary

Fixes 4 real bugs in the \codebase\ hub (write-actions dry-run, move import, safe_delete callers, git log routing) and syncs AGENTS.md §2 tool registry.

Fixes

  • codebase(action=git): now resolves the real \project_root\ and dispatches \log/\history/\�ranch/\ ile\ to the dedicated git tools instead of hitting a non-existent \log\ path (was: \Path does not exist: .../log).
  • write_tools.safe_delete: counts real usages via \ ind_all_references\ (was \ ind_references, which missed cross-file callers and reported a false \

- codebase(action=git) resolves the real project_root and dispatches log/history/branch/file to dedicated git tools instead of a non-existent log path (fixes 'Path does not exist: .../log').
- write_tools safe_delete counts real usages via find_all_references (was find_references, which missed cross-file callers -> false '0 usages').
- write_tools move infers a valid dotted module relative to project root instead of a Windows absolute path (fixes 'from D:\... import ...').
- replace/insert dry-run already guarded by apply=False (regression added).
- docs: AGENTS.md §2 synced to registered tool set; KNOWN_ISSUES updated.

Regression: tests/test_codebase_hub.py, tests/test_write_tools.py (67 passed).
@coderabbitai

coderabbitai Bot commented Aug 28, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: a6973c93-4aed-4b8a-b12c-049413dff0c6


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ManSio
ManSio merged commit fc170e6 into main Aug 28, 2026
12 of 13 checks passed
@ManSio
ManSio deleted the fix/codebase-hub-write-actions branch August 28, 2026 13:25
ManSio added a commit that referenced this pull request Sep 25, 2026
…11) (#40)

* style(search): strip trailing whitespace in graph_adapter

* fix(release): align version to 3.5.0 in extension.toml and __init__

* docs: sync version references to 3.5.0

stale_detector flagged 20 outdated version strings across 17 docs (en/ru/zh). Updated to the pyproject version; stale_detector now reports 0 drift.

* test(e17): add signal-value experiment harness

Measures what execution-derived TESTS edges buy an LLM across four tasks.

- e17_extract.py: AST extraction of functions/tests (fixes qualified-name
  match that returned FUNC NOT FOUND for every class method).
- e17_pilot_experiment.py: specificity- and complexity-selected panel,
  unique/foreign/coverage-matched decoy, edge-case and impact questions.
- tooling: line-budget prompt shards for isolated subagent generation,
  blind test-hidden judging, objective impact recall/precision.
- report v1-v4: reasoning tasks hit a ceiling (B never beats A); impact
  ('which tests cover F') is decisive for the signal (A=0.00, B=1.00).
- 13 extraction tests including adversarial cases.

Result recorded in AGENT_DIARY.md and KNOWN_ISSUES.md.

* test(e17): validate TESTS-signal as a verifier via mutation testing

Negate the first single-line if inside 10 panel functions, run the graph-linked
tests (arm B) and coverage-matched decoy tests (arm D), restore via git checkout.

- B killed the mutant 6/10, D 0/10 (Fisher exact p=0.011).
- 4/10 linked tests did not catch a strong mutation (weak assertions, e.g.
  TestCypherLexer asserts only len(tokens) > 0).
- Conclusion: the signal identifies verifying tests far better than decoys, but
  its value is bounded by assertion strength, not coverage.

* test(e17): profile verification strength across multiple mutations

10 functions x 4 mutations (negate-if, flip-compare, bool-flip, int+1,
str-append); every graph-linked test and decoy run under every mutant.

- 24 tests profiled; decoy tests killed 0/104 mutants (control clean).
- kill-rate predictors (Spearman): assertion specificity rho=+0.54 p=0.006,
  assert count rho=+0.47 p=0.019, density/LOC not significant.
- What an assertion compares against matters more than how many it has.
  TESTS edge = reach; mutation kill-rate = verification.

* test(e17): scale strength profiling; AST proxies do not predict kill-rate

24 functions (providers/mcp excluded), 55 tests, 4 mutations each; split by
function; decoys killed 0/55.

- v6 correlation (exact_compare rho=+0.54) did NOT replicate: train rho=+0.13,
  held-out rho=-0.01; rule precision 0.59 / recall 0.67.
- AST features (assert count/density/specificity/LOC) do not predict whether a
  test catches a behaviour change; TEST_VERIFIES needs real mutation runs.
- Report marks v6 CORRECTED; history preserved.

* docs(known-issues): archive autosynced experiment entries

* fix(search): revive idle embedder, unblock hot-reload in search

Two independent causes of search_code >15s timeouts (2026-09-22):

- llama_runner watchdog unloads the embedder after EMBEDDER_IDLE_TIMEOUT
  (120s, free RAM) but RemoteEmbedder never brought it back -> :8080
  connection refused (WinError 10061) -> per-chunk retries -> timeout.
  Add LlamaRunner.is_port_up()/ensure_embedder_started() and revive the
  server once on embed failure, before per-item retries.
- _maybe_hot_reload awaited a full reindex inside search_code; a batch of
  new/changed files (22) pushed the call past the 15s budget even with a
  live embedder. Run it as a deduped background task.
- Also fix the inverted model key in the (currently unreferenced)
  LlamaRunner.embed (bge-m3 -> DEFAULT_EMBEDDING_MODEL).

Tests: tests/test_embedder_idle_recovery.py (10).

* docs(experiments): record E6 commit-gate and E7 symptom-index eval

* docs(experiments): finalize E7 eval (3 models, arrival-symptom NONE 10/10)

* feat(graph): add quiet-break isolation gate (delta A+B)

Detect commits that silently isolate a graph node, read-only (git diff + persisted PropertyGraph, no index mutation). Absolute zero-inbound-CALLS is unusable (1535/4081 = 38% of repo), so two delta signals: removed_last_caller (a symbol still defined loses its last caller) and new_orphan (a newly added function nothing calls). Exposed as graph_query(action=isolation) / CLI; fail-open (unavailable/empty) when git or graph is absent. Controls: unit 10/10, real-git E2E 4/4, CLI 1899ms 0 false positives.

* feat(cli): redact credentials at delivery boundary

Delivery notes (hooks -> python -m src.cli -> model context) bypass the commit secret-gate, so a key pasted into a note as evidence would leak into every agent context the note fires on. Add an in-house prefix-anchored redactor (src/core/redact.py) and apply it at the single CLI choke point (success and error JSON). Anchored on structural prefixes, not entropy, so paths, commit hashes and digests survive byte-identical. Not a security boundary (novel/split-line formats pass) - documented as such. Design credited to Tom Jones, crystal-memory scripts/redact.py (Apache-2.0, not vendored). Tests: 8/8 (unit + CLI integration).

* feat(delivery): add anti-numbing restraint for advisory notes

Repeated identical advisory notes numb the reader (Tom Jones crystal-memory; our E3/E4). Add src/core/restraint.py: a findings signature (kind+symbol+file) with a per-project cooldown and 2^strikes backoff, state in data_root (atomic replace). Opt-in via graph_query isolation kwargs={restraint:true}; blocking gates never use it (a blocked commit is a hard stop, not nagging). Tests 4/4 (22/22 with gate+redact); CLI E2E: call1 deliver, immediate call2 suppressed (cooldown 597s).

* docs(experiments): record E11 frozen list vs new arrival index

Blind mapper against Tom's catalogue arrival index (tjonesit/crystals @3e30ed2): the dispositional arrival symptom (#16 agent won't use tools) now reaches family A in 5/5 runs including both VALID runs (longcat-2.0, qwen3.7-plus, controls 6/6). deepseek x3 invalid (NC3 UI-spacing lure, same NONE-control failure as E7). New risk noted: the arrival layer can manufacture a false positive on an adjacent domain.

* docs: clean-state verification 44d451f (1800 passed, all gates green)

Local-clone clean state via verify_clean_state.sh --no-clone on committed revision 44d451f: 1800 passed / 0 failed / 13 skipped / 98 deselected (822s); lock-drift negative control PASSED; guard inventory ALL PROVEN (3); revision gate VALID. Note: when a branch is not pushed, the default GitHub clone verifies origin, not this branch - clone the local repo and use --no-clone.

* test(graph): make TESTS-signal fixture platform-independent

test_graph_stage_e4 hardcoded file_path D:/Project/... which is not absolute on POSIX, so get_tests_for_symbol normalized it to <root>/D:/... and found no node -> 0 covering tests. Passed on Windows, failed on ubuntu (2 tests). Build absolute paths from tmp_path.as_posix() instead; fixture is now platform-independent.

* docs: record CI green on PR #40 (ubuntu+windows+clean-state)

Fix 808864c turned the ubuntu CI green; full run success. Only remaining red is GitHub Advanced Security (Copilot Autofind), which fails with CAPIError 400 model-not-supported on GitHub's side, not our code.

* feat(opencode): add commit-time gate hook (stale block + isolation advisory)

Project-local opencode plugin: on git commit, run stale_detector (block on objective doc drift) and graph_query isolation (advisory; graph gate has FP risk). Advisory uses before-stash/after-deliver by callID (E8 lesson). CLI invoked as python -m src.cli with cwd=project (CLI ignores --project). Env switches MSCODEBASE_PY/STALE_GATE/ISOLATION_GATE/GATE_LOG. Activated after Zed reload; global install is out of workspace and needs explicit approval.

* fix(opencode): invoke CLI via node:child_process, not plugin context

In opencode 1.18.23 the plugin factory context does not provide a callable
shell helper (docs describe v2 where it exists as BunShell). The hook loaded
(init logged) but threw "TypeError: ... is not a function" on the shell call.

Call python through node:child_process.execFile with an argv array (no shell
quoting) instead of the context shell. node:fs already worked, so node builtins
are available in the plugin runtime.

* docs: record opencode gate hook live verification

Fresh init after restart; git commit --dry-run -> stale drift=0, isolation count=0.
Negative control (staged new_orphan added in a separate git add step) -> isolation
count=1 and the advisory note delivered into the command output; test file removed,
tree clean. Note: the hook runs before the command, so git add && git commit in one
command leaves the file unstaged at hook time and the gate cannot see it.

* fix(docs): verify_references no longer crashes on a lone dollar in backticks

auto_update_docs(action=verify) raised IndexError "string index out of range":
the empty-content guard ran before the $ prefix was stripped, so a lone $
(or $ ) in backticks left content empty and content[0] blew up. Re-check for
empty content after stripping $. Adds a regression test.

* feat(docs): deterministic L1 doc-reference checker (near-zero FP)

The old verifier collected only definitions and scanned venv via rglob, so it
reported 1366 broken with ~3% real (params/fields/stdlib). Replace the reference
check with src/core/doc_reference_l1.py: vocabulary = identifiers + string
literals from src/tests/scripts/tools (+root scripts); scope = live docs (root
docs minus ledgers, docs/** minus archive/generated/research/blog/ISSUES/
investigations); whitelist stdlib/typing, external APIs, env vars, model names,
commit scopes. auto_update_docs(action=verify) now delegates to it.

Effect (measured): broken 1366 -> 0. Also fixes WISDOM naming a non-existent
_symbol_resolver (actual: _build_symbol_resolver). Tests 4/4.
Limitation: a stale tool name that still appears anywhere in code/tests is
masked; tool names need a registry check, not a vocabulary check.

* docs(experiments): record E13 test-time measurement (xdist ~15%)

xdist -n 4: 1817 passed in 169s vs 200s serial (~15%), no failures.
Not worth default complexity; CI cold clean-state (~13m) is dominated by
install/caches, not CPU. Named next levers: cache pip in clean-state and
investigate test_temporal_facts_generator (~42s).

* docs: sync tool counts to code (65 = 32 core + 16 intel + 13 inline + 4 dev)

Authoritative count derived from code (server_tools.py tool_classes, tools_reg.py and
dev_tools.py decorators, server_tools.py inline decorators): 65 registered tools.
Fix contradicting numbers in README (61/66), AGENTS (64, 31 core) and
docs/{en,ru,zh}/ARCHITECTURE (64/61/58/46, 28/31/20 core); describe the visible set
via the MSCODEBASE_MCP_TOOLS allowlist instead of a stale magic number. L1 checker
clean (0 broken); tests 12/12.

* docs(readme): semantic audit fixes against code

Verified claims against code and corrected: tests badge 1856 -> 1889
(_count_tests), last-updated date, DI services 18 -> 14 (di_container.py has 14
add_singleton calls), tests 1773 -> 1889, and the tool-count line (32+16+13+4=65,
+execute_script -> 66). Confirmed correct: 29 EdgeType members and 13
structural_search patterns. Open: AGENTS.md lists get_repo_map/... as "not
registered" while they are registered tool_classes (hidden by allowlist) and
README lists them - needs a decision.

* docs(agents): semantic audit fixes against code

Fixed: source-of-truth path src/core/intelligence_layer.py does not exist ->
src/core/intelligence/layer.py; CI matrix "(3.10-3.12)" -> Python 3.14-only.
Confirmed correct: RuntimeCoordinator/can_execute, ProjectContext.capture()
(project_context.py:164), apply_file_move (indexer.py:659), root files.
Open: the "Not registered (consolidated)" block lists get_repo_map/get_hotspots/
verify_action/get_task_status etc. as not exposed, but they are registered
tool_classes (hidden by the default allowlist) - needs an owner decision.

* docs+fix: resolve "not registered" tools; correct server_tools log counts

Mechanism: only core tool_classes in the MSCODEBASE_MCP_TOOLS allowlist are
registered via mcp.tool(); intel(16)/inline(13)/dev(4) are always registered.
So the AGENTS claim was right in spirit (hidden -> tool not found) but the list
was wrong (verify_action/get_task_status are in the default allowlist; bootstrap_
pipeline was missing). Rewrite the block as "Hidden core tools (default allowlist)";
add a visibility note to README; fix the startup log hardcode in server_tools.py
(intel 14->16, inline 12->13) that printed "62 total" while 65 register.

* docs(architecture): semantic audit fixes (en/ru/zh) against code

- Layer 2 "Bridge" marked deprecated (code: not used, LSP server removed 2026-07-20).
- server.py "~220 lines" -> 68 (actual line count).
- core "30 files" -> 54 top-level (123 incl. subpackages).
- DI container "15+/18 services" -> 14.
- README: lsp_main.py "does not work" -> "removed" (file deleted).
Confirmed correct: tool counts 32/16/13/4=65 and module paths.

* docs(architecture): fix services/labels/tool-group counts against code

- remote_embedder described as ONNX-only -> llama.cpp GGUF primary with ONNX/LM
  Studio/Ollama fallback.
- "15 node labels" -> 16 (NodeLabel members).
- DI "Registered Services (11)" -> 14 (add Path, Indexer, GitUrlSourceFactoryKey).
- Tool groups: Intelligence 14 -> 16 (add restore/supersede), Lifecycle 3 -> 4
  (add get_action_receipt).

---------

Co-authored-by: MSCodeBase Agent <mscodebase@intelligence.local>
ManSio pushed a commit that referenced this pull request Sep 26, 2026
…production)

Blind mapper via opencode (longcat-2.0 x5, qwen3.7-plus x3, deepseek-v4.1-flash x3), isolated dir, MCP off, --pure, tools denied. Valid 5/11; in all valid runs #16 -> NONE (arrival symptom unreachable); the only failure is the same known false positive (#11 Safari/CSS -> a-generated-document-is-unverified), concentrated in deepseek (0/3 valid). Qualitative reproduction of E7 confirmed; numeric rate NOT reproduced (5/11 vs ~10/11) because --variant was not pinned. Raw runs and manifest under results/recovered_e7/.
ManSio pushed a commit that referenced this pull request Sep 26, 2026
Replace ad-hoc PowerShell+Tee with scripts/f4_blind_run.py (subprocess argv-list, UTF-8, ANSI-strip); --variant is required so an unpinned run fails loudly. Rerun with --variant high: valid 10/11 (longcat 4/5, qwen 3/3, deepseek 3/3), controls 6/6 in all valid, #16 -> NONE everywhere. The earlier 'deepseek 0/3' was an unpinned-budget confound, not a model property. Pitfall #19 added to the registry.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant