Skip to content

Repository files navigation

Mycelium Banner

Mycelium

Mycelium is an inspectable memory system for local AI agents. It keeps the original conversation as a durable record, turns useful information into an organized Markdown wiki, and brings the relevant parts back when they are needed later.

It is designed for users who want a local assistant that can build context over time with inspectable evidence and an organized wiki. You can chat with it through the included web app, inspect the complete evidence-to-claim pipeline, review proposed memory updates, or add the Python library to another agent.

Why use it?

Most chat assistants either forget everything between sessions or require the entire conversation history to be sent again. Mycelium takes a different approach:

  • Memory persists across conversations. Projects, preferences, decisions, research, and prior discussions can carry into a new session.
  • Everything stays inspectable. Raw logs and wiki views are Markdown; canonical sources, claims, decisions, and chat state are stored in SQLite and available through the inspector or JSON export.
  • Local models do the work. Chat, retrieval, and memory consolidation run through Ollama on your machine.
  • You stay in control. The UI shows the memories used for a response and requires review before contradictions or replacements change canonical claims.

Features

  • Multi-session chat with a local Ollama model
  • Hybrid automatic retrieval plus assistant-directed follow-up memory search
  • Durable short-term memory that is retrievable before wiki consolidation
  • Plain-text episodic logs and an Obsidian-compatible Markdown wiki
  • Evidence-triggered, claim-level reconsolidation with human review
  • Deterministic, read-only wiki projections with exact source provenance
  • Meeting ingestion pipeline - upload meeting audio to have it transcribed, diarized, and consolidated into the memory system
  • A Python API for adding Mycelium memory to other agents and frameworks

Storage and fresh stores

Canonical memory and chat state live in memory.sqlite3. Markdown under wiki/ and logs/ is a generated, inspectable view; edit memory through the application, not by modifying these files. LanceDB under indexes/ is rebuildable.

This version requires a fresh store. Existing JSON stores and benchmark runs are not migrated or modified. For the web app, select a new directory before running your normal launch command:

MYCELIUM_STORE=./mycelium_store_sqlite ./start.sh

One process owns each writable store. A second server, worker, or library process using that directory fails clearly; separate benchmark stores can run independently. Library owners should use Mycelium as a context manager or call close().

The memory inspector shows pending Markdown publication and offers Retry publication. Canonical edits commit atomically; retrying publication requires no model call. Invalid build plans are recorded as failed without partial canonical updates, and a later build can recompute them.

Export canonical records into a fresh directory for offline inspection:

.venv/bin/python -m mycelium.snapshots mycelium_store_sqlite memory-export

The export is JSONL by record collection, not another writable backend. Existing Engram meeting storage remains separate; old meetings are not automatically imported into a fresh memory store.

Quick start

Requirements

  • Python 3.11 or newer
  • uv
  • Node.js and npm
  • Ollama running locally

Clone and install dependencies:

git clone https://github.com/nsadras/mycelium.git
cd mycelium
uv sync
cd ui
npm install
cd ..

Mycelium is currently configured to use gemma4:12b for language tasks and embeddinggemma for memory retrieval. Download both models with Ollama, or change them in mycelium.toml:

ollama pull gemma4:12b
ollama pull embeddinggemma

With Ollama running, start the backend and frontend together:

./start.sh

Open http://localhost:5173 to use the app. The FastAPI backend is available at http://localhost:8000.

Using the app

The UI is organized around five main areas:

Area What it is for
Chat Create, rename, resume, and continue conversations. Each answer can show which memory pages were loaded and which tools were called.
Memory Inspect sources, segments, claims, Dream audits, and pending reconciliation proposals.
Wiki Browse and curate entity-owned views deterministically generated from canonical claims.
Logs Inspect the original episodic records that serve as source evidence for the wiki.
Engram Upload meeting audio, review the transcript and speakers, then save the finished meeting into memory.

A typical workflow is simple:

  1. Start a chat and use the assistant normally.
  2. Completed turns are saved automatically as sources, without extraction or embedding calls.
  3. Click Build Memory to extract statements from pending sources and organize the wiki.
  4. Start another chat to retrieve built memories and inspect their exact cited sources.
  5. Open Memory to inspect processing status and review proposed changes.

Capture alone does not make information searchable across sessions. Current-chat context works normally; unbuilt sources remain browsable and await Build Memory. There is no Flush control or scheduled build.

How memory works

Build Memory uses a stable snapshot of captured sources, resumes unfinished extraction, then runs the current organizer. New arrivals remain pending for the next build. Saved transcripts are durable retry inputs; repeated capture does not duplicate source records. A failed meeting summary does not prevent source admission.

The wiki distinguishes people, organizations, ongoing projects, recurring series, individual events, artifacts, places, and abstract topics. A meeting, tool, or deliverable can remain part of its larger context without creating an unnecessary standalone page. Responsibilities shared between a person and a project appear on both pages while remaining one source-backed memory.

Every wiki fact links back to its source. When new information conflicts with existing memory, Mycelium creates a review proposal instead of silently overwriting either version. The Memory and Wiki views let you inspect evidence, correct organization, merge duplicate subjects, and approve or reject proposed changes.

For the detailed lifecycle, entity model, retrieval design, and validation rules, see DESIGN.md.

Web search

To let the chat assistant use Ollama's web search and fetch tools, add an Ollama API key to a .env file in the project root:

OLLAMA_API_KEY=your_api_key_here

Tool calls and the result seen by the model are visible in the chat and retained as source observations for later consolidation.

Meeting memory with Engram

Engram is optional. Install its speech-processing dependencies separately:

uv sync --group engram

For speaker diarization, accept the terms for pyannote/speaker-diarization-community-1 on Hugging Face and provide a token:

export HF_TOKEN=your_hugging_face_token

Upload a recording from the Engram tab, click Process, review the generated transcript and speaker labels, then finalize it. Mycelium saves the transcript and structured meeting summary into the same memory system used by chat.

GPU acceleration is detected automatically when available. See DESIGN.md for model, device, and testing options.

Use Mycelium as a library

The web app is optional. The Python API exposes the same three-stage lifecycle used by the server:

Operation Input Output
ingest_source SourceInput with a transcript, source kind, participants, segments, and idempotency key IngestionResult with captured (or empty) status and durable source/log/episode/operation IDs; capture does not extract claims
retrieve_context RetrievalRequest with a query and context budget RetrievalResult with real wiki page references (page_references), typed evidence records and sources, authoritative Markdown/pseudo-XML rendering, and a retrieval trace
consolidate ConsolidationRequest with dry-run and deferred-claim policy ConsolidationResult with the build report and processed extraction episode IDs

For an ordinary agent turn, retrieve memory before generation and ingest the completed exchange afterward:

import asyncio

import mycelium


async def main():
    memory = mycelium.Mycelium(
        store_path="./agent_memory",
        ollama_model="gemma4:12b",
    )

    question = "What did we decide about the project architecture?"
    retrieval = await memory.retrieve_context(
        mycelium.RetrievalRequest(query=question)
    )

    # Supply retrieval.rendered_context as runtime evidence alongside the
    # question. Keep behavioral instructions in the model's system prompt.
    answer = "We chose a plain-text wiki backed by source logs."

    await memory.ingest_source(mycelium.SourceInput(
        transcript=f"USER: {question}\nASSISTANT: {answer}",
        session_id="architecture-chat",
        idempotency_key="architecture-chat:1",
    ))
    consolidation = await memory.consolidate(
        mycelium.ConsolidationRequest()
    )
    print(consolidation.report)


if __name__ == "__main__":
    asyncio.run(main())

Mycelium.session() remains an ergonomic wrapper around retrieval and ingestion for conversational agents. Web and library integrations use the same budgeted memory-context renderer. Call consolidate() explicitly to build captured sources into searchable memories and wiki pages. See examples/basic_session.py for a runnable example and examples/langgraph_integration.py for a LangGraph integration pattern.

Configuration

The main settings live in mycelium.toml:

[store]
path = "./mycelium_store"

[llm]
model = "gemma4:12b"
url = "http://localhost:11434"
context_window_tokens = 32768

[session]
context_budget_tokens = 32768

[retrieval]
embedding_model = "embeddinggemma:latest"
candidate_limit = 20
initial_result_limit = 5
tool_result_limit = 6
tool_search_limit = 3
tool_evidence_budget_tokens = 6000

The default memory store is ./mycelium_store. It consists primarily of Markdown and JSON, so it can be inspected with ordinary text tools or opened as a wiki outside the app. Every saved chat message carries its own timestamp, allowing one conversation to span multiple days without losing temporal context. session.context_budget_tokens is the total input budget shared by the assistant system prompt, recent transcript, initial memory, and follow-up memory evidence; it is capped by llm.context_window_tokens. Retrieval tool limits are per assistant response. During a response, the runtime accumulates initial retrieval and follow-up tool discoveries in one read-only evidence workspace. The model only chooses whether to search records or inspect a record's sources; the runtime handles merging, deduplication, and replacement of older workspace snapshots. The final workspace is persisted with the assistant message and is available from the chat's collapsed Evidence workspace inspector.

Architecture, storage contracts, retrieval details, migrations, development checks, and benchmark workflows are documented in DESIGN.md. The Daily Driver fixture has its own benchmark guide.

Replay the chat-to-memory smoke test

With the models configured in mycelium.toml available in your running Ollama instance:

MYCELIUM_RUN_CHAT_REPLAY=1 .venv/bin/pytest -q -s tests/test_chat_memory_replay.py

This opt-in integration test starts a fresh temporary store, captures the saved fried-rice and alignment conversations verbatim, runs Build Memory, checks that the single You page owns facts from both conversations, and asks the original cooking question in an empty third chat. It checks actual retrieved source citations, not answer keywords. Model calls and embeddings are real; the third chat has memory tools but no web tools. It may take several minutes and model outputs can vary. The normal test suite skips it.

The printed temporary directory contains the store, build report, rendered pages, and third-chat response for inspection (including on failure, up to the stage reached). Your live store is untouched. The fixture at tests/fixtures/chat_memory_replay.json contains personal conversation text saved with permission; review it before publishing or sharing the repository.

Additional focused real-model probes and capture/build/retrieval replays:

MYCELIUM_RUN_EXTRACTION_REPLAYS=1 .venv/bin/pytest -q -s tests/test_extraction_replays.py

These cover source-only conversation, unaccepted assistant suggestions, cross-turn acceptance/refusal with original-context citations, and facts embedded in questions versus purely informational questions. Exact accounting and provenance are checked deterministically; an evaluation-only model judge checks meaning without requiring exact output wording. The judge is not an independent quality oracle.

License

Mycelium is available under the MIT License. See LICENSE.

About

provenance-first memory system for local LLM agents

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages