Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

ย 

History

155 Commits
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

Second Brain Logo

Second Brain

Personal Knowledge Management & Learning System

"Tell me and I forget, teach me and I may remember, involve me and I learn." โ€” Benjamin Franklin

A comprehensive system for ingesting, organizing, connecting, and actively learning from personal and professional knowledge sourcesโ€”powered by LLMs and graph-based knowledge representation.


๐Ÿ“‘ Table of Contents


๐ŸŽฏ Vision

Transform passive information consumption into active knowledge acquisition through:

  1. Automated ingestion of diverse data sources
  2. Intelligent summarization and connection discovery
  3. Deliberate practice via AI-generated exercises and spaced repetition

๐Ÿง  Core Philosophy

The Two-Fold Challenge

Challenge Focus Solution Approach
Extraction & Summarization LLM-powered Automated pipelines that distill raw sources into structured, interconnected notes
Learning & Retention Human-centered Active exercises, spaced repetition, and deliberate practice systems

Key Insight

These challenges can be addressed independently, but solving extraction in service of learning maximizes value. Every piece of ingested content should feed into the learning loop.


๐Ÿ“Š System Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                              DATA SOURCES                                    โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚   Papers    โ”‚  Articles   โ”‚   Books     โ”‚    Code     โ”‚   Ideas & Notes     โ”‚
โ”‚  (Books.app)โ”‚ (Raindrop)  โ”‚ (Physical)  โ”‚ (Git repos) โ”‚   (Manual input)    โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
       โ”‚             โ”‚             โ”‚             โ”‚                 โ”‚
       โ–ผ             โ–ผ             โ–ผ             โ–ผ                 โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                         INGESTION LAYER                                      โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  โ€ข PDF Parser + Highlight Extractor + Handwriting OCR (Vision LLM)           โ”‚
โ”‚  โ€ข Raindrop API Client                                                       โ”‚
โ”‚  โ€ข Book Photo OCR Pipeline (Mistral Vision API)                              โ”‚
โ”‚  โ€ข Git/GitHub API Integration                                                โ”‚
โ”‚  โ€ข Manual/CLI Entry Tools                                                    โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                    โ”‚
                                    โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                       PROCESSING LAYER (LLM-Powered)                         โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  โ€ข Summarization Engine                                                      โ”‚
โ”‚  โ€ข Key Concept Extraction                                                    โ”‚
โ”‚  โ€ข Tag & Topic Classification                                                โ”‚
โ”‚  โ€ข Connection Discovery (semantic similarity)                                โ”‚
โ”‚  โ€ข Follow-up Task Generation                                                 โ”‚
โ”‚  โ€ข Exercise & Quiz Generation                                                โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                    โ”‚
                                    โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                         KNOWLEDGE HUB (Obsidian)                             โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  ๐Ÿ“ vault/                                                                   โ”‚
โ”‚  โ”œโ”€โ”€ ๐Ÿ“ sources/           # Raw ingested content organized by type          โ”‚
โ”‚  โ”‚   โ”œโ”€โ”€ ๐Ÿ“ papers/                                                          โ”‚
โ”‚  โ”‚   โ”œโ”€โ”€ ๐Ÿ“ articles/                                                        โ”‚
โ”‚  โ”‚   โ”œโ”€โ”€ ๐Ÿ“ books/                                                           โ”‚
โ”‚  โ”‚   โ”œโ”€โ”€ ๐Ÿ“ code/                                                            โ”‚
โ”‚  โ”‚   โ””โ”€โ”€ ๐Ÿ“ ideas/                                                           โ”‚
โ”‚  โ”œโ”€โ”€ ๐Ÿ“ topics/            # Topic-based index notes (auto-generated)        โ”‚
โ”‚  โ”œโ”€โ”€ ๐Ÿ“ projects/          # Active learning projects                        โ”‚
โ”‚  โ”œโ”€โ”€ ๐Ÿ“ exercises/         # Generated practice problems                     โ”‚
โ”‚  โ”œโ”€โ”€ ๐Ÿ“ reviews/           # Spaced repetition queue                         โ”‚
โ”‚  โ””โ”€โ”€ ๐Ÿ“ meta/              # System config, templates, scripts               โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                    โ”‚
                                    โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                         KNOWLEDGE GRAPH (Neo4j)                              โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  Nodes: Concepts, Sources, Topics, Authors, Tags                             โ”‚
โ”‚  Edges: RELATES_TO, CITES, CONTRADICTS, EXTENDS, PREREQUISITE_FOR           โ”‚
โ”‚  Queries: "What do I know about X?", "What connects A to B?"                 โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                    โ”‚
                    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                    โ–ผ                               โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚     LEARNING SYSTEM (FSRS)       โ”‚  โ”‚       AI ASSISTANT (LLM Agent)       โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค  โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  โ€ข Spaced repetition scheduling  โ”‚  โ”‚  โ€ข Natural language chat interface   โ”‚
โ”‚  โ€ข Flashcard management          โ”‚  โ”‚  โ€ข RAG over vault & knowledge graph  โ”‚
โ”‚  โ€ข Mastery tracking              โ”‚  โ”‚  โ€ข Streaming responses with citationsโ”‚
โ”‚  โ€ข Practice session orchestrationโ”‚  โ”‚  โ€ข Context-aware follow-up questions โ”‚
โ”‚  โ€ข Weak-spot identification      โ”‚  โ”‚  โ€ข Configurable LLM model selection  โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ”‚  โ€ข Conversation history & sessions   โ”‚
                                      โ”‚  โ€ข Source linking to original notes  โ”‚
                                      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿ“ฅ Ingestion Pipelines

1. Academic Papers (PDF)

Source: MacOS Books app, Zotero, direct PDF uploads

The PDF pipeline uses a hybrid approach:

  • Mistral OCR โ€“ Single API call that extracts full document text (markdown-formatted with tables & figures) AND detects handwritten notes/diagrams via image annotations
  • PyMuPDF โ€“ Separate pass to extract PDF annotation objects (highlights, underlines, comments, sticky notes) from the PDF structure
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                         PDF INPUT                                โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                              โ”‚
           โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
           โ–ผ                                     โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”           โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚      Mistral OCR        โ”‚           โ”‚       PyMuPDF           โ”‚
โ”‚   (single API call)     โ”‚           โ”‚  (PDF structure parse)  โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค           โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
โ”‚ โ€ข Full text (markdown)  โ”‚                        โ”‚
โ”‚ โ€ข Tables & figures      โ”‚                        โ–ผ
โ”‚ โ€ข Handwritten notes     โ”‚           โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚   (via image analysis)  โ”‚           โ”‚   Digital Annotations   โ”‚
โ”‚ โ€ข Diagrams detected     โ”‚           โ”‚  โ€ข Highlights           โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜           โ”‚  โ€ข Underlines           โ”‚
             โ”‚                        โ”‚  โ€ข Comments/sticky notesโ”‚
             โ”‚                        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
             โ”‚                                     โ”‚
             โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                              โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚     Unified Content Merge      โ”‚
              โ”‚  (associate annotations with   โ”‚
              โ”‚   their context in the paper)  โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                              โ”‚
                              โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚      LLM Summarization         โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                              โ”‚
                              โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚       Markdown Note            โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Output: Structured markdown with summary, key findings, highlights, handwritten notes with context, and auto-generated follow-up questions.

2. Web Content (Articles, Blog Posts)

Source: Raindrop.io API

Raindrop Collection โ†’ API fetch โ†’ Content extraction โ†’
LLM Summarization โ†’ Markdown note with highlights preserved

Integration Points:

  • Scheduled sync (daily/hourly)
  • Preserve Raindrop collections as Obsidian folders or tags
  • Extract user highlights as blockquotes
  • Archive original content (avoid link rot)

3. Physical Books

Source: Photos of highlighted pages

Photo โ†’ Mistral Vision OCR โ†’ Highlight extraction โ†’
Text cleanup โ†’ LLM processing โ†’ Structured book notes

Workflow:

  1. Photograph highlighted pages with consistent lighting
  2. Batch process through OCR pipeline
  3. AI identifies highlighted vs. non-highlighted text
  4. Aggregate into chapter-based or theme-based notes
  5. Store original images in separate media vault

4. Code & Repositories

Source: GitHub starred repos, personal projects

Git repo โ†’ Structure analysis โ†’ README parsing โ†’
Key file identification โ†’ LLM code summarization โ†’
Markdown note with architecture overview, key patterns, learnings

Captured Elements:

  • Repository purpose and architecture
  • Notable design patterns
  • Dependencies and technology stack
  • Personal notes on why it was saved
  • Code snippets worth remembering

5. Ideas & Fleeting Notes

Source: CLI tool, mobile app, voice memos

Quick capture โ†’ Inbox folder โ†’ Daily processing โ†’
Elaboration or linking to existing notes

Quick capture sends items to an inbox folder for daily processing, elaboration, and linking to existing notes.


๐Ÿท๏ธ Organization Strategy

Primary: Content Type Hierarchy

sources/
โ”œโ”€โ”€ papers/      # Academic papers, research
โ”œโ”€โ”€ articles/    # Blog posts, news, essays
โ”œโ”€โ”€ books/       # Book notes and highlights
โ”œโ”€โ”€ code/        # Repository analyses
โ”œโ”€โ”€ ideas/       # Fleeting notes, thoughts
โ””โ”€โ”€ work/        # Meetings, proposals, slack

Secondary: Semantic Tags

Hierarchical topic tags (ml/transformers, systems/distributed) and meta tags (status/actionable, quality/foundational).

Tertiary: Bidirectional Links

Leverage Obsidian's [[wikilinks]] extensively:

  • Every note should link to related concepts
  • Use block references for granular connections
  • Auto-generate backlink summaries

๐ŸŽ“ Learning & Deliberate Practice System

๐Ÿ“– Full details: See LEARNING_THEORY.md for research foundations and citations.

Learning Science Foundation

This system is grounded in research on human memory and learning. Key insights:

Research Key Finding System Implementation
Ericsson (2008) โ€” Deliberate Practice Expertise requires structured practice with feedback, not just experience Adaptive difficulty + immediate LLM feedback
Bjork & Bjork (2011) โ€” Desirable Difficulties Spacing, interleaving, and generation enhance long-term retention Spaced repetition + varied exercises
Dunlosky et al. (2013) โ€” Learning Techniques Practice testing and distributed practice are highest utility; highlighting/rereading are lowest Retrieval-based exercises, avoid recognition tasks
Van Gog et al. (2011) โ€” Cognitive Load Worked examples before problems for novices Adaptive: examples โ†’ testing as mastery increases
Chi et al. (1994) โ€” Self-Explanation Prompting self-explanation builds correct mental models Self-explanation prompts in exercises

Core Principles

  1. Learning โ‰  Performance: Easy recall during study (retrieval strength) doesn't guarantee long-term retention (storage strength)
  2. Generation over Recognition: Producing answers from memory beats re-reading or highlighting
  3. Desirable Difficulties: Spacing, interleaving, testing, and variation slow immediate performance but enhance retention
  4. Adaptive Scaffolding: Novices get worked examples; intermediates get retrieval practice

Exercise Types

              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚    1. INGEST NEW CONTENT    โ”‚
              โ”‚   (automated pipelines)     โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                            โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚   2. UNDERSTAND & CONNECT   โ”‚
              โ”‚   (summarization, linking)  โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                            โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚    3. ACTIVE PRACTICE       โ”‚โ—„โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚  (generation, not review)   โ”‚         โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜         โ”‚
                            โ”‚                         โ”‚
                            โ–ผ                         โ”‚
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”         โ”‚
              โ”‚   4. SPACED RETRIEVAL       โ”‚         โ”‚
              โ”‚  (testing > restudying)     โ”‚โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Exercise Generation

Content Type Exercise Types Desirable Difficulty Applied
Conceptual Explain-in-own-words, compare/contrast, teach-back Generation effect (no notes allowed)
Technical Implement from scratch, debug code, extend functionality Generation + Variation
Procedural Reconstruct steps from memory, adapt to new scenario Retrieval practice + Interleaving
Analytical Case study analysis, predict outcomes, critique approaches Generation + Spacing

Spaced Repetition Integration

  • Generate Anki-compatible flashcards from key concepts
  • Schedule review sessions based on forgetting curves (FSRS algorithm)
  • Track confidence levels per concept
  • Surface weak areas for targeted practice

๐Ÿ› ๏ธ Technical Stack

Python
Python 3.11+
FastAPI
FastAPI
React
React 18
Vite
Vite
TailwindCSS
TailwindCSS
PostgreSQL
PostgreSQL
Neo4j
Neo4j
Redis
Redis
Docker
Docker
Obsidian
Obsidian
LiteLLM
LiteLLM
Celery
Celery
Mistral
Mistral OCR
Gemini
Gemini
Anthropic
Anthropic
OpenAI
OpenAI
Raindrop
Raindrop.io
GitHub
GitHub API

Core Technologies

Component Technology Rationale
Frontend React + Vite + TailwindCSS Modern, fast, great DX
Backend FastAPI + Python Async, type-safe, OpenAPI docs
Knowledge Hub Obsidian Markdown-based, local-first, extensible
Graph Database Neo4j Native graph storage, Cypher queries
Relational DB PostgreSQL Learning records, user data
Cache Redis Session state, rate limiting
Task Queue Celery Async background job processing
LLM Backbone LiteLLM (GitHub) Unified interface to 100+ LLMs (OpenAI, Anthropic, Gemini, Mistral)
Vision/OCR Mistral OCR (default for PDFs), Gemini 3 Flash Document processing, handwriting recognition

APIs & Services

Service Purpose
Raindrop.io Web bookmark sync
GitHub Repository analysis
Mistral Primary OCR for PDF/document processing
Google (Gemini) Default text LLM for summarization, exercises

๐Ÿ–ฅ๏ธ Web Application

The Second Brain web application provides a full-featured interface for knowledge management, spaced repetition learning, and analytics.

Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚    Frontend     โ”‚โ”€โ”€โ”€โ”€โ–ถโ”‚     Backend     โ”‚โ”€โ”€โ”€โ”€โ–ถโ”‚   Data Layer    โ”‚
โ”‚  (React/Vite)   โ”‚     โ”‚   (FastAPI)     โ”‚     โ”‚ Neo4j/PG/Redis  โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Frontend Pages: Dashboard, Practice Session, Exercises Catalogue, Card Catalogue, Review Queue, Knowledge Explorer, Knowledge Graph, Analytics, Follow-up Tasks, LLM Usage, Learning Assistant, Settings

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                           FRONTEND (React)                                    โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”‚
โ”‚  โ”‚   Practice  โ”‚  โ”‚   Review    โ”‚  โ”‚  Analytics  โ”‚  โ”‚   Knowledge         โ”‚  โ”‚
โ”‚  โ”‚   Session   โ”‚  โ”‚   Queue     โ”‚  โ”‚  Dashboard  โ”‚  โ”‚   Explorer          โ”‚  โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ”‚
โ”‚         โ”‚               โ”‚                โ”‚                    โ”‚              โ”‚
โ”‚  โ€ข Free recall    โ€ข Spaced cards   โ€ข Learning curves    โ€ข Graph viz         โ”‚
โ”‚  โ€ข Self-explain   โ€ข Due items      โ€ข Topic mastery      โ€ข Note browser      โ”‚
โ”‚  โ€ข Worked examplesโ€ข Confidence     โ€ข Time invested      โ€ข Connection map    โ”‚
โ”‚  โ€ข Interleaved Qs   ratings        โ€ข Weak spots         โ€ข Search            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                      โ”‚
                                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                           BACKEND (FastAPI)                                   โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  /api/practice/*        /api/review/*         /api/analytics/*               โ”‚
โ”‚  โ”œโ”€โ”€ generate-exercise  โ”œโ”€โ”€ due-items         โ”œโ”€โ”€ learning-curve            โ”‚
โ”‚  โ”œโ”€โ”€ submit-response    โ”œโ”€โ”€ update-card       โ”œโ”€โ”€ topic-mastery             โ”‚
โ”‚  โ”œโ”€โ”€ get-feedback       โ”œโ”€โ”€ schedule          โ”œโ”€โ”€ session-history           โ”‚
โ”‚  โ””โ”€โ”€ self-explain       โ””โ”€โ”€ confidence        โ””โ”€โ”€ weak-spots                โ”‚
โ”‚                                                                              โ”‚
โ”‚  /api/knowledge/*       /api/ingest/*         /api/assistant/*              โ”‚
โ”‚  โ”œโ”€โ”€ graph              โ”œโ”€โ”€ pdf               โ”œโ”€โ”€ chat                      โ”‚
โ”‚  โ”œโ”€โ”€ search             โ”œโ”€โ”€ raindrop          โ”œโ”€โ”€ generate-questions        โ”‚
โ”‚  โ”œโ”€โ”€ connections        โ”œโ”€โ”€ ocr               โ””โ”€โ”€ explain-connection        โ”‚
โ”‚  โ””โ”€โ”€ topics             โ””โ”€โ”€ github                                          โ”‚
โ”‚                                                                              โ”‚
โ”‚  /api/capture/*                                                             โ”‚
โ”‚  โ”œโ”€โ”€ text, url, photo, voice, pdf, book                                     โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                      โ”‚
                                      โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                              DATA LAYER                                       โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚        Neo4j            โ”‚       PostgreSQL        โ”‚        Redis             โ”‚
โ”‚   Knowledge Graph       โ”‚    Learning Records     โ”‚    Session Cache         โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ โ€ข Concepts & relations  โ”‚ โ€ข Practice attempts     โ”‚ โ€ข Active sessions        โ”‚
โ”‚ โ€ข Source documents      โ”‚ โ€ข Confidence ratings    โ”‚ โ€ข Temp exercise state    โ”‚
โ”‚ โ€ข Topic hierarchies     โ”‚ โ€ข Spaced rep schedule   โ”‚ โ€ข Rate limiting          โ”‚
โ”‚ โ€ข Semantic embeddings   โ”‚ โ€ข Time tracking         โ”‚                          โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Backend APIs: /api/practice/*, /api/review/*, /api/analytics/*, /api/knowledge/*, /api/ingest/*, /api/assistant/*, /api/capture/*

Screenshots

The application features a modern, dark-themed interface optimized for focused learning and knowledge management.

Dashboard

Dashboard

The Dashboard serves as your home screen, answering "What should I do today?" at a glance. It displays a stats header showing your current streak, due review cards, and daily progress toward your learning goals. Quick action cards provide one-click access to practice sessions and review queues. The page also surfaces due cards for immediate review, identifies weak spots requiring attention, shows a streak calendar for motivation, and includes a quick capture input for rapidly saving ideas or URLs.

Practice Session

Practice Session

The Practice Session page enables deep learning through structured exercises grounded in cognitive science. You can select topics hierarchically with visual mastery indicators showing your current level, configure session duration (5-30 minutes), and choose whether to reuse existing exercises or generate new AI-powered ones. Exercise types include free recall (retrieve from memory), self-explanation (explain concepts in your own words), worked examples (study solutions before attempting), code debugging, and teach-back promptsโ€”all with immediate LLM-powered feedback on your responses.

Exercises Catalogue

Exercises Catalogue

The Exercises Catalogue provides a comprehensive browsing interface for all available exercises in the system. You can search exercises with full-text search, filter by exercise type (recall, explain, apply, code) and difficulty level, group by topic, and jump directly into practice mode. Each exercise card shows its type, difficulty, associated topic, and when it was last practiced, helping you identify fresh material or areas needing review.

Card Catalogue

Card Catalogue

The Card Catalogue lets you browse and manage all spaced repetition flashcards in your knowledge base. Cards are organized by topic and state (new, learning, review, mastered), with filters for card type (definition, comparison, application, example, concept). You can search cards, view their front/back content, see scheduling information, and track mastery progress. This page complements the Review Queue by providing a library view of all your cards rather than just those currently due.

Review Queue (Spaced Repetition)

Review Queue

The Review Queue implements evidence-based spaced repetition using the FSRS (Free Spaced Repetition Scheduler) algorithm. Cards due for review are presented one at a time with active recallโ€”you type your answer before seeing the correct response, which is more effective than simple recognition. The LLM evaluates your answers for semantic correctness, allowing for variations in wording. You then rate your confidence (Again/Hard/Good/Easy) to adjust scheduling. The page also supports AI-powered card generation from your notes.

Knowledge Explorer

Knowledge Explorer

The Knowledge Explorer provides a unified interface for browsing your entire knowledge base stored in Obsidian. Toggle between tree view (folder hierarchy) and list view (flat listing), use real-time search to find notes instantly, and access the command palette with โŒ˜K for quick navigation. Selected notes render inline with full markdown support including syntax highlighting for code blocks, LaTeX math, and wiki-link navigation. Deep linking support means you can share URLs to specific notes.

Knowledge Graph

Knowledge Graph

The Knowledge Graph offers an interactive D3.js force-directed visualization of your Neo4j knowledge graph. Different node types (Content, Concepts, Notes) are color-coded, with edges representing relationships like RELATES_TO, CITES, EXTENDS, and PREREQUISITE_FOR. Click nodes to view details, drag to rearrange the layout, and scroll to zoom in/out. A statistics sidebar shows content breakdown by type. This visualization helps discover unexpected connections between ideas and identify knowledge clusters.

Analytics Dashboard

Analytics Dashboard

The Analytics Dashboard provides comprehensive insights into your learning journey. The stats grid displays total time invested, current streak, and overall mastery percentage. Activity charts show practice sessions and review activity over configurable time periods (7/30/90 days). A topic mastery radar visualizes your proficiency across different knowledge areas. Progress breakdowns show completion by topic, while weak spots analysis identifies topics with declining retention and provides "Practice Now" buttons for targeted improvement. Calculated insights surface trends and recommendations.

Follow-up Tasks

Follow-up Tasks

The Follow-up Tasks page displays actionable items generated automatically during content processing. When you ingest a paper, article, or book, the LLM identifies potential follow-up actions: research topics to explore, concepts to practice, connections to make with other notes, and applications to try. Tasks are categorized by type (research, practice, connect, apply, review) and priority (high, medium, low), with estimated time requirements. You can filter, search, and mark tasks complete as you work through them, turning passive reading into active engagement.

Ingest

Ingest

The Ingest page provides a unified interface for capturing new content directly from the desktop web UI and monitoring the entire ingestion pipeline. The top panel offers tabbed capture for text notes, URLs, and file uploads (PDFs, images, audio) with expandable options for tagging, title, and learning material generation (cards/exercises). The bottom panel displays a live, auto-refreshing ingestion queue showing all content items across every status (Pending, Processing, Completed, Failed). Click any item to expand a detail panel with full processing stage progress, error messages, cost/token stats, and a direct link to the generated note in the Knowledge Explorer.

LLM Usage

LLM Usage

The LLM Usage page provides a dashboard for monitoring your AI API usage and costs. It displays budget status with visual progress bars, spending trends over time, and breakdowns by model (GPT-4, Claude, Gemini, Mistral) and pipeline (ingestion, processing, exercises, assistant). This transparency helps you understand where AI costs go and optimize your usage. You can set monthly budgets and receive alerts when approaching limits.

Learning Assistant

Learning Assistant

The Learning Assistant is an AI-powered chat interface for exploring your knowledge base conversationally. Ask natural language questions like "What do I know about attention mechanisms?" or "How does paper X relate to paper Y?" The assistant searches your vault and knowledge graph, synthesizes information, and provides source citations linking back to your notes. You can configure which LLM model to use, and responses stream in real-time with full markdown rendering. This turns your knowledge base into an interactive, queryable resource.

Settings

Settings

The Settings page lets you customize the application to your preferences. Appearance settings include compact mode for denser information display and animation toggles for reduced motion. Learning preferences let you set default session lengths and daily practice goals. Keyboard shortcuts are configurable, with defaults like โŒ˜K for command palette and โŒ˜1-6 for page navigation. You can also manage notification preferences, configure LLM model defaults, and export your data for backup or migration purposes.


๐Ÿ“ฑ Mobile Capture (PWA)

A critical bottleneck in knowledge management is capture friction โ€” the effort required to get information into the system. The companion Progressive Web App provides a mobile-optimized interface for on-the-go capture with offline support.

Capture Types

Scenario Capture Method Processing
Physical book highlight Photo of page Vision OCR โ†’ highlight extraction โ†’ ingest
Fleeting idea Voice memo or text Transcription โ†’ LLM expansion โ†’ inbox
Interesting article Share sheet / URL Content fetch โ†’ summarize โ†’ save
Whiteboard / diagram Photo Vision LLM โ†’ describe โ†’ save with image
PDF document File upload Mistral OCR โ†’ full processing pipeline

Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                     MOBILE DEVICE                                โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”          โ”‚
โ”‚   โ”‚   ๐Ÿ“ท Camera   โ”‚  โ”‚   ๐ŸŽค Voice   โ”‚  โ”‚   ๐Ÿ“Ž Share   โ”‚          โ”‚
โ”‚   โ”‚  (book pages, โ”‚  โ”‚   (ideas,    โ”‚  โ”‚   (URLs,     โ”‚          โ”‚
โ”‚   โ”‚  whiteboards) โ”‚  โ”‚   memos)     โ”‚  โ”‚   articles)  โ”‚          โ”‚
โ”‚   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜          โ”‚
โ”‚          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                   โ”‚
โ”‚                   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                            โ”‚
โ”‚                   โ”‚   PWA / Mobile   โ”‚                            โ”‚
โ”‚                   โ”‚   Quick Capture  โ”‚                            โ”‚
โ”‚                   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                             โ”‚ Upload (queue if offline)
                             โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                      BACKEND                                     โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                                  โ”‚
โ”‚   /api/capture/photo     โ†’ Vision OCR โ†’ Text extraction          โ”‚
โ”‚   /api/capture/voice     โ†’ Whisper transcription โ†’ LLM expand    โ”‚
โ”‚   /api/capture/url       โ†’ Content fetch โ†’ Summarize             โ”‚
โ”‚   /api/capture/text      โ†’ Save to inbox โ†’ Tag suggestion        โ”‚
โ”‚   /api/capture/pdf       โ†’ Mistral OCR โ†’ Full pipeline           โ”‚
โ”‚   /api/capture/book      โ†’ Batch page OCR โ†’ Book notes           โ”‚
โ”‚                                                                  โ”‚
โ”‚                   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                         โ”‚
โ”‚                   โ”‚   Inbox Processing  โ”‚                         โ”‚
โ”‚                   โ”‚  (async via Celery) โ”‚                         โ”‚
โ”‚                   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                         โ”‚
โ”‚                             โ”‚                                    โ”‚
โ”‚              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                     โ”‚
โ”‚              โ–ผ              โ–ผ              โ–ผ                     โ”‚
โ”‚         Neo4j          Obsidian        PostgreSQL                โ”‚
โ”‚      (concepts)         (notes)       (metadata)                 โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Key Features

Feature Details
Installable "Add to Home Screen" โ€” launches like a native app
Offline capable Service worker caches assets and queues captures in IndexedDB
Background sync Automatically uploads queued captures when connection is restored
Share target Receive shared URLs, text, and images from other apps (Android)
< 3 second capture Minimal UI with large touch targets optimized for speed

The PWA runs as a separate lightweight frontend (localhost:5174) and communicates with the same backend API. All captures are processed asynchronously via Celery and flow into the standard ingestion pipeline.

See 08_mobile_capture.md for full design details.

Screenshots

Capture Home ย ย ย  Text Capture ย ย ย  URL Capture

Left to right: Main capture screen with all capture types, Quick Note text capture, URL capture for saving links


๐Ÿš€ Getting Started

Prerequisites

  • Python 3.11+
  • Docker Desktop installed and running
  • At least one LLM API key (Gemini, Mistral, OpenAI, or Anthropic)

Quick Start

# Clone the repository
git clone https://github.com/<your-username>/second-brain.git
cd second-brain

# Run the interactive setup script
python scripts/setup_project.py

The setup script guides you through:

  1. Environment configuration โ€” API keys, database credentials, data directory
  2. Vault setup โ€” Obsidian folder structure, templates, meta notes
  3. Docker services โ€” Start PostgreSQL, Neo4j, Redis, backend, frontend
  4. Database migrations โ€” Initialize schema

Setup Options

python scripts/setup_project.py                    # Full interactive setup
python scripts/setup_project.py --non-interactive # Use defaults
python scripts/setup_project.py --env-only        # Only configure .env
python scripts/setup_project.py --help-only       # Show all available commands
python scripts/setup_project.py --help-env        # Show env variable reference

Access Points

Service URL Description
Frontend http://localhost:3000 Main web application
Knowledge Graph http://localhost:3000/graph Interactive graph visualization
Mobile Capture PWA http://localhost:5174 Mobile-optimized capture app
Backend API http://localhost:8000 REST API endpoints
API Documentation http://localhost:8000/docs Swagger/OpenAPI docs
Neo4j Browser http://localhost:7474 Graph database UI

Local Development (without Docker)

Backend: cd backend && pip install -r requirements.txt && uvicorn app.main:app --reload

Frontend: cd frontend && npm install && npm run dev

Platform-Specific Setup

macOS

Prerequisites:

# Install Homebrew (if not installed)
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"

# Install Python 3.11+
brew install python@3.11

# Install Docker Desktop
# Download from: https://www.docker.com/products/docker-desktop/
# Or via Homebrew:
brew install --cask docker

# Verify installations
python3 --version    # Should be 3.11+
docker --version     # Should show Docker version
docker compose version

Notes:

  • Docker Desktop must be running before docker compose commands
  • On Apple Silicon (M1/M2/M3), Docker automatically handles ARM64 architecture
  • The ~ tilde expands correctly on macOS for local development
Linux (Ubuntu/Debian)

Prerequisites:

# Update package list
sudo apt update

# Install Python 3.11+
sudo apt install python3.11 python3.11-venv python3-pip

# Install Docker (official method)
# Remove old versions
sudo apt remove docker docker-engine docker.io containerd runc

# Install prerequisites
sudo apt install ca-certificates curl gnupg lsb-release

# Add Docker's official GPG key
sudo mkdir -p /etc/apt/keyrings
curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /etc/apt/keyrings/docker.gpg

# Set up repository
echo "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] https://download.docker.com/linux/ubuntu $(lsb_release -cs) stable" | sudo tee /etc/apt/sources.list.d/docker.list > /dev/null

# Install Docker
sudo apt update
sudo apt install docker-ce docker-ce-cli containerd.io docker-compose-plugin

# Add your user to the docker group (to run without sudo)
sudo usermod -aG docker $USER
newgrp docker

# Verify installations
python3 --version
docker --version
docker compose version

Notes:

  • Log out and back in for docker group changes to take effect
  • For systemd services, use absolute paths (tilde ~ won't expand)
  • If running in WSL2, see Windows section for additional notes
Windows (with WSL2)

Prerequisites:

  1. Install WSL2:

    # Run in PowerShell as Administrator
    wsl --install
    # Restart your computer
  2. Install Docker Desktop:

    • Download from: https://www.docker.com/products/docker-desktop/
    • During installation, enable "Use WSL 2 based engine"
    • After installation, open Docker Desktop Settings โ†’ Resources โ†’ WSL Integration
    • Enable integration with your WSL distribution
  3. In WSL2 terminal (Ubuntu):

    # Install Python
    sudo apt update
    sudo apt install python3.11 python3.11-venv python3-pip
    
    # Verify Docker (provided by Docker Desktop)
    docker --version
    docker compose version

Notes:

  • Run all commands from within WSL2, not PowerShell
  • Store your project in the WSL filesystem (/home/user/) not /mnt/c/ for better performance
  • Use absolute paths in .env file (e.g., /home/user/data not ~/data)
  • Docker Desktop manages the Docker daemon; you don't need to start it manually

Verifying Your Setup

After installation, verify everything is working:

# Check Python version (should be 3.11+)
python3 --version

# Check Docker is running
docker info

# Check Docker Compose
docker compose version

# Test Docker can run containers
docker run hello-world

# Verify GPU support (optional, for local LLM inference)
docker run --rm --gpus all nvidia/cuda:11.8.0-base-ubuntu22.04 nvidia-smi

Troubleshooting Common Issues

Docker daemon not running

macOS/Windows: Start Docker Desktop application.

Linux:

sudo systemctl start docker
sudo systemctl enable docker  # Start on boot
Permission denied when running docker

Linux:

sudo usermod -aG docker $USER
# Log out and back in, or run:
newgrp docker
Port already in use

Check what's using the port:

# macOS/Linux
lsof -i :8000  # Backend
lsof -i :3000  # Frontend
lsof -i :5432  # PostgreSQL

# Stop the process or use different ports in docker-compose.yml
Database connection refused
  1. Check if containers are running: docker compose ps
  2. Check container logs: docker compose logs postgres
  3. Verify .env file has correct credentials
  4. Wait for healthcheck to pass (can take 30 seconds)
Neo4j won't start / Out of memory

Neo4j requires significant memory. Ensure Docker Desktop has at least 4GB RAM allocated:

  • Docker Desktop: Settings โ†’ Resources โ†’ Memory โ†’ Set to 4GB+
  • Linux: Check available memory with free -h

Useful Commands

Command Purpose
python scripts/pipelines/run_pipeline.py article <URL> Import web article
python scripts/pipelines/run_pipeline.py pdf <file> Process PDF document
python scripts/pipelines/run_pipeline.py book <file> OCR book photos
python scripts/run_processing.py process-pending Process all pending content
python scripts/run_all_tests.py Run all tests
docker compose logs -f backend View backend logs
docker compose down -v Stop and remove all data

๐Ÿ“‹ Implementation Status

๐Ÿ“ Full Details: See implementation_plan/OVERVIEW.md for the complete implementation roadmap with task checklists.

Phase Focus Status
1 Foundation & Infrastructure โœ… Complete
2 Ingestion Pipelines โœ… Complete
3 LLM Processing โœ… Complete
4 Knowledge Graph (Neo4j) โœ… Complete
5 Backend API โœ… Complete
6 Frontend Application โœ… Complete
7 Learning System (Exercises + FSRS) โœ… Complete
8 Analytics Dashboard โœ… Complete
9 Mobile Capture (PWA) โœ… Complete
10 Assistant Tool Calling โฌœ Not Started
11 MCP Integration โฌœ Not Started
12 Polish & Production ๐ŸŸก In Progress

๐Ÿ“š Documentation

Design Documents

Detailed technical specifications for each system component:

Document Description
00_system_overview.md High-level architecture and component interactions
01_ingestion_layer.md Content ingestion pipelines and formats
02_llm_processing_layer.md LLM integration, prompts, and processing stages
03_knowledge_hub_obsidian.md Obsidian vault structure and templates
04_knowledge_graph_neo4j.md Neo4j schema, queries, and graph operations
05_learning_system.md Exercises, FSRS algorithm, mastery tracking
06_backend_api.md FastAPI endpoints and data models
07_frontend_application.md React components and state management
08_mobile_capture.md PWA mobile capture workflow
09_assistant_tool_calling.md LLM agent with tool calling
10_observability.md Logging, metrics, and monitoring

Implementation Plans

Step-by-step implementation guides with task checklists:

Document Description
OVERVIEW.md Master roadmap with all phases
00_foundation_implementation.md Infrastructure setup
01_ingestion_layer_implementation.md Content ingestion
02_llm_processing_implementation.md LLM processing stages
03_knowledge_hub_obsidian_implementation.md Obsidian integration
04_knowledge_graph_neo4j_implementation.md Neo4j setup and queries
05_learning_system_implementation.md Learning system
06_backend_api_implementation.md API development
07_frontend_application_implementation.md Frontend development
08_mobile_capture_implementation.md Mobile PWA
09_assistant_tool_calling_implementation.md Assistant tool calling
tech_debt.md Technical debt tracking

Other Documentation

Document Description
LEARNING_THEORY.md Learning science research foundations
TESTING.md Testing guide and best practices

๐Ÿ”ฌ Open Research Questions

  1. Human vs. Machine Connection-Making: To what extent should we outsource relationship discovery to AI vs. keeping it as a human cognitive exercise?

  2. Information Overload: How do we prevent the knowledge base from becoming overwhelming? What pruning and archival strategies work best?

  3. Exercise Quality: Can current LLMs generate exercises that genuinely challenge and teach, or do they tend toward superficial quizzes?


๐Ÿ”ฎ Future Extensions

Tool Calling for Learning Assistant

Enable the Learning Assistant to take actions through natural language requests, turning it from a Q&A interface into an interactive agent:

User: "Generate an exercise about attention mechanisms"
      โ†’ Assistant calls generate_exercise tool
      โ†’ Returns interactive exercise card inline in chat

Planned Tools:

Tool Description
generate_exercise Generate adaptive exercise for a topic based on current mastery
create_flashcard Create a spaced repetition card from conversation context
search_knowledge Search the knowledge graph with natural language
get_mastery Retrieve mastery state and learning history for a topic
get_weak_spots Identify topics with declining retention needing review

The design uses an LLM tool-calling loop: the model decides when to invoke tools, results are fed back for a synthesized response. See 09_assistant_tool_calling.md for the full design.

MCP Integration (Model Context Protocol)

Expose the knowledge base as MCP servers so any MCP-compatible LLM client (Claude Desktop, Cursor, etc.) can directly query your Second Brain:

User (in Claude Desktop): "What do I know about distributed consensus?"
      โ†’ LLM queries Second Brain MCP server
      โ†’ Server searches Obsidian vault + Neo4j graph
      โ†’ Returns relevant notes with citations

Planned MCP Servers:

Server Capabilities
Vault server Read/search/write Obsidian notes, list by topic
Knowledge graph server Cypher queries, concept lookup, relationship traversal
Learning server Exercise generation, spaced rep scheduling, mastery queries

๐Ÿšข Production Deployment

For production deployments, see the comprehensive guides in docs/deployment/:

Document Description
production.md Full production deployment guide
security.md Security hardening and best practices

Key Production Steps:

  1. Configure environment โ€” Set production values in .env (disable debug mode, set real secrets)
  2. SSL/TLS โ€” Use Let's Encrypt with Certbot for HTTPS certificates
  3. Reverse proxy โ€” Configure Nginx for rate limiting, security headers, and proxying
  4. CORS โ€” Restrict CORS_ORIGINS to your production domains
  5. Database security โ€” Strong passwords, network isolation, regular backups
  6. Container security โ€” Run as non-root, set resource limits, use read-only filesystems where possible

Quick Production Checklist:

# Required environment changes for production
DEBUG=false
SECRET_KEY=<generate-secure-random-key>
CORS_ORIGINS=https://yourdomain.com
POSTGRES_PASSWORD=<strong-password>
NEO4J_PASSWORD=<strong-password>

See production.md for complete instructions including Docker configuration, Nginx setup, backup procedures, and monitoring.


๐Ÿค Contributing

We welcome contributions! Please see our Contributing Guide for details on:

  • Development environment setup
  • Code style guidelines (Python and JavaScript/React)
  • Commit message conventions
  • Pull request process
  • Testing requirements

Quick Start for Contributors:

# Fork and clone the repository
git clone https://github.com/<your-username>/second-brain.git
cd second-brain

# Run the setup script
python scripts/setup_project.py

# Create a feature branch
git checkout -b feature/your-feature-name

# Make changes, then submit a PR

๐Ÿ”’ Security

For security-related concerns, please review:

  • Security Hardening Guide โ€” Production security best practices
  • Vulnerability Reporting โ€” If you discover a security vulnerability, please report it responsibly by emailing the maintainers directly rather than opening a public issue

Security Features:

  • Configurable CORS origins (not wildcard in production)
  • Environment-based secrets management
  • Database credential isolation
  • Rate limiting support
  • Security headers via reverse proxy

๐Ÿ“„ License

This project is licensed under the MIT License โ€” see the LICENSE file for details.

The MIT License is a permissive license that allows:

  • โœ… Commercial use
  • โœ… Modification
  • โœ… Distribution
  • โœ… Private use

With the only requirement being to include the license and copyright notice in copies.


๐Ÿ“š References

Learning Science

๐Ÿ“– See LEARNING_THEORY.md for detailed research summaries.

Key sources: Ericsson (2008) on Deliberate Practice, Bjork & Bjork (2011) on Desirable Difficulties, Dunlosky et al. (2013) on Effective Learning Techniques.

Knowledge Management

AI-Assisted Learning

Tools & Plugins


This is a living document. As the system evolves, so will this design.

About

No description, website, or topics provided.

Resources

Contributing

Security policy

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages