A grounded, queryable, auditable knowledge-graph assistant built from a small, validated snapshot of Open Food Facts.
The project is designed for four contributors working from four laptops and for a complete, reproducible live demonstration in the local Evidence Workbench or Jupyter.
The Workbench also exposes the fixture-matched, versioned competition report: the current measured result is 32/32 with evaluator v4. The report still publishes per-case anchor presence, execution mode and reproducibility hashes instead of replacing generated output with oracle answers.
Recipe Studio adds a second useful graph workflow: it composes four original fixture-bound recipes from mapped products, scales quantities, calculates label-based nutrition and shows both the preparation roadmap and auditable recipe-to-product subgraph.
The searchable recipe catalog also includes a committed, revision-pinned snapshot of 320 English Wikibooks Cookbook recipes. These are attributed external sources, not certified composition templates: source-listed diet categories and deterministic ingredient coverage are visible, while nutrition and suitability claims are withheld.
The Conversations workspace uses an ElevenLabs conversational agent for a
continuous, hands-free call. FastAPI keeps the API key server-side and gives
the authenticated browser only a short-lived signed conversation URL. The
agent's search_foodgraph client tool queries the local read-only FoodGraph
search endpoint before it makes product, ingredient, allergen, nutrition or
source claims. Configure ELEVENLABS_API_KEY, ELEVENLABS_AGENT_ID and
ELEVENLABS_VOICE_ID in the ignored local .env.
Voice uses ElevenLabs Agents for turn detection, interruption, reasoning and speech. Neo4j remains the factual source through the client tool. Voice calls are not added to the local text-chat history or activity journal, and FoodGraph does not store raw audio; provider processing and retention follow the configured ElevenLabs account settings. The separate Play answer control still streams ElevenLabs TTS for an already certified text answer.
Given a plain-language product request, FoodGraph can:
- use vector search to find semantically relevant products;
- traverse ingredients, allergens, additives, labels, brands, categories, and provenance;
- answer only from the retrieved subgraph;
- generate and display a read-only Cypher query for a structured question;
- answer "How do we know this?" with source, source agent, captured time, and extraction confidence.
FoodGraph explores crowdsourced label data. It is not dietary, allergy, or medical advice; users must verify the physical product label.
The competition reference build includes the frozen fixture, graph, offline vectors, three-tool Evidence Agent with deterministic tool-derived answers, bounded process-local conversations, a 32-case benchmark and a local evidence-first web interface. It also includes two versioned Recipe Studio templates projected into Neo4j without changing the original 130-product / 7,490-fact snapshot counts. The current live-model benchmark passes 32/32 cases: all schema-guided structured queries, unsafe-query checks, provenance audits, retrieval anchors and replay-agent cases pass. The remaining team release gate is a clean-clone rehearsal on another laptop.
| Workstream | Responsibility | Roadmap |
|---|---|---|
| P1 - Data | Open Food Facts acquisition, curation, fixture validation | Person 1 |
| P2 - Graph | Neo4j schema, idempotent import, indexes, graph queries | Person 2 |
| P3 - Retrieval & Query | embeddings, vector + traversal retrieval, safe Text2Cypher | Person 3 |
| P4 - Audit & Demo | provenance, evaluation, final notebook, presentation | Person 4 |
Install Neo4j Desktop 2.2.1 with an Enterprise Developer instance running Neo4j 2026.06.x / Cypher 25. All four laptops use that same baseline. Create/start the local instance, then create the dedicated database from the system database:
CREATE DATABASE foodgraph IF NOT EXISTS;Then, from the repository root:
conda env create -f environment.yml
conda activate environment
python -m pip install -e .
python -m ipykernel install --user --name environment --display-name "Python (environment)"
cp .env.example .env
python scripts/validate_fixture.py
python scripts/validate_artifacts.py
python scripts/import_graph.py
python scripts/import_embeddings.py
jupyter notebook notebooks/foodgraph_demo.ipynbInstall the web dependencies once, then start the API and React app together:
cd frontend
npm ci
cd ..
python scripts/run_web_demo.pyOpen http://127.0.0.1:5173. The runner binds both services to localhost,
keeps secrets in the Python process, and stops both processes on Ctrl+C.
Ground, Query, Audit and Agent all accept user input. Demo preset restores
the fixture-bound competition question for a reproducible presentation.
Conversations retains at most eight ephemeral server-side turns and supports
certified source, previous-result allergen and first-two comparison follow-ups;
New conversation starts a blank thread while preserving the signed-in
user's saved history; a saved conversation is removed only through its explicit
delete action.
When ELEVENLABS_API_KEY and ELEVENLABS_VOICE_ID are configured locally,
Start voice conversation opens one hands-free session: Scribe Realtime
detects pauses and submits turns automatically, FoodGraph answers through the
same local evidence pipeline, and streamed speech can be interrupted by
speaking again. The browser receives only a short-lived transcription token;
the permanent provider key remains on FastAPI. FoodGraph does not retain raw
audio, and typing remains available in the same composer.
Recipe Studio is available from the Workspace sidebar. Its calculations
are deterministic and local; recipe options are explicit mappings rather than
LLM-generated substitutions.
My workspace adds optional local accounts: it keeps a salted password
verifier, an opaque HttpOnly local session, draft/published recipes,
preferences and a concise activity journal in the same local Neo4j database.
It is not cloud sync or OAuth. Generated recipes must remain drafts until the
owner explicitly publishes them.
Use python scripts/run_web_demo.py --offline for the deterministic fallback
rehearsal. If your ignored environment file is elsewhere, pass its path with
--env-file; the runner never prints its values.
The Docker stack is self-contained: it starts an isolated Neo4j Enterprise Developer database, imports the committed snapshot and embeddings once, then starts the FastAPI service and production React workbench. It does not reuse a Neo4j container belonging to another project.
docker compose up --buildOpen http://127.0.0.1:5173. The initial import can take a short time; wait
until the seed service exits successfully and web is healthy. The default
ports are 5173 (Workbench), 8000 (API), 7476 (Neo4j HTTP) and 7690
(Neo4j Bolt), deliberately avoiding the ports used by the other local Neo4j
project. The localhost-only Docker demo account is admin / admin; normal
accounts still require an email address and a password of at least 10
characters. The local Docker password defaults to foodgraph-local; override
it only for this stack with FOODGRAPH_NEO4J_PASSWORD.
# Check the complete stack without reading container logs for secrets.
docker compose ps
curl http://127.0.0.1:8000/api/meta
# Stop it; add --volumes only when intentionally resetting FoodGraph data.
docker compose downThe frontend baseline is pinned in .nvmrc (Node 22.22.3) and the lockfile.
To rehearse a complete model outage without changing the notebook, execute it with
FOODGRAPH_OFFLINE=1; the exact-question vector and reviewed Cypher are loaded
from fixture-bound committed fallbacks. FOODGRAPH_QUERY_MODEL defaults to the
fast course model gemma4:4b for Text2Cypher. The live Evidence Agent uses
FOODGRAPH_AGENT_MODEL when set and otherwise the tool-capable LLM_MODEL
configured by the course endpoint. In the current local setup that is
gemma4:31b-it-q8_0; it is slower than the query model because it is larger and
may take several tool-calling turns. The current endpoint is an external course
service, so "local" describes the application, Neo4j and data—not LLM
inference. Never commit .env.
Run every non-destructive local release gate, including the offline notebook:
python scripts/preflight.py --include-notebookThe competition benchmark can also be regenerated independently with
python scripts/run_competition_evaluation.py. It always writes the report and
exits zero only when all 32 gates pass.
- Loop Spec
- Global roadmap
- Architecture
- Recipe Network & identity loop spec
- Data contract
- Live demo script
- Test plan
- Git workflow
- Issue backlog
- Decision log
- Release handoff
- Agentic AI competition plan
- Measured evaluation report
- Evidence Workbench guide
- Multi-page Workbench Loop Spec
- Reusable multi-page Loop Engineering prompt
- Conversational AI Loop Spec
- Recipe Studio Loop Spec
- Reusable Conversational AI Loop Engineering prompt
- Four reconstruction tutorials
- Live presentation deck
- Competition presentation deck:
presentation/FoodGraph_Competition_Demo.pptx - Day-one checklist
- Contributing
- Grounded GraphRAG: "Find breakfast products similar to a chocolate-hazelnut spread with no declared palm-oil marker in this snapshot; show allergens, brand, Nutri-Score, and evidence."
- Text2Cypher: "Which brands have at least two products with Nutri-Score A or B and no declared palm-oil marker in this snapshot?"
- Audit: "How do we know that this product declares this allergen in the snapshot, and when was that information captured?"
- Evidence Agent: "Find a chocolate-spread candidate, then prove one declared-allergen claim and return its deterministic Evidence Record."
- Contextual follow-up: "Show the source", "Which of them declare milk?", or "Compare the first two."
- Recipe Studio: compose the chocolate apple oat bowl, exclude declared soybeans, inspect the selected alternatives, then show per-serving nutrition, preparation path and sources.
The repository's original code and documentation will use the MIT License. Open Food Facts database content is available under the Open Database License; individual contents use the Database Contents License. Dataset attribution and snapshot metadata must remain attached to redistributed fixtures.