The Korean document (README.ko.md) is the primary reference. This English version is kept as a secondary summary.
HydraLLM is a context-aware gateway that routes requests across Gemini / Groq / Cerebras with per-provider circuit breakers, random key rotation with quota-aware cooldowns, and real-time web enrichment, all behind an OpenAI-compatible API built on a strict Clean Architecture (Domain then Services then Adapters then API).
- Version:
1.3.0(pyproject.toml) - Python:
3.10+ - Entry point:
python main.py - Unified UI:
http://localhost:8000/ui - OpenAI-compatible endpoint:
POST /v1/chat/completions
.
├── main.py # Uvicorn entry point (supports --debug, --port)
├── src/
│ ├── app.py # FastAPI factory, lifespan, static UI mount
│ ├── adapters/providers/ # gemini, openai_compat (Groq/Ollama), cerebras, local_cli
│ ├── api/v1/ # endpoints.py, dependencies.py
│ ├── core/ # config, container, exceptions, logging
│ ├── domain/ # enums, interfaces, schemas, models
│ ├── services/ # analyzer, gateway, key_manager, session_manager,
│ │ # scraper, compressor, web_context_service,
│ │ # admin_service, metrics_service, observability,
│ │ # session_orchestrator, context_manager
│ └── utils/ # ulid helpers
├── tests/
│ ├── unit/ # analyzer, key_manager, adapters, ulid, stability
│ ├── integration/ # gateway failover, auto-models, provider validation
│ └── api/ # FastAPI endpoint contract tests
├── static/ # Unified SPA (Playground + Dashboard)
├── scripts/ # validate_flow.py (end-to-end routing validator)
├── pyproject.toml # Poetry, ruff, mypy, pytest configuration
└── .env # Provider keys and runtime settings (gitignored)
- Intelligent Routing —
services/analyzer.py::ContextAnalyzerpicks a provider/model based on token count, multimodality, detected web intent, explicit model hints (provider/model), and available key tiers. - Circuit Breaker + Cloud Failover —
services/gateway.pywraps every provider with aCircuitBreaker(5-failure threshold, 60s recovery) and retries across thePROVIDER_PRIORITYchain (Gemini then Groq then Cerebras). - Final Local Fallback — When all cloud providers are exhausted, the Gateway routes to Ollama via
OpenAICompatAdapterpointing atOLLAMA_BASE_URL. - Key Rotation with Cooldowns —
services/key_manager.py::KeyManagermaintains per-provider pools, selects keys randomly from the active set, and applies longer cooldowns for quota (1h) or forbidden/403 (24h) errors. - Web Enrichment —
services/web_context_service.py+services/scraper.py::WebScraper(Playwright + Scrapling) fetch explicit URLs or perform scraping when web intent is detected, with a 24-hour SQLite cache. When the gateway successfully injects a web context block intorequest.messages[-2], it emits a stdout INFO lineWeb context injected: N chars into request.messages[-2] (session=...)so operators can confirm enrichment without reading the SQLite event store. - Context Compression —
services/compressor.py::ContextCompressoruses LLMLingua-2 (optionalcompressionextra) to prune long histories. - Session Persistence —
services/session_manager.py::SessionManagerstores messages and parts in SQLite (WAL), supports forking and compaction thresholds, and holds runtime settings. - Unified Admin UI — Single SPA at
/uicombining playground, dashboard, key status, and model catalogue; all fetches use absolute URLs for proxy stability. - OpenAI API Compatibility —
/v1/chat/completionsincluding streaming SSE (chat.completion.chunk+[DONE]). - Incremental Web-Intent Keyword Learning —
services/keyword_store.py::KeywordStorepersists per-language (ko,en) keywords to JSON files (data/web_keywords.{lang}.json);services/intent_classifier.py::IntentClassifiersubstring-matches them before falling back to embedding similarity.scripts/validate_flow.pyautomatically registers false-negative queries to/v1/admin/intent/keywords/learnto grow the lexicon.
All endpoints are mounted under /v1 via src/api/v1/endpoints.py.
| Method | Path | Purpose |
|---|---|---|
POST |
/v1/chat/completions |
Primary chat entry (streaming supported) |
GET |
/v1/models |
List all discovered models |
GET |
/v1/admin/sessions |
List persisted sessions |
POST |
/v1/admin/sessions/new |
Create a new session |
DELETE |
/v1/admin/sessions/{session_id} |
Delete a session |
GET |
/v1/admin/logs?limit=50 |
Recent system logs |
GET |
/v1/admin/stats |
Aggregate usage + health stats |
GET |
/v1/admin/dashboard |
Stats + recent logs for the UI |
GET |
/v1/admin/status |
Live provider/agent status |
POST |
/v1/admin/refresh-models |
Re-run provider model discovery |
POST |
/v1/admin/probe |
Probe all keys for health |
POST |
/v1/admin/keys |
Add runtime keys (see Known Issues) |
GET |
/v1/admin/onboarding |
Onboarding status + available models |
POST |
/v1/admin/onboarding |
Save onboarding choices |
GET |
/v1/admin/intent/keywords |
List web-intent keywords per language |
POST |
/v1/admin/intent/keywords |
{lang,keywords[]} manual keyword registration |
POST |
/v1/admin/intent/keywords/learn |
{query} learn keywords from a false-negative query (LLM extraction + regex fallback) |
Plus the root and UI routes:
| Method | Path | Purpose |
|---|---|---|
GET |
/ |
Service banner (links to /docs, /openapi.json, /ui) |
GET |
/ui |
Unified admin SPA (static/index.html) |
GET |
/ui/static/* |
Static assets |
python3.10 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
⚠️ pydantic v2 required: this project needspydantic>=2.5andpydantic-settings>=2.1. If pydantic v1 is present in~/.local, you will seeModuleNotFoundError: No module named 'pydantic._internal'. Use a venv, or upgrade withpip install --upgrade 'pydantic>=2.5'.
# (A) pip (PEP 517 / pyproject.toml)
pip install . # runtime only
pip install '.[dev]' # + pytest / pytest-asyncio / pytest-cov / mypy / ruff
pip install '.[dev,compression]' # + llmlingua (context compression)
pip install '.[compression]' # compression only
# (B) Poetry
poetry install # runtime + dev group (default)
poetry install -E compression # + context compression
[tool.poetry.extras]declares adevextra, so you can install the full test / lint / type-check toolchain frompyproject.tomlalone viapip install '.[dev]'— no Poetry required. Quote the argument because many shells treat[...]as a glob.
The web scraper (services/scraper.py) drives Chromium, so a one-time download is required.
python -m playwright install chromiumcp .env.example .env
# edit .env to set GEMINI_KEYS, GROQ_KEYS, CEREBRAS_KEYS, etc.python main.py # starts on port 8000
curl http://127.0.0.1:8000/ # {"status":"online", ...}# Run server (defaults to port 8000)
python main.py
python main.py --debug --port 8001
# Tests (current baseline: 106 passed)
pytest # full suite
pytest -m unit # unit tests only
pytest -m integration # integration tests only
pytest tests/unit/test_analyzer.py::test_auto_routing # single test
# Code quality
ruff check .
ruff check --fix .
mypy src/
# Reproducible isolated full-suite run (clones source to a temp dir,
# fresh venv, pip install '.[dev]', pytest, and parses the pass/fail count).
# If .env is missing, the script falls back to .env.example automatically.
EXPECTED_TESTS=106 scripts/isolated_test.sh --clean # Linux / macOS / Git Bash
powershell -ExecutionPolicy Bypass -File scripts/isolated_test.ps1 # Korean Windows (cp949-safe)src/core/logging.pypinsRotatingFileHandler(..., encoding="utf-8")and reconfiguressys.stdoutto UTF-8.src/services/session_manager.py::_get_project_idpassesencoding="utf-8", errors="replace"tosubprocess.run, so repo paths containing Korean characters no longer raiseUnicodeDecodeError.scripts/isolated_test.shexportsPYTHONUTF8=1 / PYTHONIOENCODING=utf-8 / LC_ALL=C.UTF-8before running pip / pytest. It also falls back tocp -awhenrsyncis missing (typical Git Bash install) and detects.venv/Scripts/activatevs.venv/bin/activatefor Windows-native venvs.- The PowerShell counterpart runs
chcp 65001+[Console]::OutputEncoding = UTF8, copies viarobocopy, and persists the pytest log through[System.IO.File]::WriteAllLines(..., UTF8Encoding(false))instead ofTee-Objectto avoid the UTF-16 LE default in Windows PowerShell 5.1. Preferisolated_test.ps1on native Korean Windows.
Settings are loaded from .env via pydantic-settings (src/core/config.py::Settings).
Key variables:
- Keys (comma-separated pools) —
GEMINI_KEYS,GROQ_KEYS,CEREBRAS_KEYS - Priority —
PROVIDER_PRIORITY=gemini,groq,cerebras,ollama,opencode,openclaw - Routing defaults —
DEFAULT_FREE_MODEL,DEFAULT_PREMIUM_MODEL,MAX_TOKENS_FAST_MODEL - Local agents —
OLLAMA_BASE_URL,OPENCODE_BASE_URL,OPENCLAW_BASE_URL - Features —
ENABLE_CONTEXT_COMPRESSION,ENABLE_AUTO_WEB_FETCH,WEB_CACHE_TTL_HOURS - Admin —
ADMIN_API_KEY(optional; unset disables admin auth) - Web-intent keyword store —
DATA_DIR(defaultdata/),KEYWORD_EXTRACTION_MODEL(Ollama small LLM name; regex fallback only when unset)
See .env.example for the full list with example values. .env is listed in .gitignore and must not be committed.
.envstores provider keys in plaintext and is local-only. It is already covered by.gitignore, so keys are not pushed to the remote repository.- Treat
.envthe same as any other secret store on disk: restrict file permissions (chmod 600 .env), never share the raw file, and never paste the contents into chat/AI assistants, issue trackers, screenshots, or pair-programming tools. Once a key leaves the machine — including into an LLM session transcript — it must be considered compromised. - If a key is ever read aloud, pasted, logged, or committed by mistake, revoke and rotate it immediately at the provider console (Gemini / Groq / Cerebras), then replace the value in
.envand restart the server. - For shared environments, prefer a proper secret manager (OS keychain, Vault, 1Password CLI, cloud KMS) and populate
.envat process start; do not check secrets into any dotfile that syncs to another host.
- pytest (full suite): 106/106 passed in both source and isolated venv (
scripts/isolated_test.sh --clean). Integration testtest_auto_models_functionalitymay still fail when the local Ollama instance returns an embedding-only model for chat (seeTROUBLESHOOTING.mdsection 12). - mypy: 0 errors (
mypy src/). Pydantic v2 mypy plugin enabled. - ruff (src/): 0 errors. Test files retain
E402import-order violations due tosys.pathmanipulation (by design). - Version:
pyproject.tomlandsrc/app.pyboth declare1.3.0. - Web context vs key exhaustion: when all cloud key pools are in cooldown (forbidden=24h, quota=1h), the gateway routes to the local Ollama fallback. The web search/scrape result IS still injected into
request.messages(look for the newWeb context injected: N charsINFO log ingateway.log); the limiting factor in that scenario is the local model's capacity, not the enrichment pipeline. UsePOST /v1/admin/probeto re-validate keys, orPOST /v1/admin/keysto inject fresh keys at runtime.
Last Updated: 2026-04-27