Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
b4de0a0
fix(adopt): match served_model_name on adoption; unschedulable uses r…
milk333445 Jul 4, 2026
cf609c8
feat(engines): converge engine knowledge — fail-loud phantom + /api/e…
milk333445 Jul 4, 2026
6d12939
docs(engines): add-a-new-engine guide + TensorRT-LLM research/validat…
milk333445 Jul 4, 2026
68da159
feat(trtllm): add TensorRT-LLM as the 4th engine (Phase 1: bring-your…
milk333445 Jul 4, 2026
04f43ff
fix(trtllm): free the port before (re)start so a crashed engine recov…
milk333445 Jul 4, 2026
b4a4e74
docs(trtllm): document the crash/kill port-teardown restart fix (§13b)
milk333445 Jul 4, 2026
7543d3a
feat(trtllm): scrape /prometheus/metrics + Grafana dashboard embedded…
milk333445 Jul 4, 2026
b44210b
feat(trtllm): reject on missing runtime (available() hook) + Grafana …
milk333445 Jul 4, 2026
01bc4e2
docs(config): ship a 4-engine example config (vLLM + SGLang + llama.c…
milk333445 Jul 4, 2026
5659620
feat(trtllm): dual backend — engine_dir -> tensorrt, model_tag -> pyt…
milk333445 Jul 4, 2026
fee0e92
feat(trtllm): convert-to-TRT backend — build engines from HF via trtl…
milk333445 Jul 4, 2026
93f9fa1
feat(trtllm): convert-to-TRT UI in the model library
milk333445 Jul 4, 2026
97b689b
feat(trtllm): cache TRT engines in the HF cache + categorize the Mode…
milk333445 Jul 5, 2026
18415ed
feat(trtllm): pin /api/trt/* to the trtllm backend in the mixed nginx
milk333445 Jul 5, 2026
d399885
feat(trtllm): TRT engine picker in the Add/Edit Model dialog
milk333445 Jul 5, 2026
5117cb6
feat(models): pick a downloaded model in the Add Model dialog
milk333445 Jul 5, 2026
d544458
fix(models): clear observed row on delete so the model leaves the das…
milk333445 Jul 5, 2026
373273c
docs: document TensorRT-LLM engine + Convert to TRT across README and…
milk333445 Jul 5, 2026
e33628a
chore(mixed): include trtllm in the default make up-mixed engines
milk333445 Jul 5, 2026
57313ac
feat(ui): UI/UX polish pass across the dashboard
milk333445 Jul 5, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 9 additions & 4 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -9,16 +9,18 @@ COMPOSE := docker compose -f deploy/docker-compose.yaml
COMPOSE_MIXED := docker compose -f deploy/docker-compose.mixed.yaml

# Which engine backends the mixed stack runs (comma-separated subset of
# vllm,sglang,llamacpp). Override on the CLI: `make up-mixed ENGINES=vllm,sglang`.
# vllm,sglang,llamacpp,trtllm). Override on the CLI: `make up-mixed ENGINES=vllm,sglang`.
# The FIRST engine listed also serves the dashboard /api (all backends share Postgres,
# so any can). Compose profiles gate the rest, so unlisted backends never start.
ENGINES ?= vllm,sglang,llamacpp
# NOTE: trtllm pulls a large (~60GB) TensorRT-LLM image; drop it from ENGINES if you
# don't need it (e.g. `make up-mixed ENGINES=vllm,sglang,llamacpp`).
ENGINES ?= vllm,sglang,llamacpp,trtllm
DASHBOARD_BACKEND := $(shell echo "$(ENGINES)" | cut -d, -f1)-backend
MIXED_ENV := COMPOSE_PROFILES=$(ENGINES) DASHBOARD_BACKEND=$(DASHBOARD_BACKEND)
# Engines NOT selected this run. `docker compose up` (even with --remove-orphans)
# leaves a profiled-but-deselected service running, so we stop+remove them explicitly
# to make switching engine sets clean. Computed as ALL − ENGINES.
_ALL_ENGINES := vllm sglang llamacpp
_ALL_ENGINES := vllm sglang llamacpp trtllm
_comma := ,
_space := $(empty) $(empty)
_SELECTED := $(subst $(_comma),$(_space),$(ENGINES))
Expand All @@ -28,7 +30,7 @@ _DESELECTED_SVCS := $(addsuffix -backend,$(_DESELECTED))
.PHONY: help test test-backend test-router test-schema \
dev-backend dev-frontend build-frontend install-frontend \
up down logs ps build up-mixed down-mixed logs-mixed \
up-vllm up-sglang up-llamacpp
up-vllm up-sglang up-llamacpp up-trtllm

help:
@echo "Targets:"
Expand Down Expand Up @@ -89,6 +91,9 @@ up-sglang:
up-llamacpp:
$(MAKE) up-mixed ENGINES=llamacpp

up-trtllm:
$(MAKE) up-mixed ENGINES=trtllm

# down removes the whole project (all profiles) regardless of the current selection.
down-mixed:
COMPOSE_PROFILES=vllm,sglang,llamacpp $(COMPOSE_MIXED) down
Expand Down
30 changes: 22 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,8 @@ becomes a routable model; the router load-balances across instances; and a bundl
## Highlights

- **One router controls the whole fleet** — a single OpenAI- & Anthropic-compatible origin fronts every model. Route by the `model` field across `/v1/chat/completions`, `/v1/messages`, `/v1/embeddings`, `/v1/rerank`, `/v1/score`, `/tokenize` and more; the router resolves the group and load-balances its instances, so clients never address an instance directly.
- **Three inference engines, one control plane — vLLM, SGLang & llama.cpp** — choose the engine per model (the *Add Model* dialog has an engine selector); an engine-aware scheduler places each model on a backend that can run it, and the same router / dashboard / monitoring front them all. Run a **vLLM-only** stack (`make up`) or a **mixed** fleet (`make up-mixed`) that adds SGLang and a llama.cpp (GGUF / CPU-offload) backend. See [docs/mixed-engine-deployment.md](docs/mixed-engine-deployment.md).
- **Four inference engines, one control plane — vLLM, SGLang, llama.cpp & TensorRT-LLM** — choose the engine per model (the *Add Model* dialog has an engine selector); an engine-aware scheduler places each model on a backend that can run it, and the same router / dashboard / monitoring front them all. Run a **vLLM-only** stack (`make up`) or a **mixed** fleet (`make up-mixed`) that adds SGLang, llama.cpp (GGUF / CPU-offload) and TensorRT-LLM backends. See [docs/mixed-engine-deployment.md](docs/mixed-engine-deployment.md).
- **TensorRT-LLM, two ways** — a `trtllm` group runs a **pre-built TRT engine** (`engine_dir` → `--backend tensorrt`, lowest TTFT) or **HF weights directly** (no `engine_dir` → `--backend pytorch`, no build step, any TRT-LLM-supported arch); monitoring is identical either way. Build engines from the UI: the **Model Library**'s *Convert to TRT* action runs `trtllm-bench` (one architecture-agnostic path, not per-model scripts) to compile a cached HF model into an engine, then pick it in *Add Model*. See [docs/trtllm-launcher-impl-design_zh-TW.md](docs/trtllm-launcher-impl-design_zh-TW.md).
- **Add a model by pasting `vllm serve …`** — parsed into a form and layered on as a dynamic overlay; the router hot-reloads, no `config.yaml` edits.
- **Lifecycle + self-healing** — per-instance state machine (`stopped → starting → ready → sleeping → failed`), VRAM pre-flight guard, GPU auto-placement, crash auto-restart with backoff.
- **Autoscaling with a warm-standby tier** — per group, keep `min_ready` replicas warm and scale up on queue depth (wake first, else cold-start) to `max_ready`; fold idle replicas back down `ready → sleep → stop`. vLLM **sleep mode** (level-1) frees a replica's VRAM but wakes in seconds, so scaling down needn't mean a minute-long cold start. Set it from config.yaml or the dashboard; a live Grafana dashboard + alerts are bundled.
Expand All @@ -41,7 +42,7 @@ becomes a routable model; the router load-balances across instances; and a bundl
- **Lifecycle alerting** — discrete model events (crash, restart-budget exhausted, recovered) pushed to Slack / Discord / a generic webhook, with per-sink severity floors and per-model cooldown; configured via env or the admin **Notifications** page (one-click test). Complements Grafana's metric alerts.
- **Playground** — OpenAI-compatible chat (streaming) / completions / embeddings / reranking, with reasoning display.
- **Benchmark & evaluate** — evalscope load tests (concurrency, arrival-rate, SLA auto-tune) plus 30+ accuracy datasets with LLM-as-judge.
- **Libraries** — browse / pre-download HF model weights & datasets from the UI; tool-calling parser helper; LoRA support.
- **Libraries** — browse / pre-download HF model weights & datasets from the UI; the Model Library groups cached weights into **Models / GGUF / TensorRT-LLM engines**, builds TRT engines from a cached model, and *Add Model* lets you pick a downloaded model (and a built TRT engine) instead of typing paths; tool-calling parser helper; LoRA support (incl. PEFT→GGUF conversion).
- **Multi-user & audit** — role-based control (`viewer`/`operator`/`admin`) via named operator credentials, with a redacted **audit log** of every change; plus mint/revoke API keys with per-key usage attribution, rate limits and **token quotas** (total / daily / monthly). The env admin token and open local-dev still work unchanged.
- **Config versioning & backup** — the dynamic-model overlay (where every runtime change lives) is snapshotted on each mutation; export it as a portable file, import to restore, and roll back to any past version with a side-by-side diff — all from the admin **Config Versions** page. `config.yaml` is never rewritten.

Expand All @@ -63,8 +64,8 @@ make up # build + start the whole stack
**Two deployment modes:**

- **`make up`** — the default **vLLM-only** stack.
- **`make up-mixed`** — a **multi-engine** fleet: vLLM, SGLang and llama.cpp backends sharing one Postgres, router, dashboard and Grafana. Add a model from *Add Model → engine: `sglang` / `llamacpp`* and it is auto-placed on a backend that can run it. `make down-mixed` / `make logs-mixed` manage it. See [docs/mixed-engine-deployment.md](docs/mixed-engine-deployment.md).
- **Pick which engines run**: `make up-mixed ENGINES=vllm,sglang` (any subset), or the shortcuts `make up-vllm` / `up-sglang` / `up-llamacpp`. Unlisted engines never start (their VRAM stays free); the first engine listed serves the dashboard API, and switching sets stops the de-selected backends automatically.
- **`make up-mixed`** — a **multi-engine** fleet: vLLM, SGLang, llama.cpp and TensorRT-LLM backends sharing one Postgres, router, dashboard and Grafana (all four by default; drop any via `ENGINES=…`). Add a model from *Add Model → engine: `sglang` / `llamacpp` / `trtllm`* and it is auto-placed on a backend that can run it. `make down-mixed` / `make logs-mixed` manage it. See [docs/mixed-engine-deployment.md](docs/mixed-engine-deployment.md).
- **Pick which engines run**: `make up-mixed ENGINES=vllm,sglang,llamacpp,trtllm` (any subset), or the shortcuts `make up-vllm` / `up-sglang` / `up-llamacpp` / `up-trtllm`. Unlisted engines never start (their VRAM stays free); the first engine listed serves the dashboard API, and switching sets stops the de-selected backends automatically. The dashboard shows the **whole fleet regardless of which backend serves the API** (nginx additionally pins the TRT engine-build endpoints to the TensorRT-LLM backend, so *Convert to TRT* works even when another engine serves the dashboard).

```bash
curl http://localhost:8887/v1/models # router: configured model groups
Expand Down Expand Up @@ -140,12 +141,13 @@ The **router only routes** — the **backend owns model lifecycle**. The fronten
backend, and Grafana sit behind nginx on a single origin; backend, router, and Prometheus
share one network namespace so the spawned vLLM instances are reachable on `localhost`.

### Mixed vLLM + SGLang + llama.cpp (`make up-mixed`)
### Mixed vLLM + SGLang + llama.cpp + TensorRT-LLM (`make up-mixed`)

Each engine runs as its own backend container (they can't share a netns), sharing one
Postgres (scheduling / desired intent), one router, one dashboard and one monitoring stack.
Each backend publishes its ready instances as **routable addresses** to a shared file_sd that
Prometheus scrapes. See [docs/mixed-engine-deployment.md](docs/mixed-engine-deployment.md).
Prometheus scrapes (TensorRT-LLM is scraped at `/prometheus/metrics` and exposes `trtllm_*`).
See [docs/mixed-engine-deployment.md](docs/mixed-engine-deployment.md).

```mermaid
flowchart LR
Expand All @@ -168,37 +170,49 @@ flowchart LR
BL["<b>backend</b> · :5073<br/>NODE_ENGINES=llamacpp"]
LINS["llama-server instances"]
end
subgraph tbe["TensorRT-LLM backend (engine-trtllm.Dockerfile)"]
BT["<b>backend</b> · :5074<br/>NODE_ENGINES=trtllm"]
TINS["trtllm-serve instances"]
end

Client --> FE
FE -->|/api| BV
FE -->|/api/trt| BT
FE -->|/v1| RT
FE -->|/grafana| GF
BV -->|launch| VINS
BS -->|launch| SINS
BL -->|launch| LINS
BT -->|launch / build engine| TINS
RT -->|route| VINS
RT -->|route| SINS
RT -->|route| LINS
RT -->|route| TINS
BV <-->|leader/schedule| PG
BS <-->|converge desired| PG
BL <-->|converge desired| PG
BT <-->|converge desired| PG
PR -->|scrape| VINS
PR -->|scrape| SINS
PR -->|scrape| LINS
PR -->|scrape /prometheus/metrics| TINS
GF -->|query| PR
```

The leader's **engine-aware scheduler** places each model on a backend that can run its engine;
a control action landing on the wrong node is deferred to the owning one. SGLang serves
OpenMetrics, so Prometheus stores its metrics as `sglang_*` (underscore), while vLLM and
llama.cpp keep colons (`vllm:*`, `llamacpp:*`).
llama.cpp keep colons (`vllm:*`, `llamacpp:*`); TensorRT-LLM exposes `trtllm_*` at
`/prometheus/metrics`. Each engine has its own bundled Grafana dashboard and alert rules.

## Documentation

| Topic | |
|---|---|
| Deployment & topology | [docs/deployment.md](docs/deployment.md) |
| Mixed-engine (vLLM + SGLang) | [docs/mixed-engine-deployment.md](docs/mixed-engine-deployment.md) |
| Mixed-engine (vLLM + SGLang + llama.cpp) | [docs/mixed-engine-deployment.md](docs/mixed-engine-deployment.md) |
| TensorRT-LLM engine (dual backend + Convert to TRT) | [docs/trtllm-launcher-impl-design_zh-TW.md](docs/trtllm-launcher-impl-design_zh-TW.md) |
| Adding a new engine | [docs/adding-a-new-engine_zh-TW.md](docs/adding-a-new-engine_zh-TW.md) |
| Configuration (`config.yaml`) | [docs/configuration.md](docs/configuration.md) |
| Features in depth | [docs/features.md](docs/features.md) |
| Monitoring (Prometheus + Grafana) | [docs/monitoring.md](docs/monitoring.md) |
Expand Down
13 changes: 8 additions & 5 deletions README_zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,8 @@
## 功能亮點

- **一個 router 掌控整個集群** — 單一 OpenAI 與 Anthropic 相容入口統管所有模型。以 `model` 欄位路由 `/v1/chat/completions`、`/v1/messages`、`/v1/embeddings`、`/v1/rerank`、`/v1/score`、`/tokenize` 等端點;router 自動解析群組並在實例間負載平衡,客戶端永遠不直接連到單一實例。
- **三種推理引擎、同一個控制平面 — vLLM、SGLang 與 llama.cpp** — 每顆模型可各自選引擎(*新增模型*對話框有引擎選擇器);engine-aware 排程器把每顆模型擺到「跑得動它」的 backend 上,並由同一個 router/控制台/監控統一前置。可只跑 **vLLM**(`make up`),或跑**混合**集群(`make up-mixed`)加上 SGLang 與 llama.cpp(GGUF/CPU offload)backend。見 [docs/mixed-engine-deployment_zh-CN.md](docs/mixed-engine-deployment_zh-CN.md)。
- **四種推理引擎、同一個控制平面 — vLLM、SGLang、llama.cpp 與 TensorRT-LLM** — 每顆模型可各自選引擎(*新增模型*對話框有引擎選擇器);engine-aware 排程器把每顆模型擺到「跑得動它」的 backend 上,並由同一個 router/控制台/監控統一前置。可只跑 **vLLM**(`make up`),或跑**混合**集群(`make up-mixed`)加上 SGLang、llama.cpp(GGUF/CPU offload)與 TensorRT-LLM backend。見 [docs/mixed-engine-deployment_zh-CN.md](docs/mixed-engine-deployment_zh-CN.md)。
- **TensorRT-LLM,兩種跑法** — `trtllm` group 可跑**預先編譯的 TRT engine**(`engine_dir` → `--backend tensorrt`,最低 TTFT),或**直接吃 HF 權重**(不給 `engine_dir` → `--backend pytorch`,免編譯、支援 TRT-LLM 所有架構);兩者監控完全一致。可從 UI 建 engine:**模型庫**的 *轉成 TRT* 用 `trtllm-bench`(一條 arch-agnostic 路徑,非逐模型腳本)把已快取的 HF 模型編成 engine,再到 *新增模型* 挑它。見 [docs/trtllm-launcher-impl-design_zh-TW.md](docs/trtllm-launcher-impl-design_zh-TW.md)。
- **貼上 `vllm serve …` 即可新增模型** — 解析成表單、以動態 overlay 疊加;router 熱重載。
- **生命週期** — 每實例狀態機(`stopped → starting → ready → sleeping → failed`)、VRAM 預檢防呆、GPU 自動擺放、崩潰指數退避自動重啟。
- **自動擴縮(含暖待命層)** — 每群組保留 `min_ready` 暖機副本,依佇列深度擴容(優先喚醒、其次冷啟)到 `max_ready`;閒置時逐階縮回 `ready → sleep → stop`。vLLM **sleep mode**(level-1)釋放副本 VRAM 但秒級喚醒,所以縮容不必付出數分鐘冷啟代價。config.yaml 或控制台皆可設定,內建即時 Grafana 面板與告警。
Expand All @@ -42,7 +43,7 @@
- **生命週期告警** — 離散的模型事件(崩潰、退避用盡、復原)推到 Slack/Discord/通用 webhook,含每個 sink 自訂嚴重度門檻與 per-model 去重;用環境變數或 admin「通知」頁(含一鍵測試)設定。與 Grafana 指標告警互補。
- **Playground** — OpenAI 相容的 chat(串流)/completions/embeddings/reranking。
- **壓測與評測** — LLM 壓測(並發、到達率、SLA 自動調優)+ 30+ 個準確度資料集與 LLM-as-judge。
- **資料庫** — 在 UI 瀏覽/預下載 HF 權重與資料集;工具調用 parser 助手;LoRA 支援。
- **資料庫** — 在 UI 瀏覽/預下載 HF 權重與資料集;模型庫把已快取權重分成 **一般模型/GGUF/TensorRT-LLM 引擎** 三區,可從已快取模型建 TRT engine,*新增模型* 也能直接挑已下載的模型(與已建的 TRT engine)而非手打路徑;工具調用 parser 助手;LoRA 支援(含 PEFT→GGUF 轉換)。
- **多使用者與稽核** — 以具名 operator 憑證做角色控管(`viewer`/`operator`/`admin`),並有脫敏的**稽核日誌**記錄每次變更;另可發行/撤銷 API 金鑰,帶 per-key 用量歸屬、速率上限與 **token 額度**(總量/每日/每月)。env 管理員權杖與本機 dev 開放模式維持不變。
- **設定版本化與備份** — 動態模型 overlay(所有 runtime 改動所在)每次變更都會自動快照;可一鍵匯出成可攜檔備份、匯入還原,也能在 admin「設定版本」頁看歷史、並排 diff 與一鍵回滾到任一版。`config.yaml` 永遠不會被改寫。

Expand All @@ -63,8 +64,8 @@ make up # 建置並啟動整套服務
**兩種啟動方式:**

- **`make up`** — 預設的**純 vLLM** 集群。
- **`make up-mixed`** — **多引擎**混合集群:vLLM、SGLang 與 llama.cpp backend 共用同一顆 Postgres、router、控制台與 Grafana。從 *新增模型 → 引擎:`sglang` / `llamacpp`* 新增的模型會自動擺到跑得動它的 backend。對應 `make down-mixed`/`make logs-mixed`。見 [docs/mixed-engine-deployment_zh-CN.md](docs/mixed-engine-deployment_zh-CN.md)。
- **可選要跑哪些引擎**:`make up-mixed ENGINES=vllm,sglang`(任意子集),或捷徑 `make up-vllm` / `up-sglang` / `up-llamacpp`。沒列到的引擎不會起(VRAM 留著);第一個引擎負責控制台 API,切換引擎組合時會自動停掉沒選到的 backend。
- **`make up-mixed`** — **多引擎**混合集群:vLLM、SGLang、llama.cpp 與 TensorRT-LLM backend 共用同一顆 Postgres、router、控制台與 Grafana(預設四個全開;用 `ENGINES=…` 可拿掉不要的)。從 *新增模型 → 引擎:`sglang` / `llamacpp` / `trtllm`* 新增的模型會自動擺到跑得動它的 backend。對應 `make down-mixed`/`make logs-mixed`。見 [docs/mixed-engine-deployment_zh-CN.md](docs/mixed-engine-deployment_zh-CN.md)。
- **可選要跑哪些引擎**:`make up-mixed ENGINES=vllm,sglang,llamacpp,trtllm`(任意子集),或捷徑 `make up-vllm` / `up-sglang` / `up-llamacpp` / `up-trtllm`。沒列到的引擎不會起(VRAM 留著);第一個引擎負責控制台 API,切換引擎組合時會自動停掉沒選到的 backend。控制台**不論由哪個 backend 服務都會顯示整個 fleet**(nginx 另外把 TRT engine 建置端點固定路由到 TensorRT-LLM backend,所以即使別的引擎在服務控制台,*轉成 TRT* 仍可用)。

```bash
curl http://localhost:8887/v1/models # router:列出設定的模型群組
Expand Down Expand Up @@ -194,7 +195,9 @@ leader 的 **engine-aware 排程器**把每顆模型擺到「跑得動它引擎
| 主題 | |
|---|---|
| 部署與架構 | [docs/deployment_zh-CN.md](docs/deployment_zh-CN.md) |
| 混合引擎(vLLM + SGLang) | [docs/mixed-engine-deployment_zh-CN.md](docs/mixed-engine-deployment_zh-CN.md) |
| 混合引擎(vLLM + SGLang + llama.cpp) | [docs/mixed-engine-deployment_zh-CN.md](docs/mixed-engine-deployment_zh-CN.md) |
| TensorRT-LLM 引擎(雙 backend + 轉成 TRT) | [docs/trtllm-launcher-impl-design_zh-TW.md](docs/trtllm-launcher-impl-design_zh-TW.md) |
| 新增一個引擎 | [docs/adding-a-new-engine_zh-TW.md](docs/adding-a-new-engine_zh-TW.md) |
| 配置(`config.yaml`) | [docs/configuration_zh-CN.md](docs/configuration_zh-CN.md) |
| 功能特色(詳細) | [docs/features_zh-CN.md](docs/features_zh-CN.md) |
| 監控(Prometheus + Grafana) | [docs/monitoring_zh-CN.md](docs/monitoring_zh-CN.md) |
Expand Down
22 changes: 22 additions & 0 deletions apps/backend/app/api/engines.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
"""The registered inference engines and their capabilities.

Single source of engine knowledge for the dashboard: the frontend reads this to
build the engine picker and to gate form fields on `capabilities` /
`inapplicable_keys`, instead of hardcoding engine names and re-deriving what each
engine can do (which drifts from the backend's launcher CAP_* flags). See
docs/multi-engine-extensibility-review_zh-TW.md §5 (P1).
"""
from fastapi import APIRouter, Depends

from app.api.deps import get_manager
from app.llmops.manager import ModelManager

router = APIRouter(prefix="/engines", tags=["engines"])


@router.get("")
async def list_engines(manager: ModelManager = Depends(get_manager)):
"""The registered LLM engines, each with:
name, capabilities (CAP_* strings), lora_endpoint_prefix, metric_prefix,
inapplicable_keys (greyed-out model_config keys), paste_example."""
return {"engines": manager.engine_catalogue()}
Loading
Loading