Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -69,3 +69,7 @@ logs/
MANIFEST
*.tar.gz
*.whl

# Rust
target/
rust/tokenizer/target/
187 changes: 180 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -125,9 +125,23 @@ pip install -e ".[dev]"
# Single benchmark
cachepilot bench --policy perc --workload mixed --requests 1000

# Same benchmark with FP8 KV tier + Prometheus export
cachepilot bench --policy perc --workload mixed --requests 1000 --kv-tier fp8 \
--prometheus-out results/cachepilot.prom --snapshots-out results/cachepilot.json

# Side-by-side comparison with traffic spike
cachepilot compare --workload mixed --requests 2000 --spike 500

# Train the admission controller with policy gradient
cachepilot rl-admission --workload mixed --requests 400 --episodes 12 --kv-tier fp8

# Emit a Grafana dashboard JSON wired to the exported Prometheus metric names
cachepilot grafana-dashboard --out docs/cachepilot_grafana.json

# Profile real token usage from Hugging Face or a downloaded Kaggle export
cachepilot profile-dataset --preset oasst1 --limit 1000 --out results/oasst1_tokens.json
cachepilot profile-dataset --path data/kaggle/chatbot_conversations.csv --out results/kaggle_tokens.json

# From YAML benchmark config
python scripts/run_bench.py benchmarks/mixed_spike.yaml --out results/mixed.json
python scripts/plot_results.py results/mixed.json --out results/mixed.png
Expand Down Expand Up @@ -158,6 +172,15 @@ evictor.record_token(block_id)

Measured improvement: **79.3% reduction in expected KV recompute cost** on heterogeneous production-like traffic.

For a real smoke path against actual vLLM installs, the repo now includes:

```bash
pytest tests/test_vllm_integration.py -m integration
```

It runs `gpt2` by default when `vllm` is installed and can target LLaMA-2 by
setting `CACHEPILOT_VLLM_LLAMA_MODEL` to a local path or accessible model ID.

---

## CUDA Kernels
Expand All @@ -183,6 +206,16 @@ nvcc -O3 -arch=sm_90 -shared -o libkvquant.so src/cuda/kv_quant.cu

Triton versions (no nvcc required) in `src/cachepilot/kernels/` with automatic NumPy fallback for CPU environments.

There is now also an opt-in native build path through `setup.py` + pybind11:

```bash
pip install -e ".[native]"
CACHEPILOT_BUILD_CUDA=1 pip install -e .
```

The current checkout machine did not have `nvcc`, so the build path is
implemented but was not compiled here.

---

## INT8 Capacity Analysis
Expand Down Expand Up @@ -216,22 +249,162 @@ Calibrated to public datasets:
## Tests

```bash
pytest # 47 tests, all pass
pytest # 50 tests pass locally, 1 optional vLLM smoke test skipped without vLLM
pytest tests/test_eviction.py -v # PERC theoretical properties
pytest tests/test_engine.py -v # eviction cost comparisons
pytest tests/test_kernels.py -v # INT8 quantization bounds + vLLM evictor
cargo test --manifest-path rust/tokenizer/Cargo.toml
```

## Native Tokenizer

There is now a built-in Rust tokenizer at `rust/tokenizer/` for fast prompt
length estimation. Python falls back automatically to
`src/cachepilot/tokenizer.py` when the native binary is absent.

The tokenizer is now boundary-aware for:
- punctuation and whitespace
- camelCase and PascalCase splits
- digit/alpha transitions
- denser non-ASCII text

This keeps the estimator cheap while behaving less like a flat
`chars / constant` rule on code and mixed-format prompts.

```bash
cargo build --release --manifest-path rust/tokenizer/Cargo.toml
```

## Model Comparison

To compare your own model against baselines on identical prompts with vLLM:

```bash
cachepilot compare-models \
--candidate path/to/your-model \
--baseline gpt2 \
--prompts prompts.txt \
--out results/model_compare.json
```

The output highlights where the candidate wins on concrete serving metrics such
as end-to-end latency and generated tokens per second.

On a CUDA host you can also source prompts directly from Hugging Face datasets
or local dataset exports:

```bash
cachepilot compare-models \
--candidate /models/your-llm \
--baseline TinyLlama/TinyLlama-1.1B-Chat-v1.0 \
--preset alpaca \
--limit 64 \
--max-tokens 64 \
--tensor-parallel-size 1 \
--out results/model_compare_cuda.json
```

## End-to-End vLLM Benchmark

For a direct CUDA-host benchmark of plain `vllm` vs `vllm+PERC` on the same
local model weights:

```bash
cachepilot vllm-benchmark \
--model /models/your-llm \
--preset alpaca \
--limit 64 \
--compare-perc \
--max-tokens 64 \
--gpu-memory-utilization 0.85 \
--out results/vllm_benchmark.json
```

This command supports:
- local model paths on the benchmark host
- prompt files (`--prompts prompts.txt`)
- Hugging Face datasets (`--hf-dataset yahma/alpaca-cleaned`)
- local CSV / JSONL / Parquet exports (`--local-dataset data/chatbot.csv`)

Install the serving stack on the CUDA host with:

```bash
pip install -e .[bench]
```

To generate a standalone Hugging Face Jobs UV script for the same benchmark:

```bash
cachepilot render-hf-vllm-job \
--model TinyLlama/TinyLlama-1.1B-Chat-v1.0 \
--preset alpaca \
--limit 48 \
--max-tokens 96 \
--gpu-memory-utilization 0.55 \
--out scripts/hf_vllm_bench.py
```

## Hardware Scorecard

For a first-principles comparison of compute, bandwidth, and effective KV cache
capacity across current GPU tiers:

```bash
cachepilot hardware-scorecard \
--model llama3_8b \
--context-tokens 2048 \
--out results/hardware_scorecard.json
```

This uses a roofline-style bound:
- `tok/s <= memory_bandwidth / decode_kv_bytes`
- `tok/s <= peak_compute / decode_flops`
- actual ceiling = `min(compute_bound, bandwidth_bound)`

The derivation is documented in [docs/roofline_proofs.md](docs/roofline_proofs.md).

## Dataset Profiling

Use the dataset profiler to measure actual prompt, response, and total token
distributions before you train or benchmark:

```bash
# Hugging Face presets
cachepilot profile-dataset --preset oasst1
cachepilot profile-dataset --preset alpaca
cachepilot profile-dataset --preset sharegpt

# Direct Hugging Face repo ID
cachepilot profile-dataset --hf-dataset OpenAssistant/oasst1 --split train

# Kaggle export after download
cachepilot profile-dataset --path data/chatbot_conversations.csv
```

The profiler handles:
- flat instruction datasets such as Alpaca (`instruction`, `input`, `output`)
- ShareGPT-style conversation lists
- Kaggle-style turn tables with `conversation_id`, `role`, and `message`

---

## Resume Bullets

- Built CachePilot, a GPU memory orchestrator for multi-model LLM serving with a provably optimal KV cache eviction algorithm (PERC), reducing expected KV recompute cost by 25% in simulation and 79% in a vLLM-compatible evictor benchmark vs LRU.
- Designed PERC (Priority Eviction with Resumption Cost) and proved its optimality via fractional knapsack reduction — jointly models context length and per-session Poisson token arrival rate to minimize expected recompute cost when freeing VRAM.
- Implemented CUDA C++ kernels for PCIe-saturating KV block eviction (250 µs per 16 MB, 99% overlap with decode) and in-place INT8 KV quantization (50% VRAM reduction, <0.4% relative error bound), plus Triton equivalents and a 50-test suite.

---

## Next Steps

- [ ] Real vLLM integration test on GPT-2 / LLaMA-2
- [ ] CUDA kernel compilation via `setup.py` with pybind11
- [ ] RL fine-tuning for AdmissionPolicy with policy gradient
- [ ] Multi-GPU NVLink-aware placement
- [ ] Grafana dashboard with live telemetry export
- [ ] FP8 KV tier (4x compression vs FP16)
- [x] Real vLLM smoke test scaffold on GPT-2 / optional LLaMA-2
- [x] CUDA kernel compilation path via `setup.py` with pybind11
- [x] RL fine-tuning for AdmissionPolicy with policy gradient
- [x] Multi-GPU NVLink-aware placement primitive
- [x] Grafana dashboard export + Prometheus-style telemetry
- [x] FP8 KV tier support (2x compression vs FP16; 4x would require INT4/NVFP4)
- [x] End-to-end vLLM benchmark path for CUDA hosts and local model weights

---

Expand Down
78 changes: 78 additions & 0 deletions docs/roofline_proofs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,78 @@
# CachePilot Roofline Notes

This repo now includes a first-principles scorecard for decode efficiency.
The goal is not to predict exact production tok/s, but to establish hard upper
bounds from bandwidth, compute, and KV-cache capacity.

## 1. KV Bandwidth Law

For one cached context token, the KV footprint across all layers is:

`B_kv = 2 * n_layers * n_heads * head_dim * bytes_per_scalar`

The factor `2` is for `K` and `V`.

If the active decode context is `C` tokens, one output token requires reading:

`B_decode = C * B_kv`

If GPU memory bandwidth is `BW` bytes/s, then no implementation can sustain:

`tok/s > BW / B_decode`

This is a conservation law on bytes moved. It is independent of scheduler
details.

## 2. Attention FLOP Law

A simplified decode-attention cost for one output token is:

`F_decode ~= 4 * n_layers * n_heads * head_dim * C`

This covers the dominant `QK` and `AV` terms.

If GPU compute is `P` FLOP/s, then:

`tok/s > P / F_decode`

is impossible.

## 3. Roofline Bound

The realizable upper bound is the lower of the compute and bandwidth limits:

`tok/s_roofline = min(P / F_decode, BW / B_decode)`

The arithmetic intensity is:

`I = F_decode / B_decode`

Substituting the two formulas above:

`I ~= 2 / bytes_per_scalar`

So:
- FP16 KV gives `~1 FLOP/byte`
- FP8 / INT8 KV gives `~2 FLOP/byte`

Modern inference GPUs have ridge points far above this, which means decode is
typically bandwidth-bound, not math-bound.

## 4. Compression Law

Halving `bytes_per_scalar` from FP16 to FP8 or INT8:
- halves `B_kv`
- halves `B_decode`
- doubles cache-token capacity
- doubles the bandwidth-bound decode ceiling

That is why KV compression is a direct throughput and cache-headroom lever.

## 5. What This Proves

These formulas prove three useful things:
- KV cache compression gives a near-linear headroom improvement before
implementation overheads.
- Large-memory, high-bandwidth GPUs dominate long-context decode workloads.
- Policy work such as PERC matters because every prevented eviction avoids
recompute that would otherwise consume the same scarce bandwidth budget.
5 changes: 5 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,8 @@ dependencies = [
[project.optional-dependencies]
plot = ["matplotlib>=3.8"]
dev = ["pytest>=8.0", "pytest-cov>=5.0", "ruff>=0.4"]
native = ["pybind11>=2.11"]
bench = ["datasets>=2.19", "huggingface-hub>=0.24", "vllm>=0.6"]

[project.scripts]
cachepilot = "cachepilot.cli:main"
Expand All @@ -31,6 +33,9 @@ where = ["src"]
[tool.pytest.ini_options]
testpaths = ["tests"]
addopts = "-v --tb=short"
markers = [
"integration: tests that require optional external runtimes or model weights",
]

[tool.ruff]
line-length = 100
Expand Down
7 changes: 7 additions & 0 deletions rust/tokenizer/Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

7 changes: 7 additions & 0 deletions rust/tokenizer/Cargo.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
[package]
name = "cachepilot-tokenizer"
version = "0.1.0"
edition = "2021"

[dependencies]

Loading
Loading