diff --git a/benchmarks/glm53-vast-screening.md b/benchmarks/glm53-vast-screening.md new file mode 100644 index 000000000..b02de1348 --- /dev/null +++ b/benchmarks/glm53-vast-screening.md @@ -0,0 +1,106 @@ +# Full GLM-5.3: H200 Serverless qualification and preliminary screening + +Measured 2026-09-07. This is full GLM-5.3, not GLM-5.3-Flash. +The H200 NVL run verifies model readiness, authenticated Vast Serverless +routing, correct semantic output, and concurrency throughput. Idle shutdown +and cached restart were not tested before the instance was deleted. + +## H200 NVL Serverless qualification + +Vast Serverless provisioned one H200 NVL worker in Japan with 143,771 MiB +reported VRAM, 32 effective CPU cores, and a 520 GiB disk. FreeToken 0.1.2 +loaded `LibertAIDAI/GLM-5.3-NVFP4`, allocated 8,384 KV tokens, and captured +CUDA graphs for batch sizes 1, 2, 4, and 8. Both worker creation and routing +through the authenticated Serverless endpoint were observed. + +The semantic smoke prompt, `What is 17 + 25? Reply with only the number.`, +returned exactly `42` with finish reason `stop` in 10.402 seconds. + +The concurrency workload requested a concise Python function for merging two +sorted iterables. It used non-streaming end-to-end timing, including routing, +queueing, and prefill, with `reasoning_effort="low"` and a 128-token output +cap. + +| Offered concurrency | Running slots | Successful requests | Output tokens | Wall seconds | Aggregate output tok/s | Mean request seconds | p95 request seconds | +| --- | --- | --- | --- | --- | --- | --- | --- | +| 8 | 8 | 8/8 | 1,016 | 57.136 | 17.782 | 57.133 | 57.135 | +| 16 | 8 | 16/16 | 2,032 | 112.747 | 18.023 | 85.600 | 112.738 | + +All 24 requests succeeded and produced assistant output. Each reached the +fixed 128-token cap, so these runs measure throughput rather than completed +answer quality. At c16, eight requests ran while the other eight queued behind +the worker's eight active-request slots. The nearly unchanged aggregate rate +and nearly doubled wall time place practical saturation at c8 for this worker +configuration. + +The [qualification record](results/glm53-full-h200-serverless-qualification-20260907.json), +[raw c8/c16 results](results/glm53-full-h200-serverless-c8-c16-20260907.json), +and [semantic smoke response](results/glm53-full-h200-serverless-smoke-20260907.json) +preserve the machine-readable evidence. + +## Preliminary RTX PRO 6000 screening + +## Configuration + +- Checkpoint: `LibertAIDAI/GLM-5.3-NVFP4`, 282 indexed shards, + 464,822,689,680 bytes. File structure was validated, not independent hashes. +- Runtime: FreeToken 0.1.2, commit + `55120d9114eb18d90e4dfe2e1454f25f0098723e`. +- GPU: one RTX PRO 6000 Blackwell Max-Q, 97,887 MiB reported VRAM, + 300 W power cap; NVIDIA driver 595.84. +- Vast machine 104296, India; CPU quota 36.48 effective cores; + container memory limit 2,057,523,167,232 bytes. +- MoE offload, Triton NVFP4, DSA attention, tensor parallelism 1, + eight running request slots, max sequence length 32,768, 48 CPU threads, + memory fraction 0.95. Expert cache: 3,461 slots; KV capacity: 8,239 tokens. +- Direct loopback HTTP, not the Vast serverless routing path. +- OS: Ubuntu 24.04.3 LTS. Host CPU reports `Genuine Intel(R) CPU 0000%@` + across two sockets; no more specific processor SKU was exposed. + +The running process was inspected with this exact command line: + +```bash +/workspace/freetoken/.venv/bin/ft serve \ + --model /workspace/models/GLM-5.3-NVFP4 \ + --served-model-name glm-5.3-nvfp4 --host 127.0.0.1 --port 1919 \ + --moe-backend auto --moe-cpu-threads 48 --memory-ratio 0.95 \ + --max-seq-len-override 32768 --max-running-requests 8 \ + --disable-moe-prefill-overlap +``` + +### Initial results + +The prompt requested a concise Python function merging two sorted iterables, +with a request index appended. Each request used a 32-token output cap and +default reasoning. The model was loaded and a smoke completion had succeeded. + +| Offered concurrency | Running slots | Output tokens | Wall seconds | Aggregate output tok/s | Mean request seconds | Mean request output tok/s | +| --- | --- | --- | --- | --- | --- | --- | +| 8 | 8 | 248 | 62.991 | 3.937 | 62.983 | 0.492 | +| 16 | 8 | 496 | 78.767 | 6.297 | 58.434 | 0.603 | + +All 24 requests returned without transport errors, but all stopped at the +output cap while still reasoning. None supplied a final code answer. These +are output-token measurements, not successful coding-task throughput. +Both results are below the requested 15 aggregate output tok/s c8 gate. +Per-request rates include queueing and prefill; they are not steady decode +rates. Prefix-cache conditions differ between the two sequential levels. + +A separate arithmetic smoke request produced the correct answer, 42, in +40.255 seconds (10 reported output tokens). With `reasoning_effort="none"`, +the unpatched runtime exposed a closing thinking tag in its answer. The +checkpoint template always opens a thinking block; disabling parsing does +not disable reasoning computation. A separately prepared parser regression fix +has local test coverage but was not deployed for these measurements. + +## Remaining lifecycle checks + +The H200 run verifies worker creation, model readiness, authenticated +Serverless routing, and a correct semantic response. Repeated concurrency +runs, idle shutdown, and cached restart remain outstanding. Do not describe +the result as a complete Serverless lifecycle qualification. The earlier +[raw screening report](https://github.com/earlvanze/FreeToken/blob/0182c7e1a4509c79906b40b0ca29ccdc44b084a1/benchmarks/results/glm53-full-c8-c16-screen-20260907.json) +retains every request result. The +[qualification record](https://github.com/earlvanze/FreeToken/blob/0182c7e1a4509c79906b40b0ca29ccdc44b084a1/benchmarks/results/glm53-full-vast-50116782-qualification-20260907.json) +records the RTX PRO 6000 hardware, checkpoint validation, and original +screening limitations. diff --git a/benchmarks/results/glm53-full-c8-c16-screen-20260907.json b/benchmarks/results/glm53-full-c8-c16-screen-20260907.json new file mode 100644 index 000000000..73c74be59 --- /dev/null +++ b/benchmarks/results/glm53-full-c8-c16-screen-20260907.json @@ -0,0 +1,330 @@ +{ + "destination": "http://127.0.0.1:1919/v1/chat/completions", + "levels": [ + { + "aggregate_output_tokens_per_s": 3.937, + "concurrency": 8, + "errors": {}, + "failures": 0, + "latency_mean_s": 62.983, + "latency_p50_s": 62.984, + "latency_p95_s": 62.989, + "output_tokens": 248, + "per_request_end_to_end_tokens_per_s_mean": 0.492, + "per_request_end_to_end_tokens_per_s_min": 0.492, + "requests": [ + { + "end_to_end_output_tokens_per_s": 0.49223106866956284, + "finish_reason": "length", + "latency_s": 62.97855209300178, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic merge operation from merge sort. Let me think about the best approach", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.4921702335061453, + "finish_reason": "length", + "latency_s": 62.98633661601343, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.4921495640537958, + "finish_reason": "length", + "latency_s": 62.98898193601053, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.4921891551813523, + "finish_reason": "length", + "latency_s": 62.98391517500568, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.4921968928358253, + "finish_reason": "length", + "latency_s": 62.98292502699769, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.4921741144299024, + "finish_reason": "length", + "latency_s": 62.9858399520017, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.4922074220686618, + "finish_reason": "length", + "latency_s": 62.98157770500984, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic merge operation from merge sort. Let me think about the best approach", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.4922283824304707, + "finish_reason": "length", + "latency_s": 62.978895786000066, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + } + ], + "successes": 8, + "wall_s": 62.991 + }, + { + "aggregate_output_tokens_per_s": 6.297, + "concurrency": 16, + "errors": {}, + "failures": 0, + "latency_mean_s": 58.434, + "latency_p50_s": 78.748, + "latency_p95_s": 78.751, + "output_tokens": 496, + "per_request_end_to_end_tokens_per_s_mean": 0.603, + "per_request_end_to_end_tokens_per_s_min": 0.394, + "requests": [ + { + "end_to_end_output_tokens_per_s": 0.8132105022173712, + "finish_reason": "length", + "latency_s": 38.12051112900372, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.8132176628103, + "finish_reason": "length", + "latency_s": 38.12017546799325, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.8132419070784376, + "finish_reason": "length", + "latency_s": 38.1190390339907, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.8132222097296122, + "finish_reason": "length", + "latency_s": 38.11996232900128, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic merge operation, like the merge step in merge sort.\n\nLet me", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.8132897467680965, + "finish_reason": "length", + "latency_s": 38.116796778995194, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic merge operation, like the merge step in merge sort.\n\nLet me", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.8132410764508141, + "finish_reason": "length", + "latency_s": 38.11907796798914, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic merge operation, like the merge step in merge sort.\n\nLet me", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.8132650815359663, + "finish_reason": "length", + "latency_s": 38.11795280998922, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic merge operation, like the merge step in merge sort.\n\nLet me", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.8132933970384179, + "finish_reason": "length", + "latency_s": 38.11662570098997, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.3936418548815157, + "finish_reason": "length", + "latency_s": 78.75178824500472, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic merge operation, like the merge step in merge sort.\n\nLet me", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.39364414520863655, + "finish_reason": "length", + "latency_s": 78.75133004599775, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic merge operation from merge sort. Let me think about the best approach", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.3936464434401844, + "finish_reason": "length", + "latency_s": 78.7508702710038, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.39364954688392906, + "finish_reason": "length", + "latency_s": 78.75024941700394, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.3936553884015476, + "finish_reason": "length", + "latency_s": 78.74908082898764, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic merge operation from merge sort. Let me think about the best way", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.39365080062907953, + "finish_reason": "length", + "latency_s": 78.74999860399112, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.39365950476823347, + "finish_reason": "length", + "latency_s": 78.74825737600622, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic merge operation from merge sort. Let me think about the best way", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + }, + { + "end_to_end_output_tokens_per_s": 0.3936625684556411, + "finish_reason": "length", + "latency_s": 78.74764451599913, + "message": { + "content": "", + "reasoning_content": "The user wants a concise Python function that merges two sorted iterables. This is a classic problem \u2014 essentially the \"merge\" step of merge sort, generalized", + "role": "assistant" + }, + "ok": true, + "output_tokens": 31 + } + ], + "successes": 16, + "wall_s": 78.767 + } + ], + "max_tokens": 32, + "measurement_mode": "nonstreaming_end_to_end_including_routing_queue_and_prefill", + "model": "glm-5.3-nvfp4", + "prompt": "Write a concise Python function that merges two sorted iterables.", + "started_at_utc": "2026-09-07T04:08:06.498640+00:00" +} diff --git a/benchmarks/results/glm53-full-h200-serverless-c8-c16-20260907.json b/benchmarks/results/glm53-full-h200-serverless-c8-c16-20260907.json new file mode 100644 index 000000000..ac5e5cbea --- /dev/null +++ b/benchmarks/results/glm53-full-h200-serverless-c8-c16-20260907.json @@ -0,0 +1,360 @@ +{ + "destination": "tekizai-glm53-full-freetoken-global", + "levels": [ + { + "aggregate_output_tokens_per_s": 17.782, + "concurrency": 8, + "errors": {}, + "failures": 0, + "final_output_responses": 8, + "latency_mean_s": 57.133, + "latency_p50_s": 57.133, + "latency_p95_s": 57.135, + "output_tokens": 1016, + "per_request_end_to_end_tokens_per_s_mean": 2.223, + "per_request_end_to_end_tokens_per_s_min": 2.223, + "requests": [ + { + "end_to_end_output_tokens_per_s": 2.222975528338122, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 57.130633414999465, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.2229172147416993, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 57.1321321180003, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.222921150056283, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 57.132030975000816, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.2228787940677743, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 57.133119600999635, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n try:\n x, y = next(it_a), next(it_b)\n while True:\n if x <= y:\n yield x\n x = next(it_a)\n else:\n yield y\n y = next(it_b)\n except StopIteration:\n # One iterable is exhausted; yield the rest of the other\n rest = it_b if x > y else it_a\n yield from (v for v in", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.2228412536012545, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 57.13408449399867, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.222889825115226, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 57.132836079006665, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.222857631502312, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 57.13366353299352, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.2227996822821874, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 57.13515302899759, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + } + ], + "successes": 8, + "truncated_responses": 8, + "wall_s": 57.136 + }, + { + "aggregate_output_tokens_per_s": 18.023, + "concurrency": 16, + "errors": {}, + "failures": 0, + "final_output_responses": 16, + "latency_mean_s": 85.6, + "latency_p50_s": 112.734, + "latency_p95_s": 112.738, + "output_tokens": 2032, + "per_request_end_to_end_tokens_per_s_mean": 1.649, + "per_request_end_to_end_tokens_per_s_min": 1.126, + "requests": [ + { + "end_to_end_output_tokens_per_s": 2.172408469480269, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 58.46046072099125, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.1724324592106656, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 58.459815153997624, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n import heapq\n return list(heapq.merge(a, b))\n```\n\nOr a manual implementation:\n\n```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x =", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.1722434441530565, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 58.464901961997384, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.172184049077388, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 58.466500595997786, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n import heapq\n return list(heapq.merge(a, b))\n```\n\nOr a manual implementation:\n\n```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x =", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.1724447804971736, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 58.45948359199974, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n import heapq\n return list(heapq.merge(a, b))\n```\n\nOr a manual implementation:\n\n```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x =", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.1724077200685157, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 58.4604808879958, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n import heapq\n return list(heapq.merge(a, b))\n```\n\nOr a manual implementation:\n\n```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x =", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.1721214444333463, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 58.468185711011756, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n import heapq\n return list(heapq.merge(a, b))\n```\n\nOr a manual two-pointer version:\n\n```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 2.172187015718534, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 58.46642074600095, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n import heapq\n return list(heapq.merge(a, b))\n```\n\nOr a manual implementation:\n\n```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x =", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 1.1265104131039854, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 112.73752867500298, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 1.1265430686301972, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 112.73426070999994, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n import heapq\n return list(heapq.merge(a, b))\n```\n\nOr a manual implementation without imports:\n\n```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result", + "reasoning_content": "Simple merge function.", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 1.1264997696660435, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 112.73859384600655, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n merged = []\n\n x = next(it_a, None)\n y = next(it_b, None)\n\n while x is not None and y is not None:\n if x <= y:\n merged.append(x)\n x = next(it_a, None)\n else:\n merged.append(y)\n y = next(it_b, None)\n\n # Append remaining items (one of these is empty)\n merged.extend(v for v in (", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 1.1265211769831036, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 112.7364514709916, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 1.1265450088364997, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 112.73406655200233, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 1.1265489183888535, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 112.73367532200064, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n import heapq\n return list(heapq.merge(a, b))\n```\n\nOr a manual implementation without imports:\n\n```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 1.1265039731652107, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 112.73817316698842, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)\n x = next(it_a, None)\n else:\n result.append(y)\n y = next(it_b, None)\n result.extend([v for v in (x, y) if v is not None])", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + }, + { + "end_to_end_output_tokens_per_s": 1.126531129999895, + "finish_reason": "length", + "has_final_output": true, + "latency_s": 112.73545543300861, + "message": { + "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n import heapq\n return list(heapq.merge(a, b))\n```\n\nOr a manual implementation without imports:\n\n```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted iterables into one sorted list.\"\"\"\n it_a, it_b = iter(a), iter(b)\n x = next(it_a, None)\n y = next(it_b, None)\n result = []\n while x is not None and y is not None:\n if x <= y:\n result.append(x)", + "role": "assistant" + }, + "ok": true, + "output_tokens": 127, + "truncated": true + } + ], + "successes": 16, + "truncated_responses": 16, + "wall_s": 112.747 + } + ], + "max_tokens": 128, + "measurement_mode": "nonstreaming_end_to_end_including_routing_queue_and_prefill", + "model": "glm-5.3-nvfp4", + "prompt": "Write a concise Python function that merges two sorted iterables.", + "reasoning_effort": "low", + "started_at_utc": "2026-09-07T05:13:12.980390+00:00" +} diff --git a/benchmarks/results/glm53-full-h200-serverless-qualification-20260907.json b/benchmarks/results/glm53-full-h200-serverless-qualification-20260907.json new file mode 100644 index 000000000..5ff0d3ed0 --- /dev/null +++ b/benchmarks/results/glm53-full-h200-serverless-qualification-20260907.json @@ -0,0 +1,80 @@ +{ + "status": "serverless_inference_and_concurrency_benchmark_verified_lifecycle_pending", + "observed_at": "2026-09-07T05:16:03Z", + "provider": "Vast.ai Serverless", + "endpoint_id": 36422, + "endpoint_name": "tekizai-glm53-full-freetoken-global", + "workergroup_id": 45647, + "worker_instance_id": 50124528, + "machine_id": 147075, + "region": "JP", + "hardware": { + "gpu": "H200 NVL", + "gpu_memory_mib": 143771, + "effective_cpu_cores": 32, + "disk_gib": 520 + }, + "checkpoint": "LibertAIDAI/GLM-5.3-NVFP4", + "served_model": "glm-5.3-nvfp4", + "runtime": "FreeToken 0.1.2", + "template_id": 696081, + "template_hash": "af9cd2b7d62b225ac623902af890bb35", + "readiness": { + "freetoken_api_ready_at_utc": "2026-09-07T05:09:16Z", + "vast_model_loaded_at_utc": "2026-09-07T05:11:34Z", + "vast_worker_idle_at_utc": "2026-09-07T05:11:44.757Z" + }, + "semantic_smoke": { + "artifact": "glm53-full-h200-serverless-smoke-20260907.json", + "transport": "authenticated_vast_serverless_endpoint", + "prompt": "What is 17 + 25? Reply with only the number.", + "actual": "42", + "answer_correct": true, + "finish_reason": "stop", + "latency_s": 10.402, + "requests_succeeded": 1, + "requests_failed": 0 + }, + "concurrency_benchmark": { + "artifact": "glm53-full-h200-serverless-c8-c16-20260907.json", + "transport": "authenticated_vast_serverless_endpoint", + "measurement_mode": "nonstreaming_end_to_end_including_routing_queue_and_prefill", + "prompt": "Write a concise Python function that merges two sorted iterables.", + "max_tokens": 128, + "reasoning_effort": "low", + "levels": [ + { + "concurrency": 8, + "requests_succeeded": 8, + "requests_failed": 0, + "wall_s": 57.136, + "aggregate_output_tokens_per_s": 17.782, + "latency_mean_s": 57.133, + "latency_p95_s": 57.135, + "output_tokens": 1016 + }, + { + "concurrency": 16, + "requests_succeeded": 16, + "requests_failed": 0, + "wall_s": 112.747, + "aggregate_output_tokens_per_s": 18.023, + "latency_mean_s": 85.6, + "latency_p95_s": 112.738, + "output_tokens": 2032 + } + ], + "note": "All responses contained assistant output and reached the fixed 128-token cap. These are single-run throughput measurements, not full-completion quality measurements or repeated statistical qualification." + }, + "lifecycle": { + "autoscaled_worker_creation_verified": true, + "serverless_routing_verified": true, + "idle_stop_verified": false, + "cached_restart_verified": false + }, + "notes": [ + "FreeToken loaded the full checkpoint, allocated 8384 KV tokens, and captured CUDA graphs for batch sizes 1, 2, 4, and 8.", + "The workergroup permits eight running requests, so the c16 result includes serverless queueing above the worker's active-request limit.", + "The semantic smoke is the completion-quality check; the capped c8/c16 workload is the throughput check." + ] +} diff --git a/benchmarks/results/glm53-full-h200-serverless-smoke-20260907.json b/benchmarks/results/glm53-full-h200-serverless-smoke-20260907.json new file mode 100644 index 000000000..ba8601550 --- /dev/null +++ b/benchmarks/results/glm53-full-h200-serverless-smoke-20260907.json @@ -0,0 +1,42 @@ +{ + "destination": "tekizai-glm53-full-freetoken-global", + "levels": [ + { + "aggregate_output_tokens_per_s": 0.288, + "concurrency": 1, + "errors": {}, + "failures": 0, + "final_output_responses": 1, + "latency_mean_s": 10.402, + "latency_p50_s": 10.402, + "latency_p95_s": 10.402, + "output_tokens": 3, + "per_request_end_to_end_tokens_per_s_mean": 0.288, + "per_request_end_to_end_tokens_per_s_min": 0.288, + "requests": [ + { + "end_to_end_output_tokens_per_s": 0.28840918986416886, + "finish_reason": "stop", + "has_final_output": true, + "latency_s": 10.401887683998211, + "message": { + "content": "42", + "role": "assistant" + }, + "ok": true, + "output_tokens": 3, + "truncated": false + } + ], + "successes": 1, + "truncated_responses": 0, + "wall_s": 10.403 + } + ], + "max_tokens": 24, + "measurement_mode": "nonstreaming_end_to_end_including_routing_queue_and_prefill", + "model": "glm-5.3-nvfp4", + "prompt": "What is 17 + 25? Reply with only the number.", + "reasoning_effort": "low", + "started_at_utc": "2026-09-07T05:12:03.652578+00:00" +} diff --git a/benchmarks/results/glm53-full-vast-50116782-qualification-20260907.json b/benchmarks/results/glm53-full-vast-50116782-qualification-20260907.json new file mode 100644 index 000000000..940a1875b --- /dev/null +++ b/benchmarks/results/glm53-full-vast-50116782-qualification-20260907.json @@ -0,0 +1,70 @@ +{ + "status": "inference_verified_concurrency_screen_below_target", + "observed_at": "2026-09-07T04:10:30Z", + "model": "LibertAIDAI/GLM-5.3-NVFP4", + "instance_id": 50116782, + "endpoint_id": 36422, + "endpoint_name": "tekizai-glm53-full-freetoken-global", + "endpoint_setup_status": "stopped_with_zero_worker_group_retained_instance_not_attached", + "workergroup_id": 45647, + "candidate_template_id": 696081, + "candidate_template_hash": "b08bd37c816e5292d52d830831186635", + "candidate_worker_commit": "a7398b3548ab6b8d392a02bffff4d142beae87c2", + "candidate_template_note": "Pinned standalone bootstrap prepared for future serverless workers; retained benchmark instance was not restarted or attached.", + "machine_id": 104296, + "country": "IN", + "gpu": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition", + "gpu_memory_mib": 97887, + "gpu_power_limit_w": 300, + "driver": "595.84", + "container_memory_limit_bytes": 2057523167232, + "container_cpu_quota_us": 3648000, + "container_cpu_period_us": 100000, + "container_cpu_equivalent_cores": 36.48, + "disk_gib": 520, + "freetoken_commit": "55120d9114eb18d90e4dfe2e1454f25f0098723e", + "planned_concurrency": [1, 2, 4, 8, 16], + "checkpoint_validation": { + "indexed_shards": 282, + "shard_bytes": 464822689680, + "missing_shards": [], + "empty_shards": [], + "header_and_tensor_offset_errors": [], + "cryptographic_hashes_verified": false, + "leftover_partial_cache_files": 8 + }, + "inference_verified": true, + "smoke_test": { + "transport": "direct_loopback_not_serverless", + "prompt": "What is 17 + 25? Reply with only the number.", + "max_tokens": 16, + "reasoning_effort": "none", + "elapsed_s": 40.255191267991904, + "completion_tokens": 10, + "content": "17 + 25 = 4242", + "finish_reason": "stop", + "answer_correct": true, + "formatting_pass": false + }, + "serverless_lifecycle_verified": false, + "benchmark_results": [{ + "artifact": "glm53-full-c8-c16-screen-20260907.json", + "qualification": false, + "running_slots": 8, + "c8_aggregate_output_tokens_per_s": 3.937, + "c16_aggregate_output_tokens_per_s": 6.297, + "transport_failures": 0, + "completed_code_answers": 0, + "reason": "Single short direct-loopback screening run, not repeated serverless qualification. All 24 responses exhausted the 32-token cap during reasoning with no final code answer. Both aggregate rates below 15 tok/s." + }], + "notes": [ + "Hardware values were read from nvidia-smi and cgroup cpu.max/memory.max inside the worker.", + "Host-wide free output overstates the memory available to this container.", + "The CPU quota is below the profile's 48-core admission target; throughput remains unqualified.", + "Downloader exited; all indexed shards have valid JSON headers, indexed tensor names and file lengths matching final tensor offsets. This is structural validation, not a cryptographic integrity check.", + "FreeToken 0.1.2 installed; full checkpoint loaded and CUDA graphs captured for batch sizes 1,2,4,8. API ready at 04:01:58 UTC; startup warmup returned HTTP 200 at 04:05:13 UTC.", + "Smoke test generated the correct answer but exposed a closing think tag despite reasoning_effort=none; formatting is not qualified.", + "Vast PyWorker startup calibration finished at 04:07:25 UTC. Its workload/s metric is not a token/s qualification result.", + "Direct-loopback c8/c16 screening completed after startup calibration. c8: 248 reported output tokens in 62.991s; c16: 496 in 78.767s. Mean per-request end-to-end rates 0.492 and 0.603 tok/s respectively. All responses are truncated reasoning, not completed coding tasks. Formal repeated benchmarks and serverless lifecycle remain pending." + ] +} diff --git a/docs/models.md b/docs/models.md index c9499163f..77439d56c 100644 --- a/docs/models.md +++ b/docs/models.md @@ -8,7 +8,7 @@ for them; other checkpoints of the same architectures work too. |---|---| | DeepSeek-V4 | [deepseek-ai/DeepSeek-V4-Flash-0731](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731) | | GLM-5.3-Flash | [RedHatAI/GLM-5.3-Flash-NVFP4](https://huggingface.co/RedHatAI/GLM-5.3-Flash-NVFP4) | -| GLM-5.2 | [nvidia/GLM-5.2-NVFP4](https://huggingface.co/nvidia/GLM-5.2-NVFP4) | +| GLM-5.3 / GLM-5.2 | [LibertAIDAI/GLM-5.3-NVFP4](https://huggingface.co/LibertAIDAI/GLM-5.3-NVFP4), [nvidia/GLM-5.2-NVFP4](https://huggingface.co/nvidia/GLM-5.2-NVFP4) | | GLM-4.7 | [nvidia/GLM-4.7-NVFP4](https://huggingface.co/nvidia/GLM-4.7-NVFP4) | | Qwen3.8-Flash-Next | [Qwen/Qwen3.8-Flash-Next-FP8](https://huggingface.co/Qwen/Qwen3.8-Flash-Next-FP8), [RadixArk/Qwen3.8-Flash-Next-NVFP4](https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4) | | Qwen3.6 / Qwen3.5 MoE | [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) ([-FP8](https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8)), [nvidia/Qwen3.6-35B-A3B-NVFP4](https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4), [Qwen/Qwen3.5-35B-A3B](https://huggingface.co/Qwen/Qwen3.5-35B-A3B) ([-FP8](https://huggingface.co/Qwen/Qwen3.5-35B-A3B-FP8)) | @@ -39,5 +39,8 @@ for them; other checkpoints of the same architectures work too. FreeToken's fast-load format, and `ft serve --model` auto-detects the result. - DeepSeek-V4 checkpoints must keep the `inference/config.json` subdir — the authoritative model args are read from there. +- Full GLM-5.3 uses the existing `glm_moe_dsa` path with the + `LibertAIDAI/GLM-5.3-NVFP4` checkpoint. The official `zai-org/GLM-5.3` + checkpoint is FP8 and is not accepted by the NVFP4 expert loader. - Qwen3.8-Flash-Next keeps a 47.7 GiB PLE n-gram table pinned in host RAM. - Multimodal checkpoints are served text-only. diff --git a/python/freetoken/kernel/aot_models.py b/python/freetoken/kernel/aot_models.py index 0ce249de8..c982cf4bf 100644 --- a/python/freetoken/kernel/aot_models.py +++ b/python/freetoken/kernel/aot_models.py @@ -266,13 +266,17 @@ def expert_bank_row_bytes(fmt: str, hidden_size: int, moe_intermediate_size: int AotModel( # MLA latent-KV: the dsa backend writes the latent via torch scatter, # not store_cache (same as DSV4 below), so no paged-KV store groups. - name="zai-org/GLM-5.2-NVFP4", + # GLM-5.3 keeps GLM-5.2's architecture, tensor names, and geometry. The + # LibertAI NVFP4 conversion follows NVIDIA's 5.2 recipe tensor-for-tensor, + # so both checkpoints share the same offload-bank and AOT shapes. + name="LibertAIDAI/GLM-5.3-NVFP4", architecture="GlmMoeDsaForCausalLM", hidden_size=6144, kv_groups=(), top_k=8, moe_intermediate_size=2048, expert_formats=_NVFP4_FORMATS, + aliases=("nvidia/GLM-5.2-NVFP4", "zai-org/GLM-5.2-NVFP4"), ), AotModel( # GLM-5.3-Flash: hybrid KDA + NoPE-MLA/DSA (kpool indexer). Latent writes diff --git a/tests/models/test_glm_moe_dsa_config.py b/tests/models/test_glm_moe_dsa_config.py new file mode 100644 index 000000000..1a124ba1f --- /dev/null +++ b/tests/models/test_glm_moe_dsa_config.py @@ -0,0 +1,69 @@ +from types import SimpleNamespace + + +def _glm53_nvfp4_config(): + """Small projection of LibertAIDAI/GLM-5.3-NVFP4 config.json.""" + return SimpleNamespace( + model_type="glm_moe_dsa", + architectures=["GlmMoeDsaForCausalLM"], + quantization_config={"quant_algo": "NVFP4", "quant_method": "modelopt"}, + hidden_size=6144, + num_hidden_layers=78, + vocab_size=154880, + intermediate_size=12288, + hidden_act="silu", + rms_norm_eps=1e-5, + tie_word_embeddings=False, + max_position_embeddings=1048576, + rope_parameters={"rope_theta": 8000000, "rope_type": "default"}, + num_attention_heads=64, + q_lora_rank=2048, + kv_lora_rank=512, + qk_nope_head_dim=192, + qk_rope_head_dim=64, + v_head_dim=256, + rope_interleave=True, + indexer_rope_interleave=True, + index_n_heads=32, + index_head_dim=128, + index_topk=2048, + indexer_types=[ + "full" if layer < 3 or (layer - 2) % 4 == 0 else "shared" + for layer in range(78) + ], + n_routed_experts=256, + num_experts_per_tok=8, + moe_intermediate_size=2048, + norm_topk_prob=True, + first_k_dense_replace=3, + n_shared_experts=1, + routed_scaling_factor=2.5, + n_group=1, + topk_group=1, + attention_bias=False, + ) + + +def test_glm53_nvfp4_maps_existing_glm_moe_dsa_path(): + from freetoken.models.glm_moe_dsa.config import parse_config + + config = parse_config(_glm53_nvfp4_config()) + + assert config.architectures == ["GlmMoeDsaForCausalLM"] + assert (config.num_layers, config.hidden_size) == (78, 6144) + assert (config.num_experts, config.num_experts_per_tok) == (256, 8) + assert config.expert_quant == "nvfp4" + assert config.glm_dsa_args.qk_head_dim == 256 + assert config.glm_dsa_args.index_topk == 2048 + full = config.attention_groups[0] + assert full.mla and full.head_dim == 576 + assert full.num_index_layers == 21 + + +def test_glm53_is_primary_full_glm_aot_checkpoint(): + from freetoken.kernel.aot_models import SUPPORTED_MODELS + + entry = next(m for m in SUPPORTED_MODELS if m.architecture == "GlmMoeDsaForCausalLM") + assert entry.name == "LibertAIDAI/GLM-5.3-NVFP4" + assert "nvidia/GLM-5.2-NVFP4" in entry.aliases + assert (entry.hidden_size, entry.moe_intermediate_size, entry.top_k) == (6144, 2048, 8)