Recipe · 4× RTX PRO 6000 · with Aevonix Research
GLM 5.3 Flash on four RTX PRO 6000s, with TensorFold.
Serve GLM-5.3-Flash from one Linux host with four 96 GB RTX PRO 6000 Blackwell GPUs, through an OpenAI-compatible API, with 40 concurrent requests, the full 1,048,576-token context, tool calling, structured outputs, and image and video input. 264.2 tok/s for one request. A collaboration with Marc Seal of Aevonix Research: my EXL3 quant and recipe patches, with Aevonix's single-host PCIe work on top. One command builds, downloads, tunes and starts it.
The collaboration
Two labs, one recipe.
This recipe is not mine alone. It was made with Marc Seal of Aevonix Research, and it lives in their GitHub. My GLM 5.3 Flash recipe for DGX Sparks was the starting point: its checkpoint and patches, ported to TensorFold v0.6.2 and taken to one PC with four GPUs over PCIe.
From my lab
- The checkpoint: Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold, my EXL3 quant: routed experts at 4 bits a weight, with the dense layers converted to FP8 at load time.
- 55 of my recipe patches, ported to TensorFold v0.6.2: dense quantization, the FP8 KV cache, drafting and prompt reuse, concurrent serving, vision, tool calling and API diagnostics.
- The benchmark: every speed below was measured with sparkDash.
From Aevonix Research
- The single-host release and 32 patches of its own.
- One host over PCIe: CUDA IPC exchanges, separate NCCL communicators, prompt lanes and faster prompt kernels.
- One command for one host:
./start.shchecks the host with a clear code for every failure, builds at a pinned commit, downloads, autotunes the kernels and runs a smoke test, built on my two-Spark startup flow.
Performance
Measured, not promised.
Four RTX PRO 6000 Blackwell Max-Q GPUs at 250 W each, one PCIe host. Measured with sparkDash v1.8.9 on 3 October 2026: greedy, temperature 0, top_p 1, thinking off, 400-token replies, with FP8 dense layers, the FP8 KV cache, DFlash2 and the default settings.
Decode
Aggregate across the concurrent requests, median per request, and median time to first token.
| Requests | Prose | Per request | TTFT | Structured | Per request | TTFT |
|---|---|---|---|---|---|---|
| 1 | 264.2 | 264.2 | 36 ms | 539.5 | 539.5 | 35 ms |
| 2 | 434.3 | 221.8 | 48 ms | 877.9 | 442.3 | 44 ms |
| 4 | 585.7 | 147.9 | 71 ms | 1,018.0 | 304.9 | 67 ms |
| 8 | 821.3 | 107.2 | 112 ms | 1,487.0 | 221.4 | 103 ms |
| 16 | 1,060.7 | 69.4 | 220 ms | 2,041.4 | 136.9 | 274 ms |
tok/s. Within a fixed configuration, replies served together equal the same requests served alone, and drafted replies equal serial ones.
Code
The same decode measurement, on code.
| Requests | Code | Per request | TTFT |
|---|---|---|---|
| 1 | 430.3 | 430.3 | 39 ms |
| 2 | 613.3 | 311.8 | 59 ms |
| 4 | 760.3 | 214.4 | 92 ms |
| 8 | 1,013.6 | 141.9 | 140 ms |
| 16 | 1,267.8 | 85.1 | 278 ms |
Prefill
Cold prompts, with a unique prefix per size.
| Prompt | Prefill | First token |
|---|---|---|
| 8,216 tokens | 8,416 tok/s | 0.98 s |
| 16,409 tokens | 6,903 tok/s | 2.38 s |
| 32,791 tokens | 6,508 tok/s | 5.04 s |
| 65,563 tokens | 5,527 tok/s | 11.86 s |
| 131,098 tokens | 8,272 tok/s | 15.85 s |
| 262,168 tokens | 7,842 tok/s | 33.43 s |
Prompt reuse
A conversation's next turn resumes from its kept prompt state instead of reading the whole conversation again. Measured with a separate client workload: the first-time values are the median for three fresh prompts, the next-time values the median for turns 2 to 5 of a 130K-token conversation. The two sets have different prompt sizes; no matched first turn at 130K was recorded.
| Prompt | First time | Next time |
|---|---|---|
| 32K-token prompt | 3.87 s | not recorded |
| 100K-token prompt | 12.09 s | not recorded |
| 140K-token prompt | 17.12 s | not recorded |
| A 130K-token conversation's next turn, thinking off | not recorded | 0.3695 s |
| A 130K-token conversation's next turn, low effort | not recorded | 0.3870 s |
Quality checks
Small, separate checks, not a claim that the model is unchanged.
| Check | Score |
|---|---|
| Tool calls | 40/40 |
| JSON, instruction only | 20/20 |
JSON with response_format | 20/20 |
| Needles at 100K / 200K / 250K | 3/3 at each |
| Images | 3/3 |
Tool bursts, default choice / required | 80/90 / 90/90 |
Why FP8 dense layers
Against the checkpoint's own BF16 dense layers, on 100 fixed greedy prompts.
| FP8 dense | 4-bit dense | |
|---|---|---|
| First token agrees | 86/100 | 73/100 |
| Mean identical prefix | 12.9 tokens | 7.0 tokens |
| Whole 32-token reply agrees | 19/100 | 6/100 |
FP8 costs about 2% single-stream speed, 11% cold and 7% warm time to first token against 4-bit; DENSE=q4 stays the speed option.


At 300 W, throughput sampled higher up to 16 streams, but the Max-Q cards overheated in the tested chassis, so every number here and the operating configuration use 250 W. The scripts warn about a different power limit and never change it.
Speed preview
What 264.2 tok/s feels like.
Sample text played at the measured decode speed, at about four characters a token: one request at 264.2 tok/s, or four at once at 147.9 tok/s each (585.7 combined).
The engine
Powered by TensorFold.
TensorFold, by Ash Hart, is an open-source engine that serves language models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API. Each model family brings its own kernels and its own draft verification, tuned for that model instead of one generic path for all.
Exact decoding
Drafts make it fast without changing a word: a draft is accepted only when it equals the token the same engine would produce one at a time. Send a request with "draft": false to compare.
Apple Silicon and NVIDIA
MLX on Macs, and CUDA on NVIDIA GPUs from compute capability 8.9: Ada (RTX 40), Hopper and Blackwell, including the RTX PRO 6000, the DGX Spark's GB10 and the RTX 50 series.
Many model families
GLM-5.3-Flash, Qwen3.8-27B, Qwen3.8 Flash Next, Nemotron 3.5 Lightning, Gemma 4 and DeepSeek-V4-Flash, among others, each with its own kernels.
In this recipe
TensorFold v0.6.2 at a pinned commit, four ranks on one host, with 87 patches on top: 55 of mine and 32 from Aevonix Research. Batched drafting, wider verify windows and round caps take it to 40 concurrent requests.
What you need
Four GPUs, one host.
Four RTX PRO 6000s
Four 96 GB RTX PRO 6000 Blackwell GPUs in one Linux x86_64 host, working NVIDIA drivers, and CUDA peer access between every pair. Keep them idle for the first compile and the kernel autotune.
Free GPU memory
At least 90,000 MiB free on each GPU at launch (about 88 GiB). The defaults keep 7.5 GiB in reserve and cap the shared KV cache at 60 GiB on each GPU.
Docker
Docker with the NVIDIA Container Toolkit, and your user in the docker group. Plus Bash, Python 3, Git, GNU patch, curl, gzip and flock. A Hugging Face token is optional.
Disk
Budget 180 GiB for the checkpoint, 10 GiB for DFlash2, 40 GiB for the image and build and 10 GiB for kernels. The script adds them up when they share a filesystem.
Quick start
One command, start to finish.
Clone and start
git clone https://github.com/Aevonix/GLM-5.3-Flash-EXL3-4x-RTX-PRO-6000-TensorFold.git
cd GLM-5.3-Flash-EXL3-4x-RTX-PRO-6000-TensorFold
./start.sh
The first run checks the host, builds TensorFold at the pinned commit with all 87 patches, downloads the checkpoint and the drafter, compiles the kernels and autotunes the launch table. That takes a while and needs idle GPUs; later runs reuse what is ready. It sends a short smoke request and prints GLM-5.3-Flash-EXL3 is now LIVE! on port 8020.
DRY_RUN=1 ./start.sh # see every step and command first, without Docker or GPUs
DRAFTER=mtp ./start.sh # without the non-commercial DFlash2 drafter (one request at a time)

Talk to it
Any OpenAI client works with base_url = "http://<host>:8020/v1" and the model glm-5.3-flash. Thinking is off by default; a request turns it on with "chat_template_kwargs": {"enable_thinking": true}.
curl -s http://127.0.0.1:8020/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "glm-5.3-flash",
"messages": [{"role": "user", "content": "Write a Python fibonacci function."}],
"temperature": 0,
"max_tokens": 2000
}'
./start.sh restart # restart, e.g. after changing a setting
./stop.sh # stop, and save the logs under logs/
docker logs -f glm53-tf # follow the server
curl -s http://127.0.0.1:8020/health
If start.sh stops at a check
Every failed check prints a code. The ones you are most likely to see:
| Code | What to do |
|---|---|
E_GPU_BUSY / E_GPU_MEMORY | Other work is on the GPUs: stop it. Each GPU needs at least 90,000 MiB free. |
E_P2P | CUDA peer access failed: check the driver and the platform's PCIe topology, ACS and IOMMU settings. Every GPU pair must support peer access. |
E_DOCKER_GROUP | Add your user to the docker group and log in again. |
W_POWER | The power limit isn't 250 W, so results may differ. The script never changes it. |
E_AUTOTUNE | Read the compiler error, leave the GPUs idle and rerun. Incomplete tuning is never marked ready. |
The full list, with every code, is in the README.
Images and video
It can see, too.
Vision is on by default. Send images and videos as OpenAI-style content parts in a user message: an image_url part, or a video_url part for a clip. The vision tower runs on rank 0.
| Input | Content part |
|---|---|
| Image | {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}} |
| Video | {"type": "video_url", "video_url": {"url": "data:video/mp4;base64,..."}} |
Only data URLs are accepted; outside media URLs are off. VISION=0 serves text only.
Memory
A million tokens a request, forty requests.
All requests share one FP8 cache pool. Any one request can use the full 1,048,576-token window, but forty full windows at once are not promised. /health shows pool_tokens and pool_free_tokens, and the engine refuses a window that doesn't fit instead of shrinking it.
| Setting | Default |
|---|---|
Concurrent requests (PARALLEL) | 40 |
Prompt plus reply (CONTEXT) | 1,048,576 tokens a request |
KV cache (KV) | FP8 |
| Memory reserve / KV cap | 7.5 GiB / 60 GiB a GPU |
Settings
The ones you might change.
Settings come from the environment, then scripts/local.sh, then .env, then the defaults in scripts/config.sh. Set one for a single run (DENSE=q4 ./start.sh restart), or keep it in scripts/local.sh. The full list is in the README.
| Variable | Default | Meaning |
|---|---|---|
PARALLEL | 40 | Concurrent requests, 1 to 64 with DFlash2 (1 with MTP) |
CONTEXT | 1048576 | Prompt plus reply tokens per request |
DENSE | fp8 | fp8, q4 (the speed option) or the checkpoint's bf16 dense layers |
KV | fp8 | KV cache; bf16 needs more memory |
DRAFTER | dflash2 | Inco AI's DFlash2 drafter (non-commercial use only), or mtp, the checkpoint's own MTP head, one request at a time |
THINKING / REASONING_EFFORT | 0 / high | Thinking off by default, and its effort when a request turns it on |
MAX_TOKENS | 32768 | The reply budget when a request doesn't set one |
VISION | 1 | Image and video input |
GPUS / PORT / HOST | 0,1,2,3 / 8020 / 0.0.0.0 | The four GPUs, and the API's port and address |
The API
/v1/chat/completions with the model glm-5.3-flash, /v1/models, /tokenize and /detokenize, /health and Prometheus /metrics. Tool calling (complete calls streamed, arguments typed by their schema) and structured outputs (json_object or json_schema, enforced with xgrammar). A sampled request without top_k is served with 20; send top_k to compare with another engine.
Exact replies
The checks compare drafted replies with serial ones, concurrent requests with solo ones, and long prompts with saved token hashes, all within one fixed configuration. FP8 dense layers, the FP8 KV cache and chunked KDA change the arithmetic, so passing the checks doesn't mean the output matches BF16 or the original model; the quality numbers above are measured for that.
Licenses
Free, with a few rules.
- The recipe, its code and documentation, is Apache License 2.0. Upstream code and model weights keep their own licenses.
- The checkpoint, Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold, is under the Apache License 2.0.
- The DFlash2 drafter is CC BY-NC-ND 4.0, non-commercial use only, and is not redistributed by the recipe.
DRAFTER=mtpserves without it. - TensorFold is Apache 2.0; its earlier code keeps its MIT notices.
- The base model, GLM-5.3-Flash, is by Z.ai.
Credits
Built together, on the work of others.
- Marc Seal and Aevonix Research: the single-host release and its 32 patches
- My lab: the EXL3 checkpoint, the original recipe and 55 of its patches, and sparkDash
- TensorFold by Ash Hart
- GLM-5.3-Flash by Z.ai
- The DFlash2 drafter by Inco AI
- The EXL3 format (ExLlamaV3) by turboderp
- b12x by local-inference-lab, and glm53-tensorfold-spark by Jay Leaton
The full list is in the recipe's CREDITS.md.
Run it on your four GPUs.
Free and open, with every patch and every number published.