Recipe · 4× RTX PRO 6000 · with Aevonix Research

GLM 5.3 Flash on four RTX PRO 6000s, with TensorFold.

Serve GLM-5.3-Flash from one Linux host with four 96 GB RTX PRO 6000 Blackwell GPUs, through an OpenAI-compatible API, with 40 concurrent requests, the full 1,048,576-token context, tool calling, structured outputs, and image and video input. 264.2 tok/s for one request. A collaboration with Marc Seal of Aevonix Research: my EXL3 quant and recipe patches, with Aevonix's single-host PCIe work on top. One command builds, downloads, tunes and starts it.

Get the free recipe Aevonix Research

Install with an AI agent

Install this model

Paste this into an AI coding agent that can run commands on the machine that will serve the model (the head node, for a Spark cluster), such as Claude Code or Codex. It reads the recipe, adapts it to your setup, starts the server and checks that it works.


        
Aevonix Research: a dotted ribbon beside GLM 5.3 Flash EXL3 on TensorFold, 4x RTX PRO 6000, in collaboration with Mia's AI Lab

The collaboration

Two labs, one recipe.

This recipe is not mine alone. It was made with Marc Seal of Aevonix Research, and it lives in their GitHub. My GLM 5.3 Flash recipe for DGX Sparks was the starting point: its checkpoint and patches, ported to TensorFold v0.6.2 and taken to one PC with four GPUs over PCIe.

From my lab

  • The checkpoint: Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold, my EXL3 quant: routed experts at 4 bits a weight, with the dense layers converted to FP8 at load time.
  • 55 of my recipe patches, ported to TensorFold v0.6.2: dense quantization, the FP8 KV cache, drafting and prompt reuse, concurrent serving, vision, tool calling and API diagnostics.
  • The benchmark: every speed below was measured with sparkDash.

From Aevonix Research

  • The single-host release and 32 patches of its own.
  • One host over PCIe: CUDA IPC exchanges, separate NCCL communicators, prompt lanes and faster prompt kernels.
  • One command for one host: ./start.sh checks the host with a clear code for every failure, builds at a pinned commit, downloads, autotunes the kernels and runs a smoke test, built on my two-Spark startup flow.

Performance

Measured, not promised.

One request264.2tok/s, prose
Sixteen at once1,060.7tok/s combined
Code430.3tok/s, one request
Prefill8,416tok/s, 8k prompt

Four RTX PRO 6000 Blackwell Max-Q GPUs at 250 W each, one PCIe host. Measured with sparkDash v1.8.9 on 3 October 2026: greedy, temperature 0, top_p 1, thinking off, 400-token replies, with FP8 dense layers, the FP8 KV cache, DFlash2 and the default settings.

Decode

Aggregate across the concurrent requests, median per request, and median time to first token.

RequestsProsePer requestTTFTStructuredPer requestTTFT
1264.2264.236 ms539.5539.535 ms
2434.3221.848 ms877.9442.344 ms
4585.7147.971 ms1,018.0304.967 ms
8821.3107.2112 ms1,487.0221.4103 ms
161,060.769.4220 ms2,041.4136.9274 ms

tok/s. Within a fixed configuration, replies served together equal the same requests served alone, and drafted replies equal serial ones.

Code

The same decode measurement, on code.

RequestsCodePer requestTTFT
1430.3430.339 ms
2613.3311.859 ms
4760.3214.492 ms
81,013.6141.9140 ms
161,267.885.1278 ms

Prefill

Cold prompts, with a unique prefix per size.

PromptPrefillFirst token
8,216 tokens8,416 tok/s0.98 s
16,409 tokens6,903 tok/s2.38 s
32,791 tokens6,508 tok/s5.04 s
65,563 tokens5,527 tok/s11.86 s
131,098 tokens8,272 tok/s15.85 s
262,168 tokens7,842 tok/s33.43 s

Prompt reuse

A conversation's next turn resumes from its kept prompt state instead of reading the whole conversation again. Measured with a separate client workload: the first-time values are the median for three fresh prompts, the next-time values the median for turns 2 to 5 of a 130K-token conversation. The two sets have different prompt sizes; no matched first turn at 130K was recorded.

PromptFirst timeNext time
32K-token prompt3.87 snot recorded
100K-token prompt12.09 snot recorded
140K-token prompt17.12 snot recorded
A 130K-token conversation's next turn, thinking offnot recorded0.3695 s
A 130K-token conversation's next turn, low effortnot recorded0.3870 s

Quality checks

Small, separate checks, not a claim that the model is unchanged.

CheckScore
Tool calls40/40
JSON, instruction only20/20
JSON with response_format20/20
Needles at 100K / 200K / 250K3/3 at each
Images3/3
Tool bursts, default choice / required80/90 / 90/90

Why FP8 dense layers

Against the checkpoint's own BF16 dense layers, on 100 fixed greedy prompts.

FP8 dense4-bit dense
First token agrees86/10073/100
Mean identical prefix12.9 tokens7.0 tokens
Whole 32-token reply agrees19/1006/100

FP8 costs about 2% single-stream speed, 11% cold and 7% warm time to first token against 4-bit; DENSE=q4 stays the speed option.

sparkDash decode benchmark card: prose, 1 to 16 concurrent requests, 4x RTX PRO 6000 at 250 W
sparkDash prefill benchmark card: 8k to 256k cold prompts, 250 W

At 300 W, throughput sampled higher up to 16 streams, but the Max-Q cards overheated in the tested chassis, so every number here and the operating configuration use 250 W. The scripts warn about a different power limit and never change it.

Speed preview

What 264.2 tok/s feels like.

Sample text played at the measured decode speed, at about four characters a token: one request at 264.2 tok/s, or four at once at 147.9 tok/s each (585.7 combined).

TensorFold · GLM-5.3-Flash-EXL3 · 4× RTX PRO 6000
264.2 tok/sSample text, played at the measured speed

The engine

Powered by TensorFold.

TensorFold, by Ash Hart, is an open-source engine that serves language models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API. Each model family brings its own kernels and its own draft verification, tuned for that model instead of one generic path for all.

Exact decoding

Drafts make it fast without changing a word: a draft is accepted only when it equals the token the same engine would produce one at a time. Send a request with "draft": false to compare.

Apple Silicon and NVIDIA

MLX on Macs, and CUDA on NVIDIA GPUs from compute capability 8.9: Ada (RTX 40), Hopper and Blackwell, including the RTX PRO 6000, the DGX Spark's GB10 and the RTX 50 series.

Many model families

GLM-5.3-Flash, Qwen3.8-27B, Qwen3.8 Flash Next, Nemotron 3.5 Lightning, Gemma 4 and DeepSeek-V4-Flash, among others, each with its own kernels.

In this recipe

TensorFold v0.6.2 at a pinned commit, four ranks on one host, with 87 patches on top: 55 of mine and 32 from Aevonix Research. Batched drafting, wider verify windows and round caps take it to 40 concurrent requests.

What you need

Four GPUs, one host.

Four RTX PRO 6000s

Four 96 GB RTX PRO 6000 Blackwell GPUs in one Linux x86_64 host, working NVIDIA drivers, and CUDA peer access between every pair. Keep them idle for the first compile and the kernel autotune.

Free GPU memory

At least 90,000 MiB free on each GPU at launch (about 88 GiB). The defaults keep 7.5 GiB in reserve and cap the shared KV cache at 60 GiB on each GPU.

Docker

Docker with the NVIDIA Container Toolkit, and your user in the docker group. Plus Bash, Python 3, Git, GNU patch, curl, gzip and flock. A Hugging Face token is optional.

Disk

Budget 180 GiB for the checkpoint, 10 GiB for DFlash2, 40 GiB for the image and build and 10 GiB for kernels. The script adds them up when they share a filesystem.

Quick start

One command, start to finish.

1

Clone and start

git clone https://github.com/Aevonix/GLM-5.3-Flash-EXL3-4x-RTX-PRO-6000-TensorFold.git
cd GLM-5.3-Flash-EXL3-4x-RTX-PRO-6000-TensorFold
./start.sh

The first run checks the host, builds TensorFold at the pinned commit with all 87 patches, downloads the checkpoint and the drafter, compiles the kernels and autotunes the launch table. That takes a while and needs idle GPUs; later runs reuse what is ready. It sends a short smoke request and prints GLM-5.3-Flash-EXL3 is now LIVE! on port 8020.

DRY_RUN=1 ./start.sh          # see every step and command first, without Docker or GPUs
DRAFTER=mtp ./start.sh        # without the non-commercial DFlash2 drafter (one request at a time)
start.sh command deck: AEVONIX, the system (TensorFold v0.6.2 + 87 patches, 4x RTX PRO 6000, PCIe, 250 W), the measured decode bars and the start sequence
2

Talk to it

Any OpenAI client works with base_url = "http://<host>:8020/v1" and the model glm-5.3-flash. Thinking is off by default; a request turns it on with "chat_template_kwargs": {"enable_thinking": true}.

curl -s http://127.0.0.1:8020/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "glm-5.3-flash",
  "messages": [{"role": "user", "content": "Write a Python fibonacci function."}],
  "temperature": 0,
  "max_tokens": 2000
}'
./start.sh restart           # restart, e.g. after changing a setting
./stop.sh                    # stop, and save the logs under logs/
docker logs -f glm53-tf      # follow the server
curl -s http://127.0.0.1:8020/health

If start.sh stops at a check

Every failed check prints a code. The ones you are most likely to see:

CodeWhat to do
E_GPU_BUSY / E_GPU_MEMORYOther work is on the GPUs: stop it. Each GPU needs at least 90,000 MiB free.
E_P2PCUDA peer access failed: check the driver and the platform's PCIe topology, ACS and IOMMU settings. Every GPU pair must support peer access.
E_DOCKER_GROUPAdd your user to the docker group and log in again.
W_POWERThe power limit isn't 250 W, so results may differ. The script never changes it.
E_AUTOTUNERead the compiler error, leave the GPUs idle and rerun. Incomplete tuning is never marked ready.

The full list, with every code, is in the README.

Images and video

It can see, too.

Vision is on by default. Send images and videos as OpenAI-style content parts in a user message: an image_url part, or a video_url part for a clip. The vision tower runs on rank 0.

InputContent part
Image{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
Video{"type": "video_url", "video_url": {"url": "data:video/mp4;base64,..."}}

Only data URLs are accepted; outside media URLs are off. VISION=0 serves text only.

Memory

A million tokens a request, forty requests.

All requests share one FP8 cache pool. Any one request can use the full 1,048,576-token window, but forty full windows at once are not promised. /health shows pool_tokens and pool_free_tokens, and the engine refuses a window that doesn't fit instead of shrinking it.

SettingDefault
Concurrent requests (PARALLEL)40
Prompt plus reply (CONTEXT)1,048,576 tokens a request
KV cache (KV)FP8
Memory reserve / KV cap7.5 GiB / 60 GiB a GPU

Settings

The ones you might change.

Settings come from the environment, then scripts/local.sh, then .env, then the defaults in scripts/config.sh. Set one for a single run (DENSE=q4 ./start.sh restart), or keep it in scripts/local.sh. The full list is in the README.

VariableDefaultMeaning
PARALLEL40Concurrent requests, 1 to 64 with DFlash2 (1 with MTP)
CONTEXT1048576Prompt plus reply tokens per request
DENSEfp8fp8, q4 (the speed option) or the checkpoint's bf16 dense layers
KVfp8KV cache; bf16 needs more memory
DRAFTERdflash2Inco AI's DFlash2 drafter (non-commercial use only), or mtp, the checkpoint's own MTP head, one request at a time
THINKING / REASONING_EFFORT0 / highThinking off by default, and its effort when a request turns it on
MAX_TOKENS32768The reply budget when a request doesn't set one
VISION1Image and video input
GPUS / PORT / HOST0,1,2,3 / 8020 / 0.0.0.0The four GPUs, and the API's port and address

The API

/v1/chat/completions with the model glm-5.3-flash, /v1/models, /tokenize and /detokenize, /health and Prometheus /metrics. Tool calling (complete calls streamed, arguments typed by their schema) and structured outputs (json_object or json_schema, enforced with xgrammar). A sampled request without top_k is served with 20; send top_k to compare with another engine.

Exact replies

The checks compare drafted replies with serial ones, concurrent requests with solo ones, and long prompts with saved token hashes, all within one fixed configuration. FP8 dense layers, the FP8 KV cache and chunked KDA change the arithmetic, so passing the checks doesn't mean the output matches BF16 or the original model; the quality numbers above are measured for that.

Licenses

Free, with a few rules.

  • The recipe, its code and documentation, is Apache License 2.0. Upstream code and model weights keep their own licenses.
  • The checkpoint, Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold, is under the Apache License 2.0.
  • The DFlash2 drafter is CC BY-NC-ND 4.0, non-commercial use only, and is not redistributed by the recipe. DRAFTER=mtp serves without it.
  • TensorFold is Apache 2.0; its earlier code keeps its MIT notices.
  • The base model, GLM-5.3-Flash, is by Z.ai.

Credits

Built together, on the work of others.

  • Marc Seal and Aevonix Research: the single-host release and its 32 patches
  • My lab: the EXL3 checkpoint, the original recipe and 55 of its patches, and sparkDash
  • TensorFold by Ash Hart
  • GLM-5.3-Flash by Z.ai
  • The DFlash2 drafter by Inco AI
  • The EXL3 format (ExLlamaV3) by turboderp
  • b12x by local-inference-lab, and glm53-tensorfold-spark by Jay Leaton

The full list is in the recipe's CREDITS.md.

Run it on your four GPUs.

Free and open, with every patch and every number published.