Recipe · 3 DGX Sparks
The full GLM 5.3, on three DGX Sparks.
Serve Z.ai's full GLM-5.3 (78 layers of MLA with DeepSeek sparse attention, 256 routed experts, an MTP head) from three NVIDIA DGX Sparks through an OpenAI-compatible API, on my own 2.75-bit EXL3 quant. A 499,712-token context by default, a 314k-token prompt read and answered correctly, and every drafted reply equal to the serial one. One command sets up all three Sparks and starts the server.
The checkpoint
My own EXL3 quant, a third on each Spark.
The recipe serves Mia-AiLab/GLM-5.3-EXL3-2.75bpw-TensorFold, my own EXL3 quantization of GLM-5.3, calibrated for how TensorFold serves it: routed experts at 2.75 bits a weight on average, each expert at its own width from 2 to 4 bits, BF16 elsewhere.
How it is split
- Attention heads 22 / 21 / 21 across the three Sparks.
- Every expert in 16 blocks of 128 columns, 6 / 5 / 5, the extra block rotating by layer, so each Spark holds a third of the 233 GiB of routed experts.
- Mixed-width experts load through TensorFold's universal grouped kernels, one device buffer a layer.
- The workers keep no copy: each reads its share of every tensor from the head's cache over a read-only NFS export.
The model
GLM-5.3 is Z.ai's full model, not the Flash one: multi-head latent attention with DeepSeek-V3.2's sparse-attention indexer (the top 2,048 tokens), 256 routed experts and an MTP head. It thinks before it answers.
Performance
Measured, not promised.
Three DGX Sparks in the default long-context mode (CP=1, 499,712-token window, fp4 KV cache, 3,072-row prompt chunks, 4-bit dense weights, DSpark plus copy drafts, RoCE all-gathers), one request at a time, on the release build (boot hcp, 2026-10-04). Decode and prefill measured with sparkDash through the OpenAI API, greedy; the 94k and 314k prompts with the recipe's tools/needle.py.
Decode
One request, greedy.
| Text | Decode | TTFT |
|---|---|---|
| Prose | 31.3 tok/s | 5 ms |
| Code | 41.7 tok/s | 412 ms |
Drafts land more often on predictable text: on worked arithmetic (thinking on) the drafter's acceptance was 87% and a request averaged ~45 tok/s, with bursts above 70. Sampled replies (temperature 1.0, top_p 0.95) decode slower, since fewer drafts survive: prose 22.9-24.0, code 25.7-29.4 tok/s (an earlier build of the same mode).
Prefill
How fast a prompt is read before the reply starts.
| Prompt | Prefill | First token |
|---|---|---|
| 4,120 tokens | 601.5 tok/s | 6.85 s |
| 8,218 tokens | 606.5 tok/s | 13.55 s |
| 16,407 tokens | 598.4 tok/s | 27.42 s |
| 32,792 tokens | 591.2 tok/s | 55.47 s |
| 94,317 tokens needle | 545 tok/s | 173.0 s, answered correctly |
| 314,225 tokens needle | 468 tok/s | 671.1 s, answered correctly |
Through the 314k-token prompt, the tightest Spark never had less than 5.26 GiB free.
Without context parallelism
CP=0 ./start.sh: a 163,840-token window with an FP8 KV cache. Ranges are three runs.
| Prose | Code | |
|---|---|---|
| Decode, greedy | 30.6-32.6 tok/s | 41.7-43.2 tok/s |
| Decode, sampled | 26.5-27.8 tok/s | 31.2-33.1 tok/s |
Prefill 630 / 679 / 660 tok/s at 8k / 16k / 32k tokens.
Several requests at once
PARALLEL=2 to 4 (with CP=0) decodes that many requests together, each with its own window. Greedy, FP8 KV.
| Setting | Together | Total | Each |
|---|---|---|---|
PARALLEL=2, 65,536 window | 2 | 37.8 tok/s | 18.9-20.0 |
PARALLEL=4, 32,768 window | 4 | 48.2 tok/s | 12.1-14.0 |
Each reply is the same alone and beside others (12/12). One request alone runs at about 27 tok/s on these servers, so PARALLEL pays only when requests really arrive together.
Speed preview
What 31.3 tok/s feels like.
Sample text played at the measured prose speed of one request, at about four characters a token. Code comes out faster, at 41.7 tok/s.
Long context
Half a million tokens, on a 4-bit cache.
With context parallelism (CP=1, the default) each Spark keeps every third token's caches, so the window is about three times what fits otherwise, and decode is as fast as without it. The KV cache is 4-bit (KV=fp4, 0.69x FP8's bytes a token), and on a paired quality suite it showed no measurable difference from FP8.
fp4 against FP8 KV
The same items on both, greedy, the recipe's tools/quality.py; exact two-sided McNemar test.
| Category | FP8 | FP4 | p |
|---|---|---|---|
| Short questions | 123/150 | 124/150 | 1.00 |
| Chained arithmetic (10 steps) | 40/40 | 39/40 | 1.00 |
| Ledger tracking (25 transfers) | 35/40 | 35/40 | 1.00 |
| Python tasks, hidden tests | 25/25 | 25/25 | 1.00 |
Every ledger miss in both runs was a reply that ran past the token limit; every finished reply was right. On fp4 the word problems scored 40/40.
Recall and reuse
- Recall at length: 16 keys, 4 of them corrected later, asked back at ~39k, ~157k and ~314k tokens: 112 of 112 keys right, none stale.
- Needles: found at 9.9k, 94k and 314k tokens.
- Prompt cache on NVMe (
TF_GLM_DISK_CACHE): kept prompt states also go to each Spark's own disk. A 94,317-token prompt that took 236.1 s the first time resumed in 3.0 s after a restart, with the right answer. - More room:
KV=fp4x(opt-in) holds ~24% more tokens a GiB; its long-context recall has not been measured yet.
The engine
Powered by TensorFold.
TensorFold, by Ash Hart, is an open-source engine that serves language models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API. Each model family brings its own kernels and its own draft verification, tuned for that model instead of one generic path for all.
Exact decoding
Drafts make it fast without changing a word: a draft is accepted only when it equals the token the same engine would produce one at a time. Send a request with "draft": false to compare.
Apple Silicon and NVIDIA
MLX on Macs, and CUDA on NVIDIA GPUs from compute capability 8.9: Ada (RTX 40), Hopper and Blackwell, including the DGX Spark's GB10 and the RTX 50 series.
Many model families
GLM-5.3-Flash, Qwen3.8-27B, Qwen3.8 Flash Next, Nemotron 3.5 Lightning, Gemma 4 and DeepSeek-V4-Flash, among others, each with its own kernels.
In this recipe
TensorFold v0.6.0 runs one rank on each Spark, with 110 patches on top: the GLM-5.3-Flash recipe's 68, and 42 for full GLM-5.3, among them the 3-rank layout, a loader for mixed-width EXL3 experts, our own prompt kernels, DSpark drafts, context parallelism, the 4-bit KV cache, 2 to 4 requests at once and the NVMe prompt cache.
What you need
Three Sparks, three cables.
Three DGX Sparks
GB10 with 128 GB unified memory each, with nothing else large on their GPUs: each rank holds ~86 GiB of weights, and start.sh plans every Spark's memory before it starts, refusing a start that would leave one under the floor.
A triangle of ConnectX-7 cables
One direct QSFP cable per pair of Sparks, each cable its own subnet, with RoCE v2: both CX7 ports of each Spark are used, each toward one peer. Plus one network all three share for ssh and the ranks' rendezvous.
Key-based ssh
From the head (rank 0, which runs ./start.sh and the API) to both workers: ssh-copy-id user@<worker>, then check with ssh -o BatchMode=yes user@<worker> true. Docker with the NVIDIA container runtime on each Spark.
Disk on the head, NFS to the workers
About 273 GiB for the checkpoint and 2.4 GiB for DSpark on the head, plus 100 GB left free. The head exports its Hugging Face cache read-only over NFS to both workers, which keep only the image (about 25 GB).
Quick start
One command, on the head.
Share the weights with the workers
Once, on the head (it needs root there; the workers need nothing installed), one entry per cable subnet:
sudo apt install nfs-kernel-server
echo "$HOME/.cache/huggingface 10.0.22.0/24(ro,no_subtree_check) 10.0.23.0/24(ro,no_subtree_check)" | sudo tee -a /etc/exports
sudo exportfs -ra
Clone and point it at the workers
git clone https://github.com/MiaAI-Lab/GLM-5.3-EXL3-3x-DGX-Sparks-TensorFold.git
cd GLM-5.3-EXL3-3x-DGX-Sparks-TensorFold
cp scripts/local.sh.example scripts/local.sh # WORKER=user@<rank 1>, WORKER2=user@<rank 2>, NFS settings
DRY_RUN=1 ./start.sh # the memory plan and every rank's docker command; nothing changes
If a worker is reached over another network than its cable, set FABRIC_PEER / FABRIC_PEER2 to its CX7 address. The workers need no copy of the repository.
Start it
./start.sh
The first run sets up all three Sparks: the image on each, the checkpoint (about 273 GiB) and DSpark downloaded on the head, the workers' NFS mounts checked, then the CUDA kernels compile once. Later starts take 3 to 5 minutes. It runs a smoke test through all three ranks and prints GLM-5.3-EXL3 is now LIVE! on port 8888.

Talk to it
Any OpenAI client works with base_url = "http://<head-address>:8888/v1" and the model GLM-5.3-EXL3. The model thinks before it answers (reasoning_content), so give replies enough max_tokens.
curl -s http://<head-address>:8888/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "GLM-5.3-EXL3",
"messages": [{"role": "user", "content": "Write a Python fibonacci function."}],
"max_tokens": 2000
}'
./start.sh restart # restart all three ranks, e.g. after changing a setting
CP=0 ./start.sh restart # without context parallelism: 163,840 tokens, FP8 KV
CP=0 PARALLEL=2 CONTEXT=65536 ./start.sh restart # two requests decoded together (up to 4; needs CP=0)
./stop.sh # stop all three ranks and free their GPU memory
Settings
The ones you might change.
Every setting lives in scripts/config.sh, with the reason for its default. Set one for a single run from the environment (CONTEXT=32768 ./start.sh restart), or keep it in scripts/local.sh or a .env file. The full list is in the README.
| Variable | Default | Meaning |
|---|---|---|
CP | 1 | Context parallelism: each Spark keeps every third token's caches (~3x the window, exact, one request at a time); 0: every Spark keeps every token |
CONTEXT | 499712 (163840 with CP=0) | Prompt + reply window, 4,096 to 1,048,576; the memory plan checks every start |
KV | fp4 (fp8 with CP=0) | fp4: 4-bit latent rows, the long-context setting; fp8; bf16: exact, 1.7x FP8's bytes; fp4x: opt-in, ~24% more context than fp4 |
PARALLEL | 1 | Requests decoded together, 1 to 4, each with its own window; needs CP=0 |
DRAFTER | dspark | Red Hat AI's DSpark speculator, up to 8 drafts a round; or mtp, the checkpoint's own MTP head |
DENSE | q4 | Non-expert weights as 4-bit groups; fp8 takes ~2.8 GiB more a rank |
THINKING | 1 | Think before answering by default |
TF_GLM_DISK_CACHE | off | A folder inside the container for the NVMe prompt cache, e.g. /cache/pcache (PARALLEL=1 only) |
SERVED_NAME / PORT | GLM-5.3-EXL3 / 8888 | The model id in /v1/models and replies, and the API's port |
Memory, watched
On a Spark the GPU and the CPU share one 128 GB pool, and a Spark that runs out of it freezes instead of failing. So start.sh plans every rank's memory with each Spark's free memory of the moment and refuses a start that would leave one under the floor, and a guard on every Spark stops that Spark's rank below 3 GiB free: a failed rank beats a frozen Spark.
Exact replies
Drafts only propose: every drafted token is checked against the model's own sample, so a drafted reply equals TensorFold's serial one (12/12 on every measured boot, tools/exact.py), and requests decoded together get the replies they get alone. Two defaults are lossy against the checkpoint in BF16, for speed and window: the 4-bit dense weights and the fp4 KV cache (its quality above).
Licenses
Free, with a few rules.
- The recipe is Apache License 2.0. Its patches modify TensorFold v0.6.0, whose code stays under TensorFold's licenses (Apache 2.0 from v0.6.0, and the MIT notice of earlier code).
- The checkpoint, Mia-AiLab/GLM-5.3-EXL3-2.75bpw-TensorFold, derives from Z.ai's GLM-5.3 and is under the GLM-5.3 License, which ships with it.
- The DSpark speculator, RedHatAI/GLM-5.3-speculator.dspark, is Red Hat AI's, under the glm-5.3 license its model card states.
- The image is based on NVIDIA's PyTorch container; by pulling or running it you accept NVIDIA's terms, which the container prints at every start.
Credits
Built on the work of others.
- TensorFold by Ash Hart and the TensorFold contributors
- drowzeys: the EXL3 prompt-expert kernels of patch 0106, and the context-parallel scheme of 0112-0113, from drowzeys/TensorFold
- GLM-5.3 by Z.ai
- Red Hat AI's DSpark speculator; the drafter is adapted from vllm-project/speculators and vLLM
- b12x's RoCE transport by local-inference-lab
- Code from glm53-tensorfold-spark by Jay Leaton (tool calling, L2 prefetch, expert loads)
The full list, including the runtime stack, is in the recipe's CREDITS.md.
Run it on your Sparks.
Free and open, with every patch and every number published.