Recipe · 2 or 3 DGX Sparks · Mia's favourite
GLM 5.3 Flash on two DGX Sparks, with TensorFold.
Serve GLM-5.3-Flash from two NVIDIA DGX Sparks through an OpenAI-compatible API, with 4 concurrent requests, the model's full 1,048,576-token context and image and video input, on my own EXL3 quant, closer to the original model. One command sets up both Sparks and starts the server, and a third Spark takes it to 77.6 tok/s.
The checkpoint
My own EXL3 quant, closer to the original.
Since v1.4 the recipe serves Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold, my own EXL3 quantization of GLM-5.3-Flash: routed experts at 4 bits a weight, BF16 elsewhere, built for TensorFold. Same format, size, speed and memory as the widely used TR3-4bpw quant it replaces; only the calibration is new, and it is closer to the original model on every test set.
Against TR3-4bpw
Both quants on the same build, served by TensorFold with its own 4-bit dense weights.
| TR3-4bpw | My quant | |
|---|---|---|
| KL to the original, as served (wiki / workload) | 0.0990 / 0.3271 | 0.0903 / 0.3150 |
| GSM8K (250, greedy) | 98.0% | 98.8% |
| HumanEval (164, greedy) | 97.6% | 95.7% |
| HumanEval+ / MBPP+ (542, thinking on) | 86.3% | 86.5% |
| Size | 175.7 GB | 175.7 GB |
The benchmark gaps are a few problems each and go both ways; none is statistically significant. The measured gain is the fidelity (KL divergence 2-18% lower as served).
What that means
- Less drift from the original: KL divergence measures how far a quant's next-token probabilities drift from the unquantized model's. This one lowers it on all six test sets.
- Fewer confident mistakes: where the original model is sure of its next token, this quant picks a different one 16-37% less often (experts only).
- Same coding, shorter replies: 469 of 542 EvalPlus problems against TR3's 468, with about 10% shorter replies.
- Apache 2.0, free to use.
Performance
Measured, not promised.
Two DGX Sparks at the default configuration (4 streams, 1,048,576-token window, FP8 KV cache, 4-bit dense weights, DFlash2 plus copy drafts, vision on), GPU clocks capped at 2,200 MHz. Decode and prefill measured with sparkDash through the OpenAI API, from another machine on the network.
Decode
Aggregate across the concurrent requests, per request, and time to first token.
| Requests | Prose | Per request | TTFT | Structured | Per request | TTFT |
|---|---|---|---|---|---|---|
| 1 | 60.4 | 60.4 | 170 ms | 114.7 | 114.7 | 149 ms |
| 2 | 79.2 | 40.4 | 269 ms | 147.6 | 77.1 | 295 ms |
| 3 | 89.5 | 30.6 | 330 ms | 196.3 | 67.4 | 317 ms |
| 4 | 108.8 | 27.9 | 340 ms | 227.9 | 61.1 | 415 ms |
tok/s. Replies served 4 at a time are identical to the same requests served one at a time (11 of 11 cases staggered, and 11 of 11 sent in a burst).
Prefill
How fast a prompt is read before the reply starts.
| Prompt | Prefill | First token |
|---|---|---|
| 8,219 tokens | 1,952.2 tok/s | 4.21 s |
| 16,407 tokens | 1,973.5 tok/s | 8.31 s |
| 32,790 tokens | 1,978.9 tok/s | 16.57 s |
| 65,563 tokens | 1,942.5 tok/s | 33.75 s |
| 131,099 tokens | 1,837.9 tok/s | 71.33 s |
| 262,170 tokens | 1,641.7 tok/s | 159.69 s |
| 981,841 tokens needle in a haystack | 1,015 tok/s | 967 s, needle found |
Prompt reuse
The server resumes from a kept prompt state instead of prefilling again, with the same reply.
| Prompt | First time | Next time |
|---|---|---|
| An identical 64k-token prompt, sent again | 34 s | under 0.07 s |
| A new conversation with the same 7.9k-token system prompt | 4.24 s | 0.13 s |
Quality
My quant with the FP8 KV cache and 4-bit dense weights (the defaults); against TR3-4bpw above.
Speed preview
What 60.4 tok/s feels like.
Sample text played at the measured decode speed, at about four characters a token. On two Sparks: one request at 60.4 tok/s, or four at once at 27.9 tok/s each (108.8 combined). On three Sparks: 77.6 tok/s, or 37.3 each (146.2 combined). Switch between them in the window.
3 Sparks · experimental
Got a third Spark? 77.6 tok/s.
./start-tp3.sh runs the same recipe as tensor parallel over three Sparks. TensorFold v0.6.0 itself serves two ranks only; the engine for three comes from this recipe's own patches (0066-0068), in the same published image. Tested: concurrent replies equal replies served one at a time, drafted replies equal serial ones, the 195k needle found, images and tool calls whole.
Decode on three Sparks
Measured with sparkDash, the default configuration otherwise (4 streams, 1M window, FP8 KV cache, DFlash2 plus copy drafts, vision on).
| Requests | Prose | Per request | TTFT | Code | Per request | TTFT |
|---|---|---|---|---|---|---|
| 1 | 77.6 | 77.6 | 144 ms | 104.3 | 104.3 | 192 ms |
| 2 | 95.7 | 49.5 | 264 ms | 139.0 | 72.7 | 293 ms |
| 3 | 121.2 | 42.0 | 249 ms | 158.7 | 55.5 | 379 ms |
| 4 | 146.2 | 37.3 | 262 ms | 169.6 | 45.7 | 348 ms |
tok/s. Prefill is about the same as on two Sparks (1-4% faster): 2,000.6 tok/s at 8k, 2,064.2 at 32k, 1,654.5 at 262k.
How to wire it
- Cabling: a triangle, one direct QSFP cable per pair of Sparks, each cable its own subnet. Wire it as a directed ring, each Spark's port 0 to the next Spark's port 1.
- Settings: in
scripts/local.sh,WORKER(rank 1) andWORKER2(rank 2), withFABRIC_PEER/FABRIC_PEER2if needed. - Check first:
DRY_RUN=1 ./start-tp3.shshows every rank's command and the links it found, without starting anything. - Memory: the tested runs used
KV_POOL_GIB=27; at the default 32, the head had less free memory than this recipe keeps under load.
The engine
Powered by TensorFold.
TensorFold, by Ash Hart, is an open-source engine that serves language models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API. Each model family brings its own kernels and its own draft verification, tuned for that model instead of one generic path for all.
Exact decoding
Drafts make it fast without changing a word: a draft is accepted only when it equals the token the same engine would produce one at a time. Send a request with "draft": false to compare.
Apple Silicon and NVIDIA
MLX on Macs, and CUDA on NVIDIA GPUs from compute capability 8.9: Ada (RTX 40), Hopper and Blackwell, including the DGX Spark's GB10 and the RTX 50 series.
Many model families
GLM-5.3-Flash, Qwen3.8-27B, Qwen3.8 Flash Next, Nemotron 3.5 Lightning, Gemma 4 and DeepSeek-V4-Flash, among others, each with its own kernels.
In this recipe
TensorFold v0.6.0 runs one rank on each Spark, with 68 patches on top: DFlash2 and copy drafts, 4-bit dense weights, an FP8 KV cache, faster prompt kernels, a one-shot RoCE all-gather between the Sparks, four requests over one shared cache pool, and three patches that run it on a third Spark.
What you need
Two Sparks, one cable.
Two DGX Sparks
Or two GB10 systems with 128 GB unified memory, with nothing else large on their GPUs: each needs about 110 GiB free memory when the server starts (start.sh warns below that).
A direct ConnectX-7 link
A QSFP cable between the CX7 ports and an IPv4 address on each end in one private subnet (ping must work), with a RoCE v2 GID. With both ports cabled and addressed, both are used: a prompt chunk's all-gather is about 1.8x faster on two rails.
Key-based ssh
From the first Spark (the head, which runs ./start.sh and the API) to the second (the worker): ssh-copy-id user@<worker>, then check with ssh -o BatchMode=yes user@<worker> true.
Docker and disk
Docker with the NVIDIA container runtime, your user in the docker group, and rsync, on both. About 205 GB on each Spark: about 176 GB for the checkpoint, 2.3 GB for DFlash2 and 25 GB for the image.
Quick start
One command, on the head.
Clone and point it at the worker
git clone https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold.git
cd GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold
cp scripts/local.sh.example scripts/local.sh # set [email protected] in it
WORKER is the worker's ssh target. If it's on another network than the link, also set FABRIC_PEER to the worker's CX7 address. The worker needs no copy of the repository.
Start it
./start.sh
The first run sets up both Sparks: the image (about 25 GB) on each, the checkpoint (about 176 GB) and DFlash2 downloaded on the head and copied to the worker, then the CUDA kernels compile once. Later starts take 2 to 6 minutes. It runs a smoke test through both ranks and prints GLM-5.3-Flash-EXL3 is now LIVE! on port 8888.

Talk to it
Any OpenAI client works with base_url = "http://<head-address>:8888/v1" and the model GLM-5.3-Flash-EXL3. The model thinks before it answers (reasoning_content), so give replies enough max_tokens.
curl -s http://<head-address>:8888/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "GLM-5.3-Flash-EXL3",
"messages": [{"role": "user", "content": "Write a Python fibonacci function."}],
"max_tokens": 2000
}'
./start.sh restart # restart both ranks, e.g. after changing a setting
./stop.sh # stop both ranks and free their GPU memory
curl -s http://<head-address>:8888/health # busy flag, streams, free pool tokens
If start.sh stops at a check
| Message | What to do |
|---|---|
only N GiB memory available | Other GPU work runs on that Spark: stop it (docker ps) and restart. |
no RoCE device or RoCE v2 GID for the link / no route from this node to ... | The route to the worker doesn't go over the CX7 port: set FABRIC_PEER to the worker's CX7 address and check both ends are addressed (ip -4 addr). |
the worker's .../hub is not writable | A container left it root-owned: fix its ownership on the worker. |
Images and video
It can see, too.
GLM's own vision tower runs on the head. Send images and videos as OpenAI-style content parts in a user message: an image_url part, or a video_url part for a clip.
| Images | Videos | |
|---|---|---|
| Formats | JPEG, PNG, WebP | MP4, WebM, MOV, MKV (anything FFmpeg decodes) |
| Per request | Up to 50, 10 MB each, 64 MB in all | Up to 4, 64 MB each, 96 MB in all, up to an hour of footage each |
| Tokens | At most 2,048 a picture (a 1080p picture takes 2,040); a request's pictures share 16,384 | 2 frames a second, at most 128 frames a clip, at most 16,384 tokens a clip |
By default only data URLs are accepted; VISION_URLS=1 also lets the server fetch public https:// URLs. VISION=0 serves text only.
Memory
A million tokens, four at a time.
All requests draw their per-token caches from one shared FP8 pool: about 2.9M tokens (2,922,496 at the measured start; about 2.1M to 2.9M depending on free memory). Any one request can grow to the full 1M window: for example one 1M-token conversation next to another of 1M, or next to three of about 640k. A request the pool can't place yet waits until others finish.
| Setting | Window | Note |
|---|---|---|
Default (KV=fp8) | 1,048,576 | 4 requests |
KV=bf16 | 196,608 | The exact cache |
KV=bf16 DRAFTER=mtp | 524,288 | One request at a time |
CONTEXT=0 | The largest that fits | No memory left to keep other conversations' prompts |
Settings
The ones you might change.
Every setting lives in scripts/config.sh. Set one for a single run from the environment (PARALLEL=2 ./start.sh restart), or keep it in scripts/local.sh or a .env file. The full list is in the README.
| Variable | Default | Meaning |
|---|---|---|
PARALLEL | 4 | Requests decoded together, 1 to 4 (above 1 needs DRAFTER=dflash2) |
CONTEXT | 1048576 | Prompt + reply window per request; 0: the largest that fits |
KV | fp8 | fp8, or bf16 (exact, shorter window) |
DENSE | q4 | The checkpoint's BF16 weights as 4-bit, fp8 or bf16 |
DRAFTER | dflash2 | IncoAI's DFlash2 drafter (non-commercial use only, +5-10% decode), or mtp, the checkpoint's own MTP head, one request at a time, which avoids that license |
THINKING | 1 | Think before answering by default; 0 answers directly unless a request asks to think |
VISION / VISION_URLS | 1 / 0 | Image and video input; 1 also accepts public https:// URLs |
TP / WORKER2 | 2 / empty | Sparks in all (3 through ./start-tp3.sh), and rank 2's ssh target |
SERVED_NAME / PORT | GLM-5.3-Flash-EXL3 / 8888 | The model id in /v1/models and replies, and the API's port |
The API
/v1/chat/completions, /v1/completions, /v1/responses, /v1/models, /tokenize, /health and Prometheus /metrics. Tool calling (streamed, arguments typed by their schema), structured outputs with xgrammar, and reasoning_effort low / high / max.
Exact replies
Drafts only propose: every drafted token is checked against the model's own sample, so drafted replies equal TensorFold's serial, one-token-at-a-time decoding. Three defaults trade exactness for speed and the 1M window: 4-bit dense weights and the FP8 KV cache are lossy (quality above), and the chunked KDA kernel is close to the serial one but not its bits. DENSE=bf16 KV=bf16 KDA_CHUNKED=0 serves the checkpoint as it is.
Licenses
Free, with a few rules.
- The recipe is Apache License 2.0. Its patches modify TensorFold v0.6.0, whose code stays under TensorFold's licenses (Apache 2.0 from v0.6.0, and the MIT notice of earlier code).
- The checkpoint, Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold, is under the Apache License 2.0.
- TR3-4bpw (
MODEL_ID=Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw, the default before v1.3.3) is under the ShapleyMcg License 1.0, an attribution-required license. Its notice:This work includes or was produced using ShapleyMcg, created by Brandon M. Music (https://github.com/brandonmmusic-max/shapleymcg). ShapleyMcg is licensed under the ShapleyMcg License v1.0, an attribution-required license that grants no rights to the person known as "0xSero." Use of ShapleyMcg without this attribution is unlicensed.
- The DFlash2 drafter is CC BY-NC-ND 4.0, non-commercial use only (commercial licensing: [email protected]).
DRAFTER=mtpserves without it. - The base model, GLM-5.3-Flash, is MIT-licensed by Z.AI.
- The image is based on NVIDIA's PyTorch container; by pulling or running it you accept NVIDIA's terms, which the container prints at every start.
Credits
Built on the work of others.
- TensorFold by Ash Hart
- GLM-5.3-Flash by Z.ai
- The EXL3 format and converter (exllamav3) by turboderp
- The earlier TR3 quantization by Brandon M. Music (ShapleyMcg)
- The DFlash2 drafter by IncoAI
- b12x's RoCE transport by local-inference-lab
- Code from glm53-tensorfold-spark by Jay Leaton (tool calling, L2 prefetch, expert loads)
The full list, including the runtime stack, is in the recipe's CREDITS.md.
Run it on your Sparks.
Free and open, with every patch and every number published.