Recipe · 2 or 3 DGX Sparks · Mia's favourite

GLM 5.3 Flash on two DGX Sparks, with TensorFold.

Serve GLM-5.3-Flash from two NVIDIA DGX Sparks through an OpenAI-compatible API, with 4 concurrent requests, the model's full 1,048,576-token context and image and video input, on my own EXL3 quant, closer to the original model. One command sets up both Sparks and starts the server, and a third Spark takes it to 77.6 tok/s.

Get the free recipe See the launch on X

Install with an AI agent

Install this model

Paste this into an AI coding agent that can run commands on the machine that will serve the model (the head node, for a Spark cluster), such as Claude Code or Codex. It reads the recipe, adapts it to your setup, starts the server and checks that it works.


        
GLM 5.3 Flash EXL3 on TensorFold, DGX Sparks

The checkpoint

My own EXL3 quant, closer to the original.

Since v1.4 the recipe serves Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold, my own EXL3 quantization of GLM-5.3-Flash: routed experts at 4 bits a weight, BF16 elsewhere, built for TensorFold. Same format, size, speed and memory as the widely used TR3-4bpw quant it replaces; only the calibration is new, and it is closer to the original model on every test set.

KL, wiki0.0903TR3-4bpw: 0.0990 (lower is better)
KL, workload0.3150TR3-4bpw: 0.3271
Confident mistakes16-37%fewer, experts only
Replies~10%shorter on coding tests

Against TR3-4bpw

Both quants on the same build, served by TensorFold with its own 4-bit dense weights.

TR3-4bpwMy quant
KL to the original, as served (wiki / workload)0.0990 / 0.32710.0903 / 0.3150
GSM8K (250, greedy)98.0%98.8%
HumanEval (164, greedy)97.6%95.7%
HumanEval+ / MBPP+ (542, thinking on)86.3%86.5%
Size175.7 GB175.7 GB

The benchmark gaps are a few problems each and go both ways; none is statistically significant. The measured gain is the fidelity (KL divergence 2-18% lower as served).

What that means

  • Less drift from the original: KL divergence measures how far a quant's next-token probabilities drift from the unquantized model's. This one lowers it on all six test sets.
  • Fewer confident mistakes: where the original model is sure of its next token, this quant picks a different one 16-37% less often (experts only).
  • Same coding, shorter replies: 469 of 542 EvalPlus problems against TR3's 468, with about 10% shorter replies.
  • Apache 2.0, free to use.

Performance

Measured, not promised.

One request60.4tok/s, prose
Four at once108.8tok/s combined
Context1Mtokens a request
Prefill1,978.9tok/s, 32k prompt

Two DGX Sparks at the default configuration (4 streams, 1,048,576-token window, FP8 KV cache, 4-bit dense weights, DFlash2 plus copy drafts, vision on), GPU clocks capped at 2,200 MHz. Decode and prefill measured with sparkDash through the OpenAI API, from another machine on the network.

Decode

Aggregate across the concurrent requests, per request, and time to first token.

RequestsProsePer requestTTFTStructuredPer requestTTFT
160.460.4170 ms114.7114.7149 ms
279.240.4269 ms147.677.1295 ms
389.530.6330 ms196.367.4317 ms
4108.827.9340 ms227.961.1415 ms

tok/s. Replies served 4 at a time are identical to the same requests served one at a time (11 of 11 cases staggered, and 11 of 11 sent in a burst).

Prefill

How fast a prompt is read before the reply starts.

PromptPrefillFirst token
8,219 tokens1,952.2 tok/s4.21 s
16,407 tokens1,973.5 tok/s8.31 s
32,790 tokens1,978.9 tok/s16.57 s
65,563 tokens1,942.5 tok/s33.75 s
131,099 tokens1,837.9 tok/s71.33 s
262,170 tokens1,641.7 tok/s159.69 s
981,841 tokens needle in a haystack1,015 tok/s967 s, needle found

Prompt reuse

The server resumes from a kept prompt state instead of prefilling again, with the same reply.

PromptFirst timeNext time
An identical 64k-token prompt, sent again34 sunder 0.07 s
A new conversation with the same 7.9k-token system prompt4.24 s0.13 s

Quality

My quant with the FP8 KV cache and 4-bit dense weights (the defaults); against TR3-4bpw above.

98.8%GSM8K, 250 problems, thinking off
95.7%HumanEval, 164 problems, thinking off
86.5%HumanEval+ and MBPP+, 542 problems, thinking on

Speed preview

What 60.4 tok/s feels like.

Sample text played at the measured decode speed, at about four characters a token. On two Sparks: one request at 60.4 tok/s, or four at once at 27.9 tok/s each (108.8 combined). On three Sparks: 77.6 tok/s, or 37.3 each (146.2 combined). Switch between them in the window.

TensorFold · GLM-5.3-Flash-EXL3 · 2× DGX Spark
60.4 tok/sSample text, played at the measured speed

3 Sparks · experimental

Got a third Spark? 77.6 tok/s.

./start-tp3.sh runs the same recipe as tensor parallel over three Sparks. TensorFold v0.6.0 itself serves two ranks only; the engine for three comes from this recipe's own patches (0066-0068), in the same published image. Tested: concurrent replies equal replies served one at a time, drafted replies equal serial ones, the 195k needle found, images and tool calls whole.

One request77.6tok/s, prose (two Sparks: 60.4)
Four at once146.2tok/s combined (two: 108.8)
Code104.3tok/s, one request
Prefill2,064tok/s, 32k prompt

Decode on three Sparks

Measured with sparkDash, the default configuration otherwise (4 streams, 1M window, FP8 KV cache, DFlash2 plus copy drafts, vision on).

RequestsProsePer requestTTFTCodePer requestTTFT
177.677.6144 ms104.3104.3192 ms
295.749.5264 ms139.072.7293 ms
3121.242.0249 ms158.755.5379 ms
4146.237.3262 ms169.645.7348 ms

tok/s. Prefill is about the same as on two Sparks (1-4% faster): 2,000.6 tok/s at 8k, 2,064.2 at 32k, 1,654.5 at 262k.

How to wire it

  • Cabling: a triangle, one direct QSFP cable per pair of Sparks, each cable its own subnet. Wire it as a directed ring, each Spark's port 0 to the next Spark's port 1.
  • Settings: in scripts/local.sh, WORKER (rank 1) and WORKER2 (rank 2), with FABRIC_PEER / FABRIC_PEER2 if needed.
  • Check first: DRY_RUN=1 ./start-tp3.sh shows every rank's command and the links it found, without starting anything.
  • Memory: the tested runs used KV_POOL_GIB=27; at the default 32, the head had less free memory than this recipe keeps under load.

The engine

Powered by TensorFold.

TensorFold, by Ash Hart, is an open-source engine that serves language models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API. Each model family brings its own kernels and its own draft verification, tuned for that model instead of one generic path for all.

Exact decoding

Drafts make it fast without changing a word: a draft is accepted only when it equals the token the same engine would produce one at a time. Send a request with "draft": false to compare.

Apple Silicon and NVIDIA

MLX on Macs, and CUDA on NVIDIA GPUs from compute capability 8.9: Ada (RTX 40), Hopper and Blackwell, including the DGX Spark's GB10 and the RTX 50 series.

Many model families

GLM-5.3-Flash, Qwen3.8-27B, Qwen3.8 Flash Next, Nemotron 3.5 Lightning, Gemma 4 and DeepSeek-V4-Flash, among others, each with its own kernels.

In this recipe

TensorFold v0.6.0 runs one rank on each Spark, with 68 patches on top: DFlash2 and copy drafts, 4-bit dense weights, an FP8 KV cache, faster prompt kernels, a one-shot RoCE all-gather between the Sparks, four requests over one shared cache pool, and three patches that run it on a third Spark.

What you need

Two Sparks, one cable.

Two DGX Sparks

Or two GB10 systems with 128 GB unified memory, with nothing else large on their GPUs: each needs about 110 GiB free memory when the server starts (start.sh warns below that).

A direct ConnectX-7 link

A QSFP cable between the CX7 ports and an IPv4 address on each end in one private subnet (ping must work), with a RoCE v2 GID. With both ports cabled and addressed, both are used: a prompt chunk's all-gather is about 1.8x faster on two rails.

Key-based ssh

From the first Spark (the head, which runs ./start.sh and the API) to the second (the worker): ssh-copy-id user@<worker>, then check with ssh -o BatchMode=yes user@<worker> true.

Docker and disk

Docker with the NVIDIA container runtime, your user in the docker group, and rsync, on both. About 205 GB on each Spark: about 176 GB for the checkpoint, 2.3 GB for DFlash2 and 25 GB for the image.

Quick start

One command, on the head.

1

Clone and point it at the worker

git clone https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold.git
cd GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold
cp scripts/local.sh.example scripts/local.sh   # set [email protected] in it

WORKER is the worker's ssh target. If it's on another network than the link, also set FABRIC_PEER to the worker's CX7 address. The worker needs no copy of the repository.

2

Start it

./start.sh

The first run sets up both Sparks: the image (about 25 GB) on each, the checkpoint (about 176 GB) and DFlash2 downloaded on the head and copied to the worker, then the CUDA kernels compile once. Later starts take 2 to 6 minutes. It runs a smoke test through both ranks and prints GLM-5.3-Flash-EXL3 is now LIVE! on port 8888.

start.sh banner: TensorFold ribbon and MIA AI LAB, GLM-5.3 Flash EXL3 · Dual DGX Sparks
3

Talk to it

Any OpenAI client works with base_url = "http://<head-address>:8888/v1" and the model GLM-5.3-Flash-EXL3. The model thinks before it answers (reasoning_content), so give replies enough max_tokens.

curl -s http://<head-address>:8888/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "GLM-5.3-Flash-EXL3",
  "messages": [{"role": "user", "content": "Write a Python fibonacci function."}],
  "max_tokens": 2000
}'
./start.sh restart       # restart both ranks, e.g. after changing a setting
./stop.sh                # stop both ranks and free their GPU memory
curl -s http://<head-address>:8888/health   # busy flag, streams, free pool tokens

If start.sh stops at a check

MessageWhat to do
only N GiB memory availableOther GPU work runs on that Spark: stop it (docker ps) and restart.
no RoCE device or RoCE v2 GID for the link / no route from this node to ...The route to the worker doesn't go over the CX7 port: set FABRIC_PEER to the worker's CX7 address and check both ends are addressed (ip -4 addr).
the worker's .../hub is not writableA container left it root-owned: fix its ownership on the worker.

Images and video

It can see, too.

GLM's own vision tower runs on the head. Send images and videos as OpenAI-style content parts in a user message: an image_url part, or a video_url part for a clip.

ImagesVideos
FormatsJPEG, PNG, WebPMP4, WebM, MOV, MKV (anything FFmpeg decodes)
Per requestUp to 50, 10 MB each, 64 MB in allUp to 4, 64 MB each, 96 MB in all, up to an hour of footage each
TokensAt most 2,048 a picture (a 1080p picture takes 2,040); a request's pictures share 16,3842 frames a second, at most 128 frames a clip, at most 16,384 tokens a clip

By default only data URLs are accepted; VISION_URLS=1 also lets the server fetch public https:// URLs. VISION=0 serves text only.

Memory

A million tokens, four at a time.

All requests draw their per-token caches from one shared FP8 pool: about 2.9M tokens (2,922,496 at the measured start; about 2.1M to 2.9M depending on free memory). Any one request can grow to the full 1M window: for example one 1M-token conversation next to another of 1M, or next to three of about 640k. A request the pool can't place yet waits until others finish.

SettingWindowNote
Default (KV=fp8)1,048,5764 requests
KV=bf16196,608The exact cache
KV=bf16 DRAFTER=mtp524,288One request at a time
CONTEXT=0The largest that fitsNo memory left to keep other conversations' prompts

Settings

The ones you might change.

Every setting lives in scripts/config.sh. Set one for a single run from the environment (PARALLEL=2 ./start.sh restart), or keep it in scripts/local.sh or a .env file. The full list is in the README.

VariableDefaultMeaning
PARALLEL4Requests decoded together, 1 to 4 (above 1 needs DRAFTER=dflash2)
CONTEXT1048576Prompt + reply window per request; 0: the largest that fits
KVfp8fp8, or bf16 (exact, shorter window)
DENSEq4The checkpoint's BF16 weights as 4-bit, fp8 or bf16
DRAFTERdflash2IncoAI's DFlash2 drafter (non-commercial use only, +5-10% decode), or mtp, the checkpoint's own MTP head, one request at a time, which avoids that license
THINKING1Think before answering by default; 0 answers directly unless a request asks to think
VISION / VISION_URLS1 / 0Image and video input; 1 also accepts public https:// URLs
TP / WORKER22 / emptySparks in all (3 through ./start-tp3.sh), and rank 2's ssh target
SERVED_NAME / PORTGLM-5.3-Flash-EXL3 / 8888The model id in /v1/models and replies, and the API's port

The API

/v1/chat/completions, /v1/completions, /v1/responses, /v1/models, /tokenize, /health and Prometheus /metrics. Tool calling (streamed, arguments typed by their schema), structured outputs with xgrammar, and reasoning_effort low / high / max.

Exact replies

Drafts only propose: every drafted token is checked against the model's own sample, so drafted replies equal TensorFold's serial, one-token-at-a-time decoding. Three defaults trade exactness for speed and the 1M window: 4-bit dense weights and the FP8 KV cache are lossy (quality above), and the chunked KDA kernel is close to the serial one but not its bits. DENSE=bf16 KV=bf16 KDA_CHUNKED=0 serves the checkpoint as it is.

Licenses

Free, with a few rules.

  • The recipe is Apache License 2.0. Its patches modify TensorFold v0.6.0, whose code stays under TensorFold's licenses (Apache 2.0 from v0.6.0, and the MIT notice of earlier code).
  • The checkpoint, Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold, is under the Apache License 2.0.
  • TR3-4bpw (MODEL_ID=Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw, the default before v1.3.3) is under the ShapleyMcg License 1.0, an attribution-required license. Its notice: This work includes or was produced using ShapleyMcg, created by Brandon M. Music (https://github.com/brandonmmusic-max/shapleymcg). ShapleyMcg is licensed under the ShapleyMcg License v1.0, an attribution-required license that grants no rights to the person known as "0xSero." Use of ShapleyMcg without this attribution is unlicensed.
  • The DFlash2 drafter is CC BY-NC-ND 4.0, non-commercial use only (commercial licensing: [email protected]). DRAFTER=mtp serves without it.
  • The base model, GLM-5.3-Flash, is MIT-licensed by Z.AI.
  • The image is based on NVIDIA's PyTorch container; by pulling or running it you accept NVIDIA's terms, which the container prints at every start.

Credits

Built on the work of others.

The full list, including the runtime stack, is in the recipe's CREDITS.md.

Run it on your Sparks.

Free and open, with every patch and every number published.