Recipe · One DGX Spark
Qwen3.8 Flash Next on one DGX Spark, with TensorFold.
Serve Qwen3.8 Flash Next from a single DGX Spark through an OpenAI-compatible API, with 5 concurrent requests at the full 262,144-token context and image and video input. It runs TensorFold v0.6.1 with one patch on top: video input and up to 50 images, copy drafts, SSD read-ahead, and the first token before the next draft.
Performance
Measured, not promised.
One DGX Spark, the recipe's defaults on TensorFold v0.6.1 (5 streams × 262,144 tokens, int8 KV cache, n-gram tables read from SSD, image input on, MTP drafting), measured through the OpenAI API.
Decode, prose
Aggregate across the concurrent requests, per request, and time to first token.
| Requests | Aggregate | Per request | First token |
|---|---|---|---|
| 1 | 63.6 tok/s | 63.6 tok/s | 66 ms |
| 2 | 85.6 tok/s | 43.8 tok/s | 226 ms |
| 4 | 114.5 tok/s | 29.7 tok/s | 304 ms |
Decode, code
The same measurement, on code replies.
| Requests | Aggregate | Per request | First token |
|---|---|---|---|
| 1 | 96.9 tok/s | 96.9 tok/s | 206 ms |
| 2 | 133.9 tok/s | 67.0 tok/s | 210 ms |
| 4 | 166.4 tok/s | 44.8 tok/s | 292 ms |
Prefill
How fast a prompt is read before the reply starts.
| Prompt | Prefill | First token |
|---|---|---|
| 8,228 tokens | 2,421 tok/s | 3.40 s |
| 16,423 tokens | 2,468 tok/s | 6.66 s |
| 32,805 tokens | 2,462 tok/s | 13.33 s |
| 65,571 tokens | 2,391 tok/s | 27.43 s |
| 131,110 tokens | 2,180 tok/s | 60.13 s |
What v0.6.1 changed
Against the recipe's earlier tables (TensorFold v0.3.6.3, 4 streams, 4,096-row prompt chunks), prose decode is within -5% to +7% and the first token comes 12-57% sooner; prefill is 1-3% slower, about what the default 2,048-row pieces cost (they leave room for the vision tower).
A new question on a long shared system prompt now reuses it: a 31k-token one answered in 0.18 s instead of 14 s.
Against unpatched TensorFold v0.3.6.2, the recipe's patches took prefill from ~1,350-1,490 to ~2,340-2,480 tok/s, with every reply byte-identical.
Speed preview
What 63.6 tok/s feels like.
Sample text played at the measured decode speed: one request at 63.6 tok/s, or four at once at 29.7 tok/s each (114.5 combined), at about four characters a token. Switch between them in the window.
The engine
Powered by TensorFold.
TensorFold, by Ash Hart, is an open-source engine that serves language models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API. Each model family brings its own kernels and its own draft verification, tuned for that model instead of one generic path for all.
Exact decoding
Drafts make it fast without changing a word: a draft is accepted only when it equals the token the same engine would produce one at a time. Send a request with "draft": false to compare.
Apple Silicon and NVIDIA
MLX on Macs, and CUDA on NVIDIA GPUs from compute capability 8.9: Ada (RTX 40), Hopper and Blackwell, including the DGX Spark's GB10 and the RTX 50 series.
Many model families
Qwen3.8 Flash Next, Qwen3.8-27B, GLM-5.3-Flash, Nemotron 3.5 Lightning, Gemma 4 and DeepSeek-V4-Flash, among others, each with its own kernels.
In this recipe
TensorFold v0.6.1 with one patch: video input and up to 50 images on TensorFold's own Flash Next vision, copy drafts, SSD read-ahead around the n-gram table, and the first token before the next draft. Shared system prompts are reused instead of read again.
What you need
One Spark, nothing else.
A DGX Spark
Or another GB10 system with 128 GB unified memory, with nothing else large on the GPU: the default setting needs about 103 GiB free when the server starts.
Docker
Docker with the NVIDIA container runtime, and your user in the docker group.
Disk
About 160 GB free on a fresh machine: about 125 GB for the checkpoint download under ~/.cache/huggingface and about 35 GB for the image; scripts/prepare.sh checks both.
Optional
The hf CLI on the host (a faster, resumable download) and a Hugging Face token in ~/.cache/huggingface/token or HF_TOKEN.
Quick start
Three lines, that's all.
Clone and start
git clone https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold.git
cd Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold
./start.sh
The first run sets everything up: it pulls the prebuilt image (about 11 GB) and downloads the checkpoint, then compiles the CUDA kernels for the GB10 once (a few minutes). Later starts take about 2.5 minutes to load the weights. It runs a smoke test and prints Qwen3.8-Flash-Next is now LIVE! on port 8888.
Talk to it
Any OpenAI client works with base_url = "http://<spark-address>:8888/v1" and the model Qwen3.8-Flash-Next. Streaming, tool calls, reasoning content, images and videos are supported. The model thinks before it answers, so give replies enough max_tokens; a request without one gets 32,768.
curl -s http://<spark-address>:8888/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "Qwen3.8-Flash-Next",
"messages": [{"role": "user", "content": "Write a Python fibonacci function."}],
"max_tokens": 1000
}'
Run it day to day
./start.sh restart # restart it, e.g. after changing a setting
./stop.sh # stop the server and free the GPU memory
docker logs -f qwen38-flash-next-tf # server log
curl -s http://<spark-address>:8888/health # busy flag and live token totals
Images and video
Up to 50 images a request.
The model's own vision tower (27 layers, 0.84 GiB, from the same checkpoint) turns images and video frames into tokens. Send them as OpenAI-style content parts in a user message: an image_url part, or a video_url part for a clip.
| Images | Videos | |
|---|---|---|
| Formats | JPEG, PNG, WebP | MP4, WebM, MOV, MKV (anything FFmpeg decodes) |
| Per request | Up to 50 (all of a chat's turns count), 10 MB each, 64 MB in all | Up to 2, 64 MB each, 96 MB in all, up to an hour of footage |
| Tokens | Up to 16,384 for all images, at most 4,096 an image | 2 frames a second (at most 256 a video), up to 16,384 tokens a request |
Text requests are unaffected: their replies stay byte-identical with vision on. VISION=0 ./start.sh restart serves text only.
Other languages
Faster in Chinese and Japanese.
Replies in any language come out right, but the drafts are tuned for English and code. For replies mostly in Chinese or Japanese, an opt-in image adds that language's tokens to the draft list: echo 'DRAFT_LANGUAGE=zh' >> .env, then ./start.sh restart. Output stays byte-identical; only speed changes.
| Replies in | Default | Language image | Change |
|---|---|---|---|
| Chinese, thinking off / on | 35.5 / 38.4 | 45.9 / 50.6 | +29% / +32% |
| Japanese, thinking off / on | 39.5 / 46.0 | 46.9 / 49.2 | +19% / +7% |
tok/s, one stream on one Spark. Leave it off for English or code: the larger list makes every draft step a little slower.
Memory and settings
A 1.3M-token pool on one Spark.
The KV pool
Every stream gets its own cache for a full window, so the pool is streams × window: 5 × 262,144 = 1,310,720 tokens in an int8 cache, about 23.4 GiB. The 29.8 GiB of n-gram tables stay on the SSD (PLE_ON_SSD=1), leaving that memory to the cache.
| Setting | KV pool | Estimate |
|---|---|---|
PARALLEL=4 (int8) | 1,048,576 | 97.7 GiB |
PARALLEL=5 (int8, default) | 1,310,720 | 102.5 GiB |
PARALLEL=6 CONTEXT=220000 | 1,320,000 | ~103 GiB |
PARALLEL=3 KV_DTYPE=bf16 | 786,432 | 102.0 GiB |
A setting that doesn't fit is refused at startup, before any weights load, with a window that fits. On v0.6.1, five concurrent ~236k-token prompts all completed.
The ones you might change
Every setting lives in scripts/config.sh; override one from the environment (PARALLEL=4 ./start.sh) or a .env file. The full list is in the README.
| Variable | Default | Meaning |
|---|---|---|
PARALLEL | 5 | Requests decoded together |
CONTEXT | 262144 | Prompt + reply window per stream |
KV_DTYPE | int8 | bf16, int8 or int4 KV cache |
VISION | 1 | Image and video input; 0 serves text only |
DRAFT_LANGUAGE | empty | zh or ja for replies mostly in that language |
THINKING | 1 | Think before answering; 0 answers directly |
MAX_TOKENS | 32768 | Reply length for a request without max_tokens |
TENSORFOLD_MEMORY_RESERVE_GIB | 2 | Memory kept back at startup; raise it (e.g. 6) if other workloads share the box |
PORT | 8888 | Where the API listens |
Licenses
Free, with a few rules.
- The recipe is MIT, and its LICENSE also carries TensorFold's MIT notice for the patches.
- The model weights, downloaded from Hugging Face, are under the Qwen Community License 1.0.
- The image is based on NVIDIA's PyTorch container; by pulling or running it you accept NVIDIA's terms, which the container prints at every start. It also contains Hugging Face transformers (Apache 2.0) and PyAV (BSD) with its FFmpeg libraries (LGPL).
Credits
Built on the work of others.
- TensorFold by Ash Hart
- Qwen3.8 Flash Next by Qwen
- Vontra's MLX 4-bit checkpoint, with the MTP draft head
- A prompt-chunk change by MovieMaker93 (TensorFold #40)
The full list, including the runtime stack, is in the recipe's CREDITS.md.
Run it on your Spark.
Free and open, with every patch and every number published.