Recipe · One DGX Spark
Qwen3.8 27B on one DGX Spark, with TensorFold.
Serve Qwen3.8-27B from a single DGX Spark through an OpenAI-compatible API, with up to 8 concurrent requests, a pinned 78 GiB KV pool (2.5M tokens), the full 262,144-token context (1M with YaRN), DFlash2 speculative decoding, and up to 50 images and video input.
Performance
Measured, not promised.
The official numbers, from sparkDash, the benchmark of the MiaAI-Lab Spark recipes, on one DGX Spark. Aggregate is the total across the concurrent requests, per request the rate of each one.
Decode, prose
Aggregate, per request, and time to first token.
| Requests | Aggregate | Per request | First token |
|---|---|---|---|
| 1 | 58.3 tok/s | 58.3 tok/s | 100 ms |
| 2 | 108.5 tok/s | 55.4 tok/s | 208 ms |
| 4 | 135.1 tok/s | 40.3 tok/s | 417 ms |
| 8 | 220.6 tok/s | 32.7 tok/s | 870 ms |
Decode, code
The same benchmark, on code replies.
| Requests | Aggregate | Per request | First token |
|---|---|---|---|
| 1 | 136.9 tok/s | 136.9 tok/s | 101 ms |
| 2 | 223.9 tok/s | 119.6 tok/s | 202 ms |
| 4 | 292.0 tok/s | 89.2 tok/s | 427 ms |
| 8 | 323.2 tok/s | 53.3 tok/s | 867 ms |
Prefill
How fast a prompt is read before the reply starts.
| Prompt | Prefill | First token |
|---|---|---|
| 8,231 tokens | 1,793.0 tok/s | 4.59 s |
| 16,420 tokens | 1,785.3 tok/s | 9.20 s |
| 32,802 tokens | 1,633.3 tok/s | 20.08 s |
| 65,576 tokens | 1,415.9 tok/s | 46.31 s |
| 131,111 tokens | 1,117.8 tok/s | 117.29 s |
Longer prompts
Measured with a needle-in-a-haystack prompt.
| Prompt | Time | Prefill |
|---|---|---|
| 195k tokens | 205 s | 948 tok/s |
| 255,897 tokens | 315 s | 812 tok/s |
| 491k, with YaRN | 924 s | |
| 884k, with YaRN | 2,664 s |
The first two on the default server (window 262,144); the YaRN rows with a 1M-token window (below).
Speed preview
What 58.3 tok/s feels like.
Sample text played at the measured prose speed: one request at 58.3 tok/s, or four at once at 40.3 tok/s each (135.1 combined), at about four characters a token. Switch between them in the window.
The engine
Powered by TensorFold.
TensorFold, by Ash Hart, is an open-source engine that serves language models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API. Each model family brings its own kernels and its own draft verification, tuned for that model instead of one generic path for all.
Exact decoding
Drafts make it fast without changing a word: a draft is accepted only when it equals the token the same engine would produce one at a time. Send a request with "draft": false to compare.
Apple Silicon and NVIDIA
MLX on Macs, and CUDA on NVIDIA GPUs from compute capability 8.9: Ada (RTX 40), Hopper and Blackwell, including the DGX Spark's GB10 and the RTX 50 series.
Many model families
Qwen3.8-27B, Qwen3.8 Flash Next, GLM-5.3-Flash, Nemotron 3.5 Lightning, Gemma 4 and DeepSeek-V4-Flash, among others, each with its own kernels.
In this recipe
TensorFold v0.6.0 with five patches: up to 50 images and video input, opt-in YaRN up to a 1,048,576-token window, an FP8 attention cache (on by default), a memory reserve of 0, and a pinned KV pool. The engine is upstream's.
What you need
One Spark, nothing else.
A DGX Spark
Or another GB10 system with 128 GB unified memory. The startup estimate is 45 GiB at the defaults; the KV pool then takes what the free memory allows, up to 78 GiB.
Docker
Docker with the NVIDIA container runtime, and your user in the docker group.
Disk
About 60 GB free: about 19 GB for the checkpoints under ~/.cache/huggingface and 25-35 GB for the image.
Optional
The hf CLI on the host, for faster, resumable downloads.
Quick start
Three lines, that's all.
Clone and start
git clone https://github.com/MiaAI-Lab/Qwen3.8-27B-DGX-Spark-TensorFold.git
cd Qwen3.8-27B-DGX-Spark-TensorFold
./start.sh
The first run pulls the prebuilt image (or builds it locally, a few minutes) and downloads the checkpoints, then compiles the CUDA kernels (a few minutes, once). Later starts take about a minute. start.sh runs a smoke test and prints the endpoint.
Talk to it
Any OpenAI client works, on port 8888 with the model Qwen3.8-27B. The model thinks before it answers, so give replies enough max_tokens. Per request you can set temperature, top_p, top_k, seed, turn thinking off with "chat_template_kwargs": {"enable_thinking": false}, or send "draft": false for the serial reference.
curl -s http://<spark-address>:8888/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "Qwen3.8-27B",
"messages": [{"role": "user", "content": "Write a Python fibonacci function."}],
"max_tokens": 1000
}'
Run it day to day
./start.sh restart # apply changed settings
./stop.sh # stop and free the GPU memory
docker logs -f qwen38-27b-tf # server log
Memory
A 2.5M-token pool, pinned.
The pinned KV pool
KV_POOL_GB (default auto) reserves up to 78 GiB of attention cache at startup, 2,555,904 tokens, and keeps it. auto sizes it from the memory free when start.sh runs (free GiB minus ~31, at most 78: a Spark with ~109 GiB free gets the full 78, one with 90 GiB free gets 59); a number sets it, 0 turns the pin off.
The server holds about 97 GiB from the first second, and that number does not move. Every stream's cache grows inside the pool, and a request that does not fit waits for others to finish. Eight full 262,144-token windows (2.1M tokens) fit in it at once.
- Measured: 8 simultaneous ~59k-token prompts, 8 of 8 ok, process memory steady at 96,693-96,715 MiB throughout.
- Not a speed feature: it avoids growth copies and keeps memory predictable; decode and prefill are unchanged.
- On a shared machine it is tight: size
KV_POOL_GBdown if other containers need room.
FP8 KV cache
KV_DTYPE=fp8, the default, stores the attention layers' keys and values in one byte each: 32 KiB a token instead of 64, so twice the tokens fit. A drafted reply still equals the serial one.
| bf16 cache | FP8 cache | |
|---|---|---|
| A full 262,144-token stream | 16 GiB | 8 GiB |
| KV pool at the admission budget (~107 GiB) | ~1.0M tokens | ~2.2M tokens |
| 1M-token windows (YaRN) | 2 streams | 4 streams |
| Prefill 8k / 126k | 1,860 / 1,149 tok/s | 1,862 / 1,147 tok/s |
| 8-request decode | 150 tok/s | 146 tok/s |
The last two rows come from a separate comparison run, not from sparkDash.
What it costs in quality
Measured in the engine on 8 sequences of 4,096 tokens (wikitext-2 and CPython source), against the bf16 cache with bf16 prompts (KL: the divergence from that reference; top-1: how often the most likely token agrees). The cache costs far less than the prompt precision: if quality matters more than prefill speed, PREFILL_FP8=0 is the setting to change first.
| Perplexity | KL | Top-1 match | |
|---|---|---|---|
| Reference: all bf16 | 3.316 | ||
| FP8 cache only | 3.315 | 0.0031 | 98.7% |
| FP8 prompts only | 3.397 | 0.0435 | 94.1% |
| Both FP8 (defaults) | 3.387 | 0.0419 | 94.0% |
A 12-question reasoning probe scored 12/12 with the FP8 cache and 11/12 with bf16 (one retrieval miss: noise at this size). Only 4k-token contexts were scored by KL; long-context quality was checked with needles only.
Longer context
A 1M-token window with YaRN.
The model's native window is 262,144 tokens, and that is what the server runs by default. YaRN stretches it by a factor; 4 gives 1,048,576 tokens, and the context follows the factor, so nothing else needs setting:
echo 'YARN_FACTOR=4' >> .env
./start.sh restart
To go back, remove the line and restart. With a 1M window, auto pins about 70 GiB instead of 78: room for two full 1M streams, or 8 streams of ~290k.
What to expect
Long prompts take time.
- Prefill is slow at length, since attention grows with it: 205 s for 195k tokens, 924 s for 491k and 2,664 s (44 minutes) for 884k. Send long prompts with a client timeout of an hour or more.
- Quality: YaRN rescales every position, so short-range replies change slightly. It was checked with needles (found at 195k, 491k and 884k tokens), not beyond. If most of your prompts are short, leave YaRN off.
- Drafts still verified: the DFlash2 drafter is not scaled, and its drafts are still checked against the model's own samples.
Images and video
Up to 50 images a request.
TensorFold's own image support takes 4 images a request and no video; this recipe raises that and adds video. Send them as OpenAI-style content parts in a user message: an image_url part, or a video_url part for a clip.
| Images | Videos | |
|---|---|---|
| Formats | JPEG, PNG, WebP | MP4, WebM, MOV, MKV (anything FFmpeg decodes) |
| Per request | Up to 50 (all of a chat's turns count), 10 MB each, 64 MB in all | Up to 2, 64 MB each, 96 MB in all, up to an hour of footage |
| Tokens | Up to 16,384 for all images, at most 4,096 an image | 2 frames a second (at most 256 a video), up to 16,384 tokens a request |
Tested: 50 full-HD photos in one request (15,735 tokens, 13 s) and a 90-second 1280×720 video (16,135 tokens, 12.5 s). Data URLs by default; VISION_URLS=1 also fetches public https:// URLs.
Tuning
What was tuned, and why.
- FP8 prompts on (
PREFILL_FP8=1): +52% prefill in this recipe's test (NVFP4, 12.6k tokens; upstream measures +35-40% on this checkpoint), at a lower prompt precision.PREFILL_FP8=0gives bf16 prompts; decode is unaffected. - 8 requests at once instead of upstream's one at a time: no cost for a single request, and 3x the aggregate decode.
PARALLEL=16also works, with the pool topping out near 72 GiB. - A memory reserve of 0: TensorFold's budget is all of the available memory instead of keeping a tenth of RAM back. Raise
TENSORFOLD_MEMORY_RESERVE_GIBif other workloads share the box: running out of unified memory can freeze the host. - MLX 4-bit over NVIDIA NVFP4: NVFP4 prefilled the same but decoded 20-25% slower and cannot take images. Quality between the two was not measured here.
Settings
The ones you might change.
Every setting lives in scripts/config.sh; override one from the environment, a .env file next to start.sh, or tensorfold serve flags (./start.sh restart --context 131072). The full list is in the README.
| Variable | Default | Meaning |
|---|---|---|
PARALLEL | 8 | Requests decoded together; 1 serves one at a time |
CONTEXT | 262144 | Prompt + reply window per stream (times YARN_FACTOR) |
YARN_FACTOR | unset | 4 for a 1,048,576-token window |
KV_POOL_GB | auto | Attention cache pinned at startup; a number sets it, 0 grows on demand |
KV_DTYPE | fp8 | fp8 (32 KiB a token) or bf16 (64 KiB) |
PREFILL_FP8 | 1 | FP8 prompts: faster prefill, lower prompt precision; 0 for bf16 |
VISION | 1 | Image and video input |
THINKING | 1 | Open a think block by default |
PORT | see file | Where the API listens (8888 in the README's examples) |
Licenses
Free, with a few rules.
- The recipe is Apache 2.0. Its patches modify TensorFold v0.6.0, which is Apache-2.0; parts of the scripts and tools come from my MIT-licensed Flash Next recipe and keep its notice.
- The model weights, downloaded from Hugging Face, are under the Qwen Community License 1.0.
- The image is based on NVIDIA's PyTorch container; by pulling or running it you accept NVIDIA's terms, which the container prints at every start. It also contains Hugging Face transformers (Apache 2.0) and PyAV (BSD) with its FFmpeg libraries (LGPL).
Credits
Built on the work of others.
- TensorFold by Ash Hart (v0.6.0)
- Qwen3.8-27B by Qwen
- Vontra's MLX 4-bit checkpoint
- z-lab's DFlash2 drafter
- The launcher, tools and video code, adapted from my Flash Next recipe
The full list, including the YaRN source and the runtime stack, is in the recipe's CREDITS.md.
Run it on your Spark.
Free and open, with every patch and every number published.