Recipe · One DGX Spark

Qwen3.8 27B on one DGX Spark, with TensorFold.

Serve Qwen3.8-27B from a single DGX Spark through an OpenAI-compatible API, with up to 8 concurrent requests, a pinned 78 GiB KV pool (2.5M tokens), the full 262,144-token context (1M with YaRN), DFlash2 speculative decoding, and up to 50 images and video input.

Get the free recipe

Install with an AI agent

Install this model

Paste this into an AI coding agent that can run commands on the machine that will serve the model (the head node, for a Spark cluster), such as Claude Code or Codex. It reads the recipe, adapts it to your setup, starts the server and checks that it works.


        
Qwen3.8 27B on TensorFold, single DGX Spark

Performance

Measured, not promised.

One request58.3tok/s, prose
Eight at once220.6tok/s combined
Code136.9tok/s, one request
Prefill1,793tok/s, 8k prompt

The official numbers, from sparkDash, the benchmark of the MiaAI-Lab Spark recipes, on one DGX Spark. Aggregate is the total across the concurrent requests, per request the rate of each one.

Decode, prose

Aggregate, per request, and time to first token.

RequestsAggregatePer requestFirst token
158.3 tok/s58.3 tok/s100 ms
2108.5 tok/s55.4 tok/s208 ms
4135.1 tok/s40.3 tok/s417 ms
8220.6 tok/s32.7 tok/s870 ms

Decode, code

The same benchmark, on code replies.

RequestsAggregatePer requestFirst token
1136.9 tok/s136.9 tok/s101 ms
2223.9 tok/s119.6 tok/s202 ms
4292.0 tok/s89.2 tok/s427 ms
8323.2 tok/s53.3 tok/s867 ms

Prefill

How fast a prompt is read before the reply starts.

PromptPrefillFirst token
8,231 tokens1,793.0 tok/s4.59 s
16,420 tokens1,785.3 tok/s9.20 s
32,802 tokens1,633.3 tok/s20.08 s
65,576 tokens1,415.9 tok/s46.31 s
131,111 tokens1,117.8 tok/s117.29 s

Longer prompts

Measured with a needle-in-a-haystack prompt.

PromptTimePrefill
195k tokens205 s948 tok/s
255,897 tokens315 s812 tok/s
491k, with YaRN924 s
884k, with YaRN2,664 s

The first two on the default server (window 262,144); the YaRN rows with a 1M-token window (below).

Speed preview

What 58.3 tok/s feels like.

Sample text played at the measured prose speed: one request at 58.3 tok/s, or four at once at 40.3 tok/s each (135.1 combined), at about four characters a token. Switch between them in the window.

TensorFold · Qwen3.8-27B · 1× DGX Spark
58.3 tok/sSample text, played at the measured speed

The engine

Powered by TensorFold.

TensorFold, by Ash Hart, is an open-source engine that serves language models on Apple Silicon and NVIDIA GPUs through an OpenAI-compatible API. Each model family brings its own kernels and its own draft verification, tuned for that model instead of one generic path for all.

Exact decoding

Drafts make it fast without changing a word: a draft is accepted only when it equals the token the same engine would produce one at a time. Send a request with "draft": false to compare.

Apple Silicon and NVIDIA

MLX on Macs, and CUDA on NVIDIA GPUs from compute capability 8.9: Ada (RTX 40), Hopper and Blackwell, including the DGX Spark's GB10 and the RTX 50 series.

Many model families

Qwen3.8-27B, Qwen3.8 Flash Next, GLM-5.3-Flash, Nemotron 3.5 Lightning, Gemma 4 and DeepSeek-V4-Flash, among others, each with its own kernels.

In this recipe

TensorFold v0.6.0 with five patches: up to 50 images and video input, opt-in YaRN up to a 1,048,576-token window, an FP8 attention cache (on by default), a memory reserve of 0, and a pinned KV pool. The engine is upstream's.

What you need

One Spark, nothing else.

A DGX Spark

Or another GB10 system with 128 GB unified memory. The startup estimate is 45 GiB at the defaults; the KV pool then takes what the free memory allows, up to 78 GiB.

Docker

Docker with the NVIDIA container runtime, and your user in the docker group.

Disk

About 60 GB free: about 19 GB for the checkpoints under ~/.cache/huggingface and 25-35 GB for the image.

Optional

The hf CLI on the host, for faster, resumable downloads.

Quick start

Three lines, that's all.

1

Clone and start

git clone https://github.com/MiaAI-Lab/Qwen3.8-27B-DGX-Spark-TensorFold.git
cd Qwen3.8-27B-DGX-Spark-TensorFold
./start.sh

The first run pulls the prebuilt image (or builds it locally, a few minutes) and downloads the checkpoints, then compiles the CUDA kernels (a few minutes, once). Later starts take about a minute. start.sh runs a smoke test and prints the endpoint.

2

Talk to it

Any OpenAI client works, on port 8888 with the model Qwen3.8-27B. The model thinks before it answers, so give replies enough max_tokens. Per request you can set temperature, top_p, top_k, seed, turn thinking off with "chat_template_kwargs": {"enable_thinking": false}, or send "draft": false for the serial reference.

curl -s http://<spark-address>:8888/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "Qwen3.8-27B",
  "messages": [{"role": "user", "content": "Write a Python fibonacci function."}],
  "max_tokens": 1000
}'
3

Run it day to day

./start.sh restart                # apply changed settings
./stop.sh                         # stop and free the GPU memory
docker logs -f qwen38-27b-tf      # server log

Memory

A 2.5M-token pool, pinned.

The pinned KV pool

KV_POOL_GB (default auto) reserves up to 78 GiB of attention cache at startup, 2,555,904 tokens, and keeps it. auto sizes it from the memory free when start.sh runs (free GiB minus ~31, at most 78: a Spark with ~109 GiB free gets the full 78, one with 90 GiB free gets 59); a number sets it, 0 turns the pin off.

The server holds about 97 GiB from the first second, and that number does not move. Every stream's cache grows inside the pool, and a request that does not fit waits for others to finish. Eight full 262,144-token windows (2.1M tokens) fit in it at once.

  • Measured: 8 simultaneous ~59k-token prompts, 8 of 8 ok, process memory steady at 96,693-96,715 MiB throughout.
  • Not a speed feature: it avoids growth copies and keeps memory predictable; decode and prefill are unchanged.
  • On a shared machine it is tight: size KV_POOL_GB down if other containers need room.

FP8 KV cache

KV_DTYPE=fp8, the default, stores the attention layers' keys and values in one byte each: 32 KiB a token instead of 64, so twice the tokens fit. A drafted reply still equals the serial one.

bf16 cacheFP8 cache
A full 262,144-token stream16 GiB8 GiB
KV pool at the admission budget (~107 GiB)~1.0M tokens~2.2M tokens
1M-token windows (YaRN)2 streams4 streams
Prefill 8k / 126k1,860 / 1,149 tok/s1,862 / 1,147 tok/s
8-request decode150 tok/s146 tok/s

The last two rows come from a separate comparison run, not from sparkDash.

What it costs in quality

Measured in the engine on 8 sequences of 4,096 tokens (wikitext-2 and CPython source), against the bf16 cache with bf16 prompts (KL: the divergence from that reference; top-1: how often the most likely token agrees). The cache costs far less than the prompt precision: if quality matters more than prefill speed, PREFILL_FP8=0 is the setting to change first.

PerplexityKLTop-1 match
Reference: all bf163.316
FP8 cache only3.3150.003198.7%
FP8 prompts only3.3970.043594.1%
Both FP8 (defaults)3.3870.041994.0%

A 12-question reasoning probe scored 12/12 with the FP8 cache and 11/12 with bf16 (one retrieval miss: noise at this size). Only 4k-token contexts were scored by KL; long-context quality was checked with needles only.

Longer context

A 1M-token window with YaRN.

The model's native window is 262,144 tokens, and that is what the server runs by default. YaRN stretches it by a factor; 4 gives 1,048,576 tokens, and the context follows the factor, so nothing else needs setting:

echo 'YARN_FACTOR=4' >> .env
./start.sh restart

To go back, remove the line and restart. With a 1M window, auto pins about 70 GiB instead of 78: room for two full 1M streams, or 8 streams of ~290k.

What to expect

Long prompts take time.

  • Prefill is slow at length, since attention grows with it: 205 s for 195k tokens, 924 s for 491k and 2,664 s (44 minutes) for 884k. Send long prompts with a client timeout of an hour or more.
  • Quality: YaRN rescales every position, so short-range replies change slightly. It was checked with needles (found at 195k, 491k and 884k tokens), not beyond. If most of your prompts are short, leave YaRN off.
  • Drafts still verified: the DFlash2 drafter is not scaled, and its drafts are still checked against the model's own samples.

Images and video

Up to 50 images a request.

TensorFold's own image support takes 4 images a request and no video; this recipe raises that and adds video. Send them as OpenAI-style content parts in a user message: an image_url part, or a video_url part for a clip.

ImagesVideos
FormatsJPEG, PNG, WebPMP4, WebM, MOV, MKV (anything FFmpeg decodes)
Per requestUp to 50 (all of a chat's turns count), 10 MB each, 64 MB in allUp to 2, 64 MB each, 96 MB in all, up to an hour of footage
TokensUp to 16,384 for all images, at most 4,096 an image2 frames a second (at most 256 a video), up to 16,384 tokens a request

Tested: 50 full-HD photos in one request (15,735 tokens, 13 s) and a 90-second 1280×720 video (16,135 tokens, 12.5 s). Data URLs by default; VISION_URLS=1 also fetches public https:// URLs.

Tuning

What was tuned, and why.

  • FP8 prompts on (PREFILL_FP8=1): +52% prefill in this recipe's test (NVFP4, 12.6k tokens; upstream measures +35-40% on this checkpoint), at a lower prompt precision. PREFILL_FP8=0 gives bf16 prompts; decode is unaffected.
  • 8 requests at once instead of upstream's one at a time: no cost for a single request, and 3x the aggregate decode. PARALLEL=16 also works, with the pool topping out near 72 GiB.
  • A memory reserve of 0: TensorFold's budget is all of the available memory instead of keeping a tenth of RAM back. Raise TENSORFOLD_MEMORY_RESERVE_GIB if other workloads share the box: running out of unified memory can freeze the host.
  • MLX 4-bit over NVIDIA NVFP4: NVFP4 prefilled the same but decoded 20-25% slower and cannot take images. Quality between the two was not measured here.

Settings

The ones you might change.

Every setting lives in scripts/config.sh; override one from the environment, a .env file next to start.sh, or tensorfold serve flags (./start.sh restart --context 131072). The full list is in the README.

VariableDefaultMeaning
PARALLEL8Requests decoded together; 1 serves one at a time
CONTEXT262144Prompt + reply window per stream (times YARN_FACTOR)
YARN_FACTORunset4 for a 1,048,576-token window
KV_POOL_GBautoAttention cache pinned at startup; a number sets it, 0 grows on demand
KV_DTYPEfp8fp8 (32 KiB a token) or bf16 (64 KiB)
PREFILL_FP81FP8 prompts: faster prefill, lower prompt precision; 0 for bf16
VISION1Image and video input
THINKING1Open a think block by default
PORTsee fileWhere the API listens (8888 in the README's examples)

Licenses

Free, with a few rules.

  • The recipe is Apache 2.0. Its patches modify TensorFold v0.6.0, which is Apache-2.0; parts of the scripts and tools come from my MIT-licensed Flash Next recipe and keep its notice.
  • The model weights, downloaded from Hugging Face, are under the Qwen Community License 1.0.
  • The image is based on NVIDIA's PyTorch container; by pulling or running it you accept NVIDIA's terms, which the container prints at every start. It also contains Hugging Face transformers (Apache 2.0) and PyAV (BSD) with its FFmpeg libraries (LGPL).

Credits

Built on the work of others.

The full list, including the YaRN source and the runtime stack, is in the recipe's CREDITS.md.

Run it on your Spark.

Free and open, with every patch and every number published.