Models

The frontier of local intelligence: free, open-source recipes for NVIDIA DGX Sparks and RTX GPUs.

Each recipe runs on my own Sparks, and the benchmarks, patches and launch scripts are all published on GitHub. Clone one and serve a frontier-class model on your own hardware.

GitHub stars
5,214
Forks
723
Contributors
80

Across 16 recipes and sparkDash

Recipes built for

DGX Spark

NVIDIA GPU

Mia’s Favorite

Needs 2–4 Sparks

GLM 5.3 Flash

vLLM · EXL3 4-bit

36.1 tok/s

Decode speed, one stream

Peak
75 tok/s ×4 streams
Prefill
1,428 tok/s 32k prompt
Context
Up to 1M

642 stars117 forks32 contributors

Last updated

Get the free recipe

Needs 2–3 Sparks

DeepSeek V4 Flash

vLLM · DSpark · vision

38.6 tok/s

Decode speed, one stream

Peak
139 tok/s ×6 streams
Prefill
18.0 s to first token, 32k
Context
1M

1,427 stars192 forks25 contributors

Last updated

Get the free recipe

Needs a 16–32 GB NVIDIA GPU

Qwen3.8 27B

ExLlamaV3 · EXL3 · Windows & Linux

16 GB VRAM

What it's built around; picks the best quant for 12–32 GB

Quant
2.0–6.0 bpw picked for your VRAM
Prefill
— not published
Context
Up to 262k on 24 GB and up

526 stars50 forks2 contributors

Last updated

Get the free recipe

Needs 1 Spark

Qwen3.8 Flash Next

vLLM · NVFP4 · vision

48.7 tok/s

Decode speed, one stream

Peak
163 tok/s ×8 streams
Prefill
1,769 tok/s 32k prompt
Context
262k · 512k with YaRN

513 stars75 forks8 contributors

Last updated

Get the free recipe

Needs 1 Spark

Qwen3.8 27B

SGLang · NVFP4 · DSpark

18.3 tok/s

Decode speed, one stream

Peak
228 tok/s ×16 streams, DFlash2
Prefill
~8.3 s to first token, 16k
Context
262k · 1M with YaRN

451 stars50 forks4 contributors

Last updated

Get the free recipe

Needs 2 Sparks

Qwen3.8 Flash Next

vLLM · NVFP4

54.4 tok/s

Decode speed, one stream

Peak
207 tok/s ×8 streams
Prefill
2,962 tok/s 32k prompt
Context
262k · 1M with YaRN

385 stars45 forks5 contributors

Last updated

Get the free recipe

Needs 2 Sparks

DeepSeek V4.1 Flash EXL3

vLLM · EXL3 2.9-bit

31.6 tok/s

Decode speed, one stream

Peak
54 tok/s ×4 streams, speculation off
Prefill
1,055 tok/s 32k prompt
Context
600k

242 stars33 forks5 contributors

Last updated

Get the free recipe

Needs 3–4 Sparks

DeepSeek V4.1 Flash

SGLang · MXFP4 + FP8

51.0 tok/s

Decode speed, one stream

Peak
85.4 tok/s ×4 streams
Prefill
~2,000 tok/s
Context
256k

177 stars29 forks5 contributors

Last updated

Get the free recipe

Needs an RTX 5090

Qwen3.8 27B

vLLM · NVFP4 · MTP

~160 tok/s

Decode speed, one stream (as published)

Context
262k one full session
Prefill
— not published
Concurrency
1 session by design; see the README

88 stars10 forks1 contributor

Last updated

Get the free recipe

1 SparkorRTX PRO 6000

Laguna S 2.1

vLLM · NVFP4 · DFlash

76.7 tok/s

Decode speed, one stream

Peak
107 tok/s ×8, server-side
Prefill
— not published
Context
256k

83 stars6 forks4 contributors

Last updated

Get the free recipe

Needs 1 Spark

Qwen3.6 35B-A3B

vLLM · NVFP4 · MTP · vision

95.1 tok/s

Decode speed, one stream

Peak
317 tok/s ×8 streams
Prefill
— not published
Context
262k

69 stars6 forks2 contributors

Last updated

Get the free recipe

Needs an RTX PRO 6000

Qwen3.8 27B

SGLang · NVFP4 · DFlash 2

240+ tok/s

Decode speed, one stream (as published)

Context
256k native
Prefill
— not published
Concurrency
8 requests

61 stars7 forks3 contributors

Last updated

Get the free recipe

Needs 2 Sparks

MiMo V2.6 Flash

SGLang · MXFP4 · multimodal

35.0 tok/s

Decode speed, one stream

Peak
120 tok/s ×8 streams
Prefill
2,317 tok/s 32k prompt
Context
Up to 1M

33 stars3 forks1 contributor

Last updated

Get the free recipe

1 SparkorRTX 5090orRTX PRO 6000

Nemotron 3.5 Lightning

SGLang · NVFP4 · DSpark

1M context

30B-A3B hybrid MoE, native 1M-token window

Concurrency
48 requests
KV pool
~4.9M tokens
Memory
0.78 MEM_FRACTION_STATIC

25 stars1 fork1 contributor

Last updated

Get the free recipe

Needs 1 Spark

Ling 3.0 Flash

SGLang · INT4 · DSpark

83 tok/s

Decode speed, one stream

Peak
136 tok/s ×6 streams
Prefill
— not published
Context
256k

13 stars2 forks1 contributor

Last updated

Get the free recipe

1 SparkorRTX 5090orRTX PRO 6000

Muse Glimmer 30B

vLLM · NVFP4 · DFlash · vision

~35 tok/s

Decode speed, one stream (measured)

Concurrency
9.2× full 256k requests
Prefill
— not published
Context
256k 131k native

8 stars0 forks1 contributor

Last updated

Get the free recipe

Decode speeds are single-stream figures published in each recipe's README (prose prompts wherever the README separates them); many were measured with sparkDash, my open-source benchmark dashboard. The meters run from 0 to 60 tok/s. GitHub numbers update every 15 minutes.

Lab tools

The dashboard behind every number.

sparkDash dashboard monitoring several DGX Sparks

For DGX Sparks & NVIDIA GPUs

sparkDash

A real-time dashboard for one or many DGX Sparks and NVIDIA GPU machines, and the benchmark behind the speeds on this page.

  • Monitor one or many DGX Sparks and NVIDIA GPU machines in one window
  • Live GPU, unified memory, storage and network metrics
  • Detects your local LLM server (vLLM, SGLang, llama.cpp and more) and shows live tok/s
  • Built-in decode and prefill benchmarks, the ones behind the numbers above
  • Prompt Showcase: stream up to 32 prompts side by side

471 stars97 forks15 contributors

Last updated

Get sparkDash

Install with an AI agent

Install this model

Paste this into an AI coding agent that can run commands on the machine that will serve the model (the head node, for a Spark cluster), such as Claude Code or Codex. It reads the recipe, adapts it to your setup, starts the server and checks that it works.