Mia’s Favorite
GLM 5.3 Flash
TensorFold · EXL3 4-bit · DFlash2
60.4 tok/s
Decode speed, one stream
- Peak
- 108.8 tok/s ×4 streams
- Prefill
- 1,978.9 tok/s 32k prompt
- Context
- 1M ×4 sharing a ~2.9M pool
21 stars0 forks1 contributor
Last updated
I run open-weight and frontier models side by side, same prompt, one shot each, and publish every raw output. See what they build, and where they break.
New recipe · my #1 for two Sparks
A million tokens of context, images and video, and four conversations at once, served from two Sparks on your desk with TensorFold. Free, with every patch and every number published.
Found the needle in a 981,841-token prompt.
See all recipes
About
I'm a developer who got hooked on running AI locally. Most days I'm deploying and optimizing local AI models to run the BEST way possible for NVIDIA DGX Sparks and other devices I have in my lab.
From GLM, DeepSeek, Qwen, and others, I deploy them all, and make them work at their best state possible on the available hardware.
I'm also comparing and benchmarking all closed frontier AI models, including Claude, GPT, Grok and others, to see how they compare to the best local has to offer.
Models offered
Each recipe runs on my own Sparks, and the benchmarks, patches and launch scripts are all published on GitHub. Clone one and serve a frontier-class model on your own hardware.
Across 19 recipes and sparkDash
Mia’s Favorite
TensorFold · EXL3 4-bit · DFlash2
60.4 tok/s
Decode speed, one stream
21 stars0 forks1 contributor
Last updated
vLLM · NVFP4
48.7 tok/s
Decode speed, one stream
576 stars86 forks8 contributors
Last updated
vLLM · DSpark
38.6 tok/s
Decode speed, one stream
1,439 stars196 forks25 contributors
Last updated
Plus 16 more recipes for DGX Sparks and NVIDIA GPUs, and sparkDash, the dashboard I benchmark with.
Services
I'll install and tune local AI on your NVIDIA DGX Sparks, the same way I run my own lab, from a single machine to a four-node cluster.

Pick the setup that matches your hardware. Every install is done personally by me.
Booked and paid securely through Whop
What customers say
Support & contact
Everything here is independent. Donations pay for GPU time, API credits and the coffee behind late-night runs.
or donate with crypto
Donate with crypto through NOWPayments. Pick an amount and a coin, and you'll finish the payment on their secure checkout.
A question, a model you want tested, a collab idea? I read every message.
or find me on