Local AI,
tested in the open.

I run open-weight and frontier models side by side, same prompt, one shot each, and publish every raw output. See what they build, and where they break.

Building with AIPushing Local AI ForwardOpen-Weight ModelsNVIDIA DGX SparksPremium Local AI RecipesQuantization Deep DivesModel Comparisons & RecommendationsBenchmarksNo Cherry-PickingTips & TricksAgentic Workflows

New recipe · my #1 for two Sparks

GLM 5.3 Flash, now on two DGX Sparks.

Powered by TensorFold

A million tokens of context, images and video, and four conversations at once, served from two Sparks on your desk with TensorFold. Free, with every patch and every number published.

One request
60.4tok/s
Four at once
108.8tok/s combined
Context
1Mtokens a request
Prefill
1,978.9tok/s, 32k prompt

Found the needle in a 981,841-token prompt.

See all recipes
TensorFold · GLM-5.3-Flash-EXL3 · 2× DGX Spark
60.4 tok/s Sample text, played at the measured speed
Mia, line-art portrait
Mia@MiaAI_lab
github.com/MiaAI-Lab

About

Hi, I'm Mia! I'm publicly developing the frontier of local and secure intelligence.

I'm a developer who got hooked on running AI locally. Most days I'm deploying and optimizing local AI models to run the BEST way possible for NVIDIA DGX Sparks and other devices I have in my lab.

From GLM, DeepSeek, Qwen, and others, I deploy them all, and make them work at their best state possible on the available hardware.

I'm also comparing and benchmarking all closed frontier AI models, including Claude, GPT, Grok and others, to see how they compare to the best local has to offer.

Models offered

The frontier of local intelligence: free, open-source recipes for NVIDIA DGX Sparks and RTX GPUs.

Each recipe runs on my own Sparks, and the benchmarks, patches and launch scripts are all published on GitHub. Clone one and serve a frontier-class model on your own hardware.

GitHub stars
5,607
Forks
789
Contributors
95

Across 19 recipes and sparkDash

Mia’s Favorite

Needs 2–3 Sparks

Vision

GLM 5.3 Flash

TensorFold · EXL3 4-bit · DFlash2

60.4 tok/s

Decode speed, one stream

Peak
108.8 tok/s ×4 streams
Prefill
1,978.9 tok/s 32k prompt
Context
1M ×4 sharing a ~2.9M pool

21 stars0 forks1 contributor

Last updated

See the full recipe

Needs 1 Spark

Vision

Qwen3.8 Flash Next

vLLM · NVFP4

48.7 tok/s

Decode speed, one stream

Peak
163 tok/s ×8 streams
Prefill
1,769 tok/s 32k prompt
Context
262k · 512k with YaRN

576 stars86 forks8 contributors

Last updated

Get the free recipe

Needs 2–3 Sparks

Vision

DeepSeek V4 Flash

vLLM · DSpark

38.6 tok/s

Decode speed, one stream

Peak
139 tok/s ×6 streams
Prefill
18.0 s to first token, 32k
Context
1M

1,439 stars196 forks25 contributors

Last updated

Get the free recipe
See all 19 models

Plus 16 more recipes for DGX Sparks and NVIDIA GPUs, and sparkDash, the dashboard I benchmark with.

Install with an AI agent

Install this model

Paste this into an AI coding agent that can run commands on the machine that will serve the model (the head node, for a Spark cluster), such as Claude Code or Codex. It reads the recipe, adapts it to your setup, starts the server and checks that it works.


        

Services

Get your local AI lab online and running flawlessly, and optimized for your hardware.

I'll install and tune local AI on your NVIDIA DGX Sparks, the same way I run my own lab, from a single machine to a four-node cluster.

Personal Install by Mia: an NVIDIA DGX Spark

Personal Install by Mia

Pick the setup that matches your hardware. Every install is done personally by me.

Booked and paid securely through Whop

The book

From Zero to Seen

Get monetized on X

How I grew my X account from 200 to over 36,000 followers in under four months, the honest way, and how to get monetized on X.

Support & contact

How to support the frontier of local intelligence.

Everything here is independent. Donations pay for GPU time, API credits and the coffee behind late-night runs.

Pushing Local AI Forward

or donate with crypto

Donate with crypto through NOWPayments. Pick an amount and a coin, and you'll finish the payment on their secure checkout.

Amount (USD)

Minimum $10 · network fees paid by sender

Contact

A question, a model you want tested, a collab idea? I read every message.

or find me on