Local AI,
tested in the open.

I run open-weight and frontier models side by side — same prompt, one shot each — and publish every raw output. See what they build, and where they break.

Building with AIPushing Local AI ForwardOpen-Weight ModelsNVIDIA DGX SparksPremium Local AI RecipesQuantization Deep DivesModel Comparisons & RecommendationsBenchmarksNo Cherry-PickingTips & TricksAgentic Workflows
Mia, line-art portrait
Mia@MiaAI_lab
github.com/MiaAI-Lab

About

Hi, I'm Mia! I'm publicly developing the frontier of local and secure intelligence.

I'm a developer who got hooked on running AI locally. Most days I'm deploying and optimizing local AI models to run the BEST way possible for NVIDIA DGX Sparks and other devices I have in my lab.

From GLM, DeepSeek, Qwen, and others, I deploy them all, and make them work at their best state possible on the available hardware.

I'm also comparing and benchmarking all closed frontier AI models, including Claude, GPT, Grok and others, to see how they compare to the best local has to offer.

Models offered

The frontier of local intelligence: free, open-source recipes for NVIDIA DGX Sparks and RTX GPUs.

Each recipe runs on my own Sparks, and the benchmarks, patches and launch scripts are all published on GitHub. Clone one and serve a frontier-class model on your own hardware.

GitHub stars
5,214
Forks
723
Contributors
80

Across 16 recipes and sparkDash

Mia’s Favorite

Needs 2–4 Sparks

GLM 5.3 Flash

vLLM · EXL3 4-bit

36.1 tok/s

Decode speed, one stream

Peak
75 tok/s ×4 streams
Prefill
1,428 tok/s 32k prompt
Context
Up to 1M

642 stars117 forks32 contributors

Last updated

Get the free recipe

Needs 2–3 Sparks

DeepSeek V4 Flash

vLLM · DSpark · vision

38.6 tok/s

Decode speed, one stream

Peak
139 tok/s ×6 streams
Prefill
18.0 s to first token, 32k
Context
1M

1,427 stars192 forks25 contributors

Last updated

Get the free recipe

Needs a 16–32 GB NVIDIA GPU

Qwen3.8 27B

ExLlamaV3 · EXL3 · Windows & Linux

16 GB VRAM

What it's built around; picks the best quant for 12–32 GB

Quant
2.0–6.0 bpw picked for your VRAM
Prefill
— not published
Context
Up to 262k on 24 GB and up

526 stars50 forks2 contributors

Last updated

Get the free recipe
See all 16 models

Plus 13 more recipes for DGX Sparks and NVIDIA GPUs, and sparkDash, the dashboard I benchmark with.

Install with an AI agent

Install this model

Paste this into an AI coding agent that can run commands on the machine that will serve the model (the head node, for a Spark cluster), such as Claude Code or Codex. It reads the recipe, adapts it to your setup, starts the server and checks that it works.


        

Services

Get your local AI lab online and running flawlessly, and optimized for your hardware.

I'll install and tune local AI on your NVIDIA DGX Sparks, the same way I run my own lab — from a single machine to a four-node cluster.

Personal Install by Mia — an NVIDIA DGX Spark

Personal Install by Mia

Pick the setup that matches your hardware. Every install is done personally by me.

Booked and paid securely through Whop

The book

From Zero to Seen

Get monetized on X

How I grew my X account from 200 to 32,000 followers in just over three months, the honest way — and how to turn that attention into income.

Support & contact

How to support the frontier of local intelligence.

Everything here is independent. Donations pay for GPU time, API credits and the coffee behind late-night runs.

Pushing Local AI Forward

or donate with crypto

Donate with crypto through NOWPayments. Pick an amount and a coin — you'll finish the payment on their secure checkout.

Amount (USD)

Minimum $10 · network fees paid by sender

Contact

A question, a model you want tested, a collab idea? I read every message.

or find me on