Insights, tools & honest experiments
I run local models, break things, and write about what actually works. No hype, no sponsored takes — just real-world results from a developer who ships.
What I cover
I work extensively with Qwen, DeepSeek, GLM, MiniMax and many more — quantizing, serving, and pushing them to their limits on real hardware.
I rely on llama.cpp, vllm, sglang, and a growing stack of open-source tools to build, benchmark, and ship local AI workflows that actually hold up in production.
Real numbers on inference speed, memory usage, and output quality — plus real agentic workflows tested end to end. No cherry-picked results.
Agents, RAG pipelines, code assistants — I run real coding tests on actual projects and share what works, what breaks, and why.