All posts

Local LLM 1 post

Every post filed under Local LLM, newest first.

Qwen3 8B can run a full coding agent on hardware you already own
4 min read

Qwen3 8B can run a full coding…

Qwen3 8B at Q4_K_M fits in about 5 GB of VRAM and hits roughly 20–50 tok/s on consumer GPUs, including older cards. Here is how to think about local agentic coding without the Mac Mini hype.