All posts

Models 18 posts

Every post filed under Models, newest first.

VGGT-Omega scales 3D reconstruction to 10B parameters with 70% less training memory
7 min read

VGGT-Omega scales 3D reconstruction to 10B parameters…

Meta and Oxford's VGGT-Omega is a CVPR 2026 oral that cuts GPU training memory by roughly 70%, scales to 10B parameters, and beats optimization pipelines on dynamic scenes. Here's what changed and how I'd evaluate it before betting a product on feed-forward 3D.

ByteDance Lance: a 3B model that reads, generates, and edits images and video in one stack
6 min read

ByteDance Lance: a 3B model that reads,…

Lance is ByteDance's Apache 2.0 unified multimodal model at 3B active parameters. It handles captioning, VQA, text-to-video, and multi-turn edits in one framework, trained on 128 A100s. Here's what that means if you ship generative media.

GLM-5.3 got 50% better at coding without changing the base model
4 min read

GLM-5.3 got 50% better at coding without…

Z.ai shipped GLM-5.3 on the same weights as GLM-5.2 and jumped from 4.6 to 28.3 on Terminal-Bench 3.0. The lesson for builders: post-training and harness fit beat another pre-training run.

OrcaRouter's abliterated Qwen3 27B is a red-team baseline, not a chatbot
3 min read

OrcaRouter's abliterated Qwen3 27B is a red-team…

Qwen3.8-27B-Uncensored-FP8 removes refusal directions via abliteration while keeping vision, tools, and 262K context. Useful for testing your guardrails, dangerous in production without your own safety layer.

Xiaomi open-sourced 100K hours of robot brain training
5 min read

Xiaomi open-sourced 100K hours of robot brain…

Xiaomi-Robotics-1 releases weights, post-training code, and benchmarks for a VLA model pretrained on 100K+ hours of embodiment-free data. Open weights skip the most expensive part of building a manipulation stack.

Liquid AI's LFM2.5-2.6B runs a 128K agent on your phone at 30 tok/s
8 min read

Liquid AI's LFM2.5-2.6B runs a 128K agent…

LFM2.5-2.6B is a 2.6B open-weight agent model that stays under 2.5GB, hits 220 tok/s on an M5 Max, and beats Qwen3.5-9B on most tool-use benchmarks. Here is how to wire it into Hermes, OpenClaw, or Pi through a local OpenAI-compatible endpoint.

Unsloth’s Dynamic GGUFs make Gemma 4 12B fit on a laptop
5 min read

Unsloth’s Dynamic GGUFs make Gemma 4 12B…

Google’s Gemma 4 12B is multimodal with a 256K context window. Unsloth’s Dynamic GGUFs squeeze a usable 4-bit build into roughly 8GB of memory without throwing quality off a cliff.

GLOSSOPETRAE proves LLMs code better in alien languages than in English
4 min read

GLOSSOPETRAE proves LLMs code better in alien…

GLOSSOPETRAE generates procedural coding languages from a seed. At full opacity, human legibility drops to ~15% while Opus and GPT hit 97-100% task accuracy. Human readability hurts model performance.

VIMPO beats GRPO on hard math benchmarks without training a critic
4 min read

VIMPO beats GRPO on hard math benchmarks…

VIMPO derives a policy-implied value function from KL-regularized RL optimality conditions. It improves over GRPO on AIME and OlympiadBench while staying critic-free. Code on GitHub.

GLM-5.2 ships 1M usable context for long coding agent runs
5 min read

GLM-5.2 ships 1M usable context for long…

Z.ai's GLM-5.2 open-weight flagship targets million-token coding trajectories with MIT weights, High and Max reasoning modes, and Terminal-Bench scores that jump from 62.0 to 81.0 versus GLM-5.1.

Cohere North Mini Code is a 30B MoE you can self-host for agentic coding
4 min read

Cohere North Mini Code is a 30B…

Cohere's first open-weight coding model activates 3B of 30B parameters per token, ships under Apache 2.0, and targets terminal agents. Here is when I would run it locally instead of a frontier API.

Frontier models are too expensive for agent loops. Specialized models are closing the gap.
7 min read

Frontier models are too expensive for agent…

Composer 2.5 scores 62 on the Coding Agent Index at $0.07 per task while Opus 4.7 costs $4.10. Here's the hybrid routing math I use when agent loops would bankrupt a frontier-only stack.

Tinker lets you fine-tune big models without owning the GPU cluster
5 min read

Tinker lets you fine-tune big models without…

Thinking Machines Lab's Tinker API runs distributed LoRA training while you write a normal Python loop on your laptop. Here's how it fits the Cursor playbook for teams that are not Cursor.