
0xSero's REAP-pruned Kimi-K2.6 trades size for agentic…
The Kimi-K2.6-519B-NVFP4 checkpoint prunes MoE experts with Cerebras REAP rules. It is strong on code, math, and tool calls, but you must keep outputs bounded to avoid repetition loops.

The Kimi-K2.6-519B-NVFP4 checkpoint prunes MoE experts with Cerebras REAP rules. It is strong on code, math, and tool calls, but you must keep outputs bounded to avoid repetition loops.

Meta and Oxford's VGGT-Omega is a CVPR 2026 oral that cuts GPU training memory by roughly 70%, scales to 10B parameters, and beats optimization pipelines on dynamic scenes. Here's what changed and how I'd evaluate it before betting a product on feed-forward 3D.

Lance is ByteDance's Apache 2.0 unified multimodal model at 3B active parameters. It handles captioning, VQA, text-to-video, and multi-turn edits in one framework, trained on 128 A100s. Here's what that means if you ship generative media.

Xiaomi-Robotics-1 is a 5B VLA model pretrained on 100K+ hours of UMI trajectories, then post-trained on real robots. Weights, inference code, and benchmarks are on Hugging Face under Apache 2.0.

Z.ai shipped GLM-5.3 on the same weights as GLM-5.2 and jumped from 4.6 to 28.3 on Terminal-Bench 3.0. The lesson for builders: post-training and harness fit beat another pre-training run.

Qwen3.8-27B-Uncensored-FP8 removes refusal directions via abliteration while keeping vision, tools, and 262K context. Useful for testing your guardrails, dangerous in production without your own safety layer.

Xiaomi-Robotics-1 releases weights, post-training code, and benchmarks for a VLA model pretrained on 100K+ hours of embodiment-free data. Open weights skip the most expensive part of building a manipulation stack.

LFM2.5-2.6B is a 2.6B open-weight agent model that stays under 2.5GB, hits 220 tok/s on an M5 Max, and beats Qwen3.5-9B on most tool-use benchmarks. Here is how to wire it into Hermes, OpenClaw, or Pi through a local OpenAI-compatible endpoint.

IMLE-based generative MPC hits competitive offline RL scores while sampling trajectories over 20x faster than Diffuser. Real robots can replan around people without waiting on denoising.

Google’s Gemma 4 12B is multimodal with a 256K context window. Unsloth’s Dynamic GGUFs squeeze a usable 4-bit build into roughly 8GB of memory without throwing quality off a cliff.

Shanghai AI Lab's 35B MoE agent reaches trillion-parameter benchmark territory by scaling trajectory length to 45K tokens and distilling six domain teachers. Here's what agent-horizon scaling actually means.

GLOSSOPETRAE generates procedural coding languages from a seed. At full opacity, human legibility drops to ~15% while Opus and GPT hit 97-100% task accuracy. Human readability hurts model performance.

VIMPO derives a policy-implied value function from KL-regularized RL optimality conditions. It improves over GRPO on AIME and OlympiadBench while staying critic-free. Code on GitHub.

Z.ai's GLM-5.2 open-weight flagship targets million-token coding trajectories with MIT weights, High and Max reasoning modes, and Terminal-Bench scores that jump from 62.0 to 81.0 versus GLM-5.1.

WeiboAI's MIT-licensed VibeThinker-3B scores 94.3 on AIME26 and 96.1% on post-cutoff LeetCode contests. It trails frontier models on knowledge-heavy GPQA by design, not by accident.

Cohere's first open-weight coding model activates 3B of 30B parameters per token, ships under Apache 2.0, and targets terminal agents. Here is when I would run it locally instead of a frontier API.

Composer 2.5 scores 62 on the Coding Agent Index at $0.07 per task while Opus 4.7 costs $4.10. Here's the hybrid routing math I use when agent loops would bankrupt a frontier-only stack.

Thinking Machines Lab's Tinker API runs distributed LoRA training while you write a normal Python loop on your laptop. Here's how it fits the Cursor playbook for teams that are not Cursor.