All posts

Models & Tooling 112 posts

Every post filed under Models & Tooling, newest first.

Apple cut 200+ jobs on Vision Pro and Siri days before the Mac mini AI push
5 min read

Apple cut 200+ jobs on Vision Pro…

Bloomberg says Apple eliminated more than 200 roles across Vision Pro gaming, Immersive Video, Siri, and Intelligent Systems Experience as Siri AI and smart glasses take priority. Vision Pro stays, but the org chart is shifting fast.

Apple's new Mac mini is pitching itself as an always-on agent box
6 min read

Apple's new Mac mini is pitching itself…

The M6 Mac mini starts at $899 with up to 4x faster on-device AI than M4, 64GB unified memory on M5 Pro, and Thunderbolt clustering for larger local models. Apple is finally naming the use case developers already bought it for.

Nvidia's 1B Nemotron embed model is built for multilingual RAG at 8K context
4 min read

Nvidia's 1B Nemotron embed model is built…

llama-nemotron-embed-1b-v2 ships Matryoshka 2048-dim vectors, 26-language eval coverage, and commercial-friendly NeMo Retriever licensing for long-document QA retrieval.

Nvidia's $6B Poolside deal is a bet on open-weight Nemotron, not another chat app
3 min read

Nvidia's $6B Poolside deal is a bet…

Nvidia licensed Poolside's Model Factory for $6 billion, invested $1 billion at a $12B valuation, and hired 109 engineers to chase frontier open-weight models that compete with DeepSeek and Kimi K3.

Ox Alpha is free on OpenRouter and nobody will say who built it
4 min read

Ox Alpha is free on OpenRouter and…

A stealth coding model with a 1M-token window landed on OpenRouter August 20 with zero lab name attached. Community forensics point at Zhipu GLM infrastructure, and that raises real routing questions for production code.

Starcloud raised $250M to put Nvidia GPUs in orbit
5 min read

Starcloud raised $250M to put Nvidia GPUs…

Starcloud's Series A extension values the orbital data center startup at $2.3B. Nvidia and Cisco joined Manhattan West on a round meant to fund manufacturing, launches, and the Vera Rubin Space-1 chip partnership.

Dario Amodei says AI trust will not come from marketing. Only real wins will.
4 min read

Dario Amodei says AI trust will not…

Anthropic's CEO made a rare X appearance to push back on critics who say his safety warnings backfired. His bet: medicine and biology results will move public opinion more than any PR campaign.

Brad Lightcap is leaving OpenAI. What the COO exit signals before the IPO
4 min read

Brad Lightcap is leaving OpenAI. What the…

OpenAI's eight-year operator Brad Lightcap announced he is starting something new in August 2026, weeks after stepping back from COO. Here is what the executive churn means for builders betting on ChatGPT at scale.

Meta Muse Code bets on the harness, not benchmark bragging rights
5 min read

Meta Muse Code bets on the harness,…

Meta shipped Muse Code, a terminal coding agent with parallel git worktrees, a replay-exact event log, and Muse Spark 1.2 co-trained on the harness. It lands second on Terminal-Bench at roughly a quarter of frontier token prices. Here is what is real and what is marketing.

FLUX 3 Video ships 20-second HD clips with native audio from one model
4 min read

FLUX 3 Video ships 20-second HD clips…

Black Forest Labs opened FLUX 3 Video for text and image to video up to 20 seconds at HD, with dialogue, lip sync, multi-shot scenes, and draft mode for cheap iteration. Open weights are still on the roadmap.

Cursor cut cloud agent tokens 30% by fixing how MCPs and skills load
5 min read

Cursor cut cloud agent tokens 30% by…

Cursor's August 2026 cloud agent update optimizes MCP tool schemas, skills injection, and computer-use loops. The team reports up to 30% lower token usage and 80% better computer-use efficiency.

DiffusionGemma hits ~1,500 TPS on one H100: when a diffusion LLM beats autoregressive serving
8 min read

DiffusionGemma hits ~1,500 TPS on one H100:…

Google's open-weight DiffusionGemma denoises 256-token blocks in about 12 passes, reaching roughly 1,500 output tokens per second on a single H100 at batch size 1. Here is where that speed matters, what you trade away, and how to serve it with vLLM.

RAMageddon is here: AI data centers are eating the chips your laptop needs
5 min read

RAMageddon is here: AI data centers are…

Apple's MacBook Air is slipping into late-August delivery windows as AI hyperscalers soak up DRAM supply. Gartner forecasts 17% higher PC prices in 2026 and says sub-$500 laptops may vanish by 2028. If you build or buy AI infra, memory is now a strategic constraint.

Qwen3.8-Max is a 2.4T MoE that codes for 16 days straight
4 min read

Qwen3.8-Max is a 2.4T MoE that codes…

Alibaba's Qwen3.8-Max packs 95B active parameters into a 2.4T MoE stack, ranks ahead of Claude Fable 5 on WebDev Arena, and prices at $2/$6 per million tokens. Weights hit Hugging Face next week.

Grok Voice Think Fast 2.0 hits 0.70s to first audio. Should you switch?
5 min read

Grok Voice Think Fast 2.0 hits 0.70s…

xAI's new speech-to-speech model tops Artificial Analysis benchmarks on quality and latency at $0.08 per audio minute. What voice agent builders should test before the August 5 default migration.

Gemini 3.5 Flash Cyber pairs cheap models with CodeMender for defender-scale scanning
4 min read

Gemini 3.5 Flash Cyber pairs cheap models…

Google DeepMind shipped Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber-specialist 3.5 Flash Cyber inside CodeMender. Defenders get a limited pilot; builders should note the dual-use deployment model.

Four AIs hit 42/42 on IMO 2026. The headline number is already saturated
5 min read

Four AIs hit 42/42 on IMO 2026.…

Claude Fable 5, GPT-5.6 Sol, Kimi K3, and AxiomProver all reported perfect IMO 2026 scores. The interesting part is cost, grading tier, and what happens when benchmarks stop separating models.

T3 Code now runs Grok Build on your SuperGrok subscription
4 min read

T3 Code now runs Grok Build on…

Theo Browne's T3 Code GUI added Grok via ACP and X or SuperGrok OAuth. No API key, same subscription you already pay for, alongside Claude Code and Codex in one dashboard.

Krea 2 shrinks to 12GB so consumer GPUs can run aesthetic image gen
4 min read

Krea 2 shrinks to 12GB so consumer…

Krea AI open-sourced Krea 2, a 12.9B-parameter diffusion transformer trained from scratch. Turbo runs in 8 steps on GGUF quants that fit 12GB VRAM. Raw is the fine-tune base; Turbo is what you ship.

AgentDeck turns Stream Deck+ into a physical Claude Code control panel
3 min read

AgentDeck turns Stream Deck+ into a physical…

AgentDeck hooks Claude Code's event stream into Elgato Stream Deck+ buttons and dials. See session status, approve permissions, and monitor token spend without tab-switching. Open source with a Marketplace plugin for macOS and Windows 11.

OpenAI trained GPT-5.5 Instant with 600 doctors and cut health errors 71%
4 min read

OpenAI trained GPT-5.5 Instant with 600 doctors…

OpenAI's June health intelligence update pairs GPT-5.5 Instant with a global physician network across 60 countries. Production monitors show 71% fewer flagged health factuality issues, with panel ratings beating older models and physician-written answers on several dimensions.

LocalAI ships ByteDance depth estimation in C++ that beats PyTorch on CPU
4 min read

LocalAI ships ByteDance depth estimation in C++…

depth-anything.cpp ports ByteDance Depth Anything 3 to ggml with no Python at inference. On CPU it runs 1.31x faster than PyTorch at q8_0, uses half the RAM, and loads 6.7x faster. LocalAI v4.5 exposes it via POST /v1/depth.

MIT's SMT trains RNNs in parallel without backpropagation through time
4 min read

MIT's SMT trains RNNs in parallel without…

Supervised Memory Training uses a Transformer teacher to label optimal memory states, then trains nonlinear RNNs with one-step supervision. You get O(1) gradient paths and time-parallel pretraining without unrolling the full sequence.

Meshy T2 turns one photo into a clean 3D mesh in six seconds
5 min read

Meshy T2 turns one photo into a…

Meshy T2 uses flow matching to generate vertices and connectivity in parallel, not autoregressive mesh tokens. Median image-to-mesh latency is six seconds with controllable face budgets and native multi-part output.

Kimi K2.7 Code thinks 30% less and still ships harder on long coding tasks
5 min read

Kimi K2.7 Code thinks 30% less and…

Moonshot's open-weight Kimi K2.7 Code keeps the 1T MoE backbone but cuts thinking tokens ~30% versus K2.6 while jumping +21.8% on Kimi Code Bench v2. Here is when I would route agents to it.

OpenAI finally lets you bank Codex rate limit resets for when you actually code
5 min read

OpenAI finally lets you bank Codex rate…

Codex rate limits used to reset on OpenAI's clock, often at 3am. Banked resets let Go, Plus, Pro, and Business users save and trigger them manually, with a 30-day expiry and referral bonuses through June 24.

Kimi Work puts 300 parallel agents on your desktop, not in a cloud sandbox
4 min read

Kimi Work puts 300 parallel agents on…

Moonshot's Kimi Work desktop agent reads local files, drives your real browser, and spins up to 300 sub-agents per task. Here is how Agent Swarm compares to cloud-only coding agents I deploy for clients.