
Claude Opus 4.8 on Azure now bills…
Anthropic brought Claude Opus 4.8 and Haiku 4.5 to Microsoft Foundry with Azure-native billing, Entra ID, and full prompt caching. Here is what enterprise teams should configure first.
Notes from the trenches on AI engineering, LLM apps, and the full-stack work that holds it all together.

Anthropic brought Claude Opus 4.8 and Haiku 4.5 to Microsoft Foundry with Azure-native billing, Entra ID, and full prompt caching. Here is what enterprise teams should configure first.

Cursor shipped a native iOS app for launching cloud agents, steering local runs with Remote Control, and merging PRs from your phone. Here is how I use it without losing my local MCP stack.

DeepSeek's DSpark speculative decoding framework adds a semi-autoregressive drafter and confidence-scheduled verification to V4 serving. Per-user generation runs 60-85% faster at matched throughput, lossless and open source.

MegaTrain stores weights in host memory and streams one layer at a time to the GPU, training 120B models on a single H200 with 1.5TB RAM. It beats DeepSpeed ZeRO-3 offload by 1.84x at 14B scale.

Brain2Qwerty v2 hits 78% word accuracy on the best participant using only a non-invasive MEG helmet. Meta open-sourced the training code. Here is what that means for applied AI and assistive tech.

Firecrawl, browser-use, Crawl4AI, MarkItDown, and Crawlee can build training and RAG pipelines without $2,000/month scraper contracts. Here is the stack I actually wire for clients.

Brian Armstrong's team runs 1,200 AI agents, defaults to GLM-5.2 and Kimi K2.7, and automated model selection. Five levers any engineering org can copy without a crypto-scale budget.

Z.ai's MIT-licensed GLM-5.2 ships 1M-token context, beats GPT-5.5 on several coding benches, and costs a fraction of Opus. Here's what production tests show, where it still breaks, and how I'd route it.