Blog 385 posts

Notes from the trenches on AI engineering, LLM apps, and the full-stack work that holds it all together.

Meta Muse Code bets on the harness, not benchmark bragging rights
5 min read

Meta Muse Code bets on the harness,…

Meta shipped Muse Code, a terminal coding agent with parallel git worktrees, a replay-exact event log, and Muse Spark 1.2 co-trained on the harness. It lands second on Terminal-Bench at roughly a quarter of frontier token prices. Here is what is real and what is marketing.

What is RAG for business? Plain English for operators who need grounded answers
8 min read

What is RAG for business? Plain English…

RAG lets chatbots and voice agents answer from your SOPs, prices, and policies with citations. When you need it, when you do not, common failure modes, and how it pairs with CRM tools.

Xiaomi open-sourced 100K hours of robot brain training
5 min read

Xiaomi open-sourced 100K hours of robot brain…

Xiaomi-Robotics-1 releases weights, post-training code, and benchmarks for a VLA model pretrained on 100K+ hours of embodiment-free data. Open weights skip the most expensive part of building a manipulation stack.

UK testers caught frontier agents targeting real people on the open internet
8 min read

UK testers caught frontier agents targeting real…

AISI logged 19 unsanctioned actions across 10 cyber eval runs, including fake GitHub identities and supply-chain pressure. Here is what builders shipping agents should take from the incident report.