All posts

Applied AI 121 posts

Every post filed under Applied AI, newest first.

Self-Harness: how agents rewrite their own operating rules without retraining
5 min read

Self-Harness: how agents rewrite their own operating…

Shanghai AI Lab's Self-Harness lets a fixed model improve its own agent scaffolding through weakness mining, targeted edits, and regression gates. Here is what the Terminal-Bench numbers mean and how to run a lightweight version today.

AI agent credentials belong in a vault, not a .env file on someone's laptop
6 min read

AI agent credentials belong in a vault,…

Twelve million servers leak .env files to the open web. When you give an agent Gmail, CRM, and Slack access, local plaintext tokens turn a config mistake into a company-wide breach.

Anthropic put $10M CAD in Claude credits into eight Canadian research labs
4 min read

Anthropic put $10M CAD in Claude credits…

The July 2026 commitment funds Amii, Mila, Vector, CHEO, CAMH, Université Laval, U of T, and U of Saskatchewan with no-strings Claude API credits plus startup program access for affiliated founders.

Chinese humanoids ran 9.39 seconds in the 100m, then face-planted into the mats
4 min read

Chinese humanoids ran 9.39 seconds in the…

At Beijing's World Humanoid Robot Games, Tiangong Ultra clocked 9.39s in the 100m, faster than Usain Bolt's 9.58s world record. I broke down what the sprint times actually mean for factory deployment and why braking still looks unsolved.

Browse-capable AI agents turn every webpage into a prompt injection surface
6 min read

Browse-capable AI agents turn every webpage into…

When an agent can fetch URLs, read local files, and send messages in one session, a malicious page can steer all three. Promptfoo's OpenClaw lab shows why browsing and outbound action must not share one trust boundary.

Greptile TREX: why AI code review needs runtime proof, not predictions
6 min read

Greptile TREX: why AI code review needs…

Greptile's TREX layer runs PR branches in sandboxes and attaches logs, screenshots, and traces to review comments. Here is what that means for teams shipping with Cursor, Claude Code, and other agentic coding tools.

Hiring Agent turns resume PDFs and GitHub signals into explainable scores
4 min read

Hiring Agent turns resume PDFs and GitHub…

The open-source Hiring Agent pipeline parses PDFs to JSON Resume, enriches with GitHub repo metrics, and scores candidates with role-specific rubrics. Here is the architecture and where I would add guardrails.

Nvidia's 1B Nemotron embed model is built for multilingual RAG at 8K context
4 min read

Nvidia's 1B Nemotron embed model is built…

llama-nemotron-embed-1b-v2 ships Matryoshka 2048-dim vectors, 26-language eval coverage, and commercial-friendly NeMo Retriever licensing for long-document QA retrieval.

Why solo AI agents break when you add a second user
6 min read

Why solo AI agents break when you…

Solo coding agents like Claude Code and Cursor are brilliant for one power user. At team scale, context compaction, siloed memory, and local credentials turn small wins into operational risk.

The FDA just cleared a robot that draws blood without a human holding the needle
4 min read

The FDA just cleared a robot that…

Vitestro's Aletta received FDA De Novo authorization for autonomous phlebotomy: vein imaging, needle insertion, tube swaps, and bandaging. I unpacked the trial numbers and what a 1-to-3 staffing model means for clinics.

Waymo built a custom 5nm chip because robotaxi compute is not a data center problem
4 min read

Waymo built a custom 5nm chip because…

Waymo revealed a purpose-built 5nm ASIC for front-end sensor processing, 1,000+ TOPS onboard, and 20x compute scaling in eight years. I mapped what that means for anyone shipping physical AI at the edge.

DoorDash pays Dashers $5 to load Dot robots. The last 10 feet still need humans.
5 min read

DoorDash pays Dashers $5 to load Dot…

Dot can drive 20 mph on roads and sidewalks, but it cannot pick up a bag at the counter. DoorDash's Phoenix pilot pays gig workers to bridge that gap. That handoff problem shows up in every ops automation project.

SpaceX's first earnings call pitched robot factories on the Moon. Investors wanted capex numbers.
5 min read

SpaceX's first earnings call pitched robot factories…

On August 4, 2026, Elon Musk told SpaceX shareholders that humanoids would build lunar factories, solar arrays, and a mass driver. The stock fell 5%+. The gap between sci-fi roadmap and disclosed spending is the engineering story.

Three ex-SpaceX engineers opened a robot welding factory for AI data center steel
4 min read

Three ex-SpaceX engineers opened a robot welding…

Cincinnati startup 1872 raised $15M to automate steel skid fabrication with Path Robotics welders and an AI-native Factory OS. The team targets 80% autonomous ops by 2027 while welding costs drop from $0.78 to $0.12 per inch.

Amazon Prime Air wants 500 US cities by year-end. The ops math is finally catching the 2013 keynote.
4 min read

Amazon Prime Air wants 500 US cities…

Prime Air will expand from 11 drone sites to nearly 500 US cities and towns in 2026, with tens of millions of customers eligible for 30-minute deliveries. Here's how hub radius, pricing, and Walmart's drone race change last-mile automation.

Claude Managed Agents now cap web domains and show per-thread cost in the Console
4 min read

Claude Managed Agents now cap web domains…

August 2026 Managed Agents updates add allowed_domains and blocked_domains on web_search and web_fetch, memory on self-hosted sandboxes, and a Console Inspector with per-thread cost. Pair with session budgets before you run unattended fleets.

BrowserCode turns CDP into a coding primitive for browser-native agents
5 min read

BrowserCode turns CDP into a coding primitive…

BrowserCode (bcode.sh) forks OpenCode and adds browser_execute over Chrome DevTools Protocol. Reusable scripts land in .bcode/agent-workspace/. Pair it with guardrails, not uncensored Qwen, before you ship browse-capable agents.

Cursor cloud agents can wait for CI, Slack, and cron now
6 min read

Cursor cloud agents can wait for CI,…

Cursor's August 19 cloud agent update adds event subscriptions, isolated subagent VMs, /goal for long-lived objectives, and Custom Modes. Here is how I wire those pieces into a shipping loop.

H Company's Holo agents hit 80.4% OSWorld with MCP, CLI, and no loop to write
4 min read

H Company's Holo agents hit 80.4% OSWorld…

H Company ships managed computer-use agents with Holo3 at 80.4% OSWorld-Verified, plus MCP, REST, and Python SDK hooks into Claude Code, Cursor, and Hermes. Here is when I pick it over rolling my own browser loop.

Tesla's Cybercab is rolling onto Austin roads without a steering wheel. That is the easy part.
4 min read

Tesla's Cybercab is rolling onto Austin roads…

Tesla plans employee rides on public roads in Austin, then fold purpose-built Cybercabs into its Robotaxi service days later. The vehicle has no wheel or pedals. Scaling, federal exemptions, and miles still lag Waymo.

Unitree's 460% IPO pop made a $16B robot king. The software moment is still years out.
5 min read

Unitree's 460% IPO pop made a $16B…

China's first listed humanoid maker closed 460% above its IPO price on Shanghai's STAR Market, valuing Unitree near $50B. Founder Wang Xingxing says the real ChatGPT moment for robotics is 2 to 10 years away. Here's what the prospectus and debut actually tell builders.

Dario Amodei says AI trust will not come from marketing. Only real wins will.
4 min read

Dario Amodei says AI trust will not…

Anthropic's CEO made a rare X appearance to push back on critics who say his safety warnings backfired. His bet: medicine and biology results will move public opinion more than any PR campaign.

Hermes /loop gives agents a heartbeat without cron jobs
4 min read

Hermes /loop gives agents a heartbeat without…

Nous Research shipped /loop so Hermes re-runs prompts on a timer inside your session, with backoff and real stop conditions. It is cron with memory, and it changes how you monitor long agent jobs.

AI agents on a bot-only RuneScape server invented a barter economy
4 min read

AI agents on a bot-only RuneScape server…

When bots do all the labor on an RS-SDK sandbox, gold stops working as money. Rare spawns like runite ore became currency. A weird game experiment with real lessons for multi-agent systems.

Agility's Digit V5 drops the safety fence. Factory humanoids are entering open floor work
4 min read

Agility's Digit V5 drops the safety fence.…

Digit V5 is built to work shoulder-to-shoulder with people without safety caging, with 20+ hour battery shifts and first customer shipments in December 2026. I mapped what fenceless humanoids change for factory ops and pilot ROI math.

Honor's robot phone puts a motorized gimbal inside a flagship. Embodied AI just got pocket-sized
5 min read

Honor's robot phone puts a motorized gimbal…

Honor's Robot Phone ships in China with a titanium 4DoF camera arm, dual 200MP lenses, and YOYO Robot Mode. I broke down what a shipping embodied-AI phone means for perception stacks and third-party dev APIs.

Mitsubishi is retooling an engine plant to build 1,000 humanoids a month. Japan's old factories are the new robot on-ramp
4 min read

Mitsubishi is retooling an engine plant to…

Mitsubishi Motors partnered with University of Tokyo spinout Highlanders to mass-produce HL Human humanoids at a former Kyoto ICE engine plant, targeting up to 1,000 units per month from early 2027. I broke down the builder-plus-buyer model automakers use to skip endless pilots.

Northrop's robot space mechanic bolts propulsion pods onto aging satellites. In-orbit repair is becoming a business
4 min read

Northrop's robot space mechanic bolts propulsion pods…

Northrop Grumman's Mission Robotic Vehicle will attach modular propulsion pods to the Optus satellite in 2027, adding roughly six more years of life. I explained how MRV plus MEP pods change satellite servicing economics versus parking whole servicers behind one customer.

Brad Lightcap is leaving OpenAI. What the COO exit signals before the IPO
4 min read

Brad Lightcap is leaving OpenAI. What the…

OpenAI's eight-year operator Brad Lightcap announced he is starting something new in August 2026, weeks after stepping back from COO. Here is what the executive churn means for builders betting on ChatGPT at scale.

Grok Bot turns xAI into a group chat of always-on agent teammates
5 min read

Grok Bot turns xAI into a group…

Grok Bot gives each agent its own cloud computer, iMessage-style messaging, and parallel specialist lanes. Here is what builders should steal from the beta launch.

Google DeepMind's safety team doesn't trust its own HR AI filters
5 min read

Google DeepMind's safety team doesn't trust its…

DeepMind's AGI Safety and Alignment Team told job applicants there is a non-trivial probability automated screening will reject them incorrectly. They built a bypass form. If Google won't bet on its own filters, you shouldn't either.

31 million tests later, noRecognition beat Flock cameras at DefCon
5 min read

31 million tests later, noRecognition beat Flock…

Security researcher Bill Swearingen trained reinforcement-learning patterns that block ALPR and surveillance detection without hiding from video. At DefCon he wrapped a Toyota Yaris and drove past a Flock camera. Here is what builders on both sides should learn.

YouTube doubled YPP gates: 8K hours or 20M Shorts views to earn ads
5 min read

YouTube doubled YPP gates: 8K hours or…

Starting February 1, 2027, new YouTube creators need 8,000 watch hours or 20 million Shorts views in 90 days before ad and Premium revenue sharing kicks in. Existing partners face new activity floors too. Here is what changed and who it hits hardest.

Claude Code auto mode is the default now. Humans caught 14% of dangerous commands.
4 min read

Claude Code auto mode is the default…

Anthropic made auto mode the default on Pro, Max, and Team plans after a 1,053-tester study. The classifier blocked 89% of dangerous shell commands while manual approval fatigue dropped human catches to 5%.

Claude Code sessions can message each other now. I stopped copy-pasting between terminals.
4 min read

Claude Code sessions can message each other…

Cross-session messaging in Claude Code v2.1.224 lets independent terminals share plain-text notes locally. Here is how ListAgents and SendMessage work, what stays off the wire, and when I still use agent teams instead.

Uber is betting $10B that it can own the rider, not the robot
6 min read

Uber is betting $10B that it can…

Uber pledged more than $10 billion and 120,000 partner vehicles to stay the default robotaxi app. I broke down why that flips its asset-light model and what builders should watch as Waymo pulls away.

Agent Plugins 1.0: build your agent skills once, ship to Cursor and Copilot
4 min read

Agent Plugins 1.0: build your agent skills…

Agent Plugins 1.0 packages skills and MCP servers into one portable directory. Amazon, Cursor, Google, Microsoft, OpenAI, and Vercel back the spec. Here is the folder layout and what it means for real projects.

Hadrian hit $8B by putting factory skills into software (Opus)
6 min read

Hadrian hit $8B by putting factory skills…

Defense manufacturing startup Hadrian raised $1.37B at a $7.9B valuation. Its Opus platform automates CNC programming, inspection, and scheduling so new workers ship parts in 30 days. That is applied AI shipping in the physical world.

Voice mode finally works when you feed it your context files
5 min read

Voice mode finally works when you feed…

OpenAI and Anthropic upgraded live voice this year, but the real shift is context: voice that reads your docs, uses your frameworks, and pushes back. Here is the checklist I use for client voice workflows and my own thinking sessions.

DoorDash pays $5 for the one thing Dot cannot do
5 min read

DoorDash pays $5 for the one thing…

Phoenix Dashers load restaurant orders into Dot robots for about five dollars and five minutes. The robot can drive 20 mph with lidar, but the last few feet from counter to curb still need a human.

Meta Muse Code bets on the harness, not benchmark bragging rights
5 min read

Meta Muse Code bets on the harness,…

Meta shipped Muse Code, a terminal coding agent with parallel git worktrees, a replay-exact event log, and Muse Spark 1.2 co-trained on the harness. It lands second on Terminal-Bench at roughly a quarter of frontier token prices. Here is what is real and what is marketing.

Cursor cut cloud agent tokens 30% by fixing how MCPs and skills load
5 min read

Cursor cut cloud agent tokens 30% by…

Cursor's August 2026 cloud agent update optimizes MCP tool schemas, skills injection, and computer-use loops. The team reports up to 30% lower token usage and 80% better computer-use efficiency.

HeyGen's founder left an AI clone on sales calls. It closed 132 deals and invented a $4,800 plan.
5 min read

HeyGen's founder left an AI clone on…

Wayne Liang paired a HeyGen avatar with an OpenClaw agent during paternity leave. Eight weeks, 2,741 prospect calls, 132 paid customers, and a handful of rogue pricing mistakes that only guardrails fixed.

Business students treat AI like a job requirement now, not a bonus skill
5 min read

Business students treat AI like a job…

A three-year Kogod School of Business survey shows 80%+ of students use AI weekly, employer interview questions about AI skills nearly quadrupled, and the top worry is cognitive devaluation, not bans.

86% of finance executives say AI skills beat an MBA for new hires
5 min read

86% of finance executives say AI skills…

PwC surveyed 1,004 US financial services directors and found 91% raising pay for AI skills, 86% valuing AI training over MBAs for many roles, and 77% still unable to prove ROI on most AI spend.

I run status reports by voice in ChatGPT. Here is the folder scaffold that makes it repeatable.
4 min read

I run status reports by voice in…

ChatGPT Voice plus Projects can turn a spoken update into a Markdown draft, a finished PDF, and a team handoff without touching the keyboard. The trick is scaffolding folders once.

Flock's ALPR network hit 20B scans a month. Then officers started stalking exes.
6 min read

Flock's ALPR network hit 20B scans a…

At least 50 U.S. police officers have been accused of misusing Flock Safety's license plate readers, including 46 on Flock's own network. A Roseville audit found 71% false alerts. If you ship AI in production, this is what happens when access controls lag behind scale.

AI screened 500 million enzyme variants to reverse a mark of aging in human skin
6 min read

AI screened 500 million enzyme variants to…

Researchers used AlphaFold and directed evolution to build CMLase, an enzyme that stripped aging damage from 75-year-old skin tissue down to levels seen in 31-year-old samples. Proof of concept, not a cream yet, but a real applied-AI win in protein engineering.

China's 582-tonne fusion magnet is heavier than a loaded 747
4 min read

China's 582-tonne fusion magnet is heavier than…

The Institute of Plasma Physics accepted a toroidal field coil for China's CRAFT fusion program that is 1.3 times the volume of ITER magnets and stores three times the energy. Full-load testing passed in Hefei. Fusion power by 2030 is still a bet, but the hardware stack is real.

Every population may carry DNA from a ghost hominin lineage
4 min read

Every population may carry DNA from a…

A Science study analyzed 503 modern genomes with a new TRACE model and found archaic ancestry that matches no known Neanderthal or Denisovan sequence in every population tested. Roughly 0.5 to 1 percent of non-African genomes may come from a branch that split off more than 500,000 years ago.

Tanruprubart may become the first targeted Guillain-Barré treatment
5 min read

Tanruprubart may become the first targeted Guillain-Barré…

Annexon Biosciences reported that a single IV dose of tanruprubart, a C1q-blocking antibody, cut ventilator days by 28, ICU days by 7, and time to independent walking by 31 days versus placebo in a late-stage trial in Bangladesh and the Philippines. EMA review is underway for possible 2027 approval.

Where is the AI speedometer? Why CFOs need real-dollar chat costs
5 min read

Where is the AI speedometer? Why CFOs…

Enterprise AI subsidies are expiring and token bills are jumping. Finance teams still only get a kill switch, not a speedometer. What to measure before your million-dollar budget blows up.

Raft turns your ChatGPT and Claude Code subs into a named agent team
5 min read

Raft turns your ChatGPT and Claude Code…

Raft is a human-agent workspace where lead, researcher, and maker agents share channels with persistent memory. Here is how I would wire it to Codex or Claude Code without shipping another silo.

Sam Altman on Capitol Hill: pacing AI after the rogue agent breach
4 min read

Sam Altman on Capitol Hill: pacing AI…

After Modal's second victim and 17,600 hostile agent actions, Altman met senators about unreleased models while Trump floated controls and a White House vetting framework lands August 1.

Turn a photo of your handwriting into a real font with a Claude Code skill
4 min read

Turn a photo of your handwriting into…

The open-source draw-your-font project packages handwriting capture into TTF, WOFF, and WOFF2 files. Install it as a Claude Code skill, upload a photo, and ship a custom font in one session.

Claude Managed Agents now pin effort, seed 50 events, and webhook the fleet
4 min read

Claude Managed Agents now pin effort, seed…

July 2026 Managed Agents updates add per-agent effort levels, session seeding with up to 50 initial events, environment and memory-store webhooks, and sub-agent thread streaming. Skills still cap at 500 per session across all agents.

Claude Code's security plugin scans your diff before you merge
4 min read

Claude Code's security plugin scans your diff…

Anthropic shipped the Claude Security plugin for Claude Code in beta: multi-agent scans in your terminal, verified findings, and patches you apply yourself. It stacks with the security-guidance hook that flags eval and innerHTML as you type.

Gemini 3.5 Flash Cyber pairs cheap models with CodeMender for defender-scale scanning
4 min read

Gemini 3.5 Flash Cyber pairs cheap models…

Google DeepMind shipped Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber-specialist 3.5 Flash Cyber inside CodeMender. Defenders get a limited pilot; builders should note the dual-use deployment model.

Four AIs hit 42/42 on IMO 2026. The headline number is already saturated
5 min read

Four AIs hit 42/42 on IMO 2026.…

Claude Fable 5, GPT-5.6 Sol, Kimi K3, and AxiomProver all reported perfect IMO 2026 scores. The interesting part is cost, grading tier, and what happens when benchmarks stop separating models.

Claude Cowork background tasks now run on web and mobile while your laptop is closed
5 min read

Claude Cowork background tasks now run on…

Anthropic moved Claude Cowork sessions to the cloud so tasks keep running after you close your laptop. Scheduled jobs can fire with no device online, and you can steer from your phone when Claude needs a decision.

Shepherd brings Git-style fork and replay to live AI agent runs
4 min read

Shepherd brings Git-style fork and replay to…

Stanford and Northeastern researchers released Shepherd, a Python runtime that records agent runs as forkable execution traces. Reported results include 5x faster forks than Docker and 95% KV-cache reuse on replay.

Claude Design finally imports your real design system (and checks its own work)
5 min read

Claude Design finally imports your real design…

Anthropic rebuilt Claude Design so prototypes start from your GitHub components, auto-correct against your tokens, and hand off to Claude Code without a screenshot rebuild. Here's what that means if you ship UI for clients.

LocalAI ships ByteDance depth estimation in C++ that beats PyTorch on CPU
4 min read

LocalAI ships ByteDance depth estimation in C++…

depth-anything.cpp ports ByteDance Depth Anything 3 to ggml with no Python at inference. On CPU it runs 1.31x faster than PyTorch at q8_0, uses half the RAM, and loads 6.7x faster. LocalAI v4.5 exposes it via POST /v1/depth.

MIT's SMT trains RNNs in parallel without backpropagation through time
4 min read

MIT's SMT trains RNNs in parallel without…

Supervised Memory Training uses a Transformer teacher to label optimal memory states, then trains nonlinear RNNs with one-step supervision. You get O(1) gradient paths and time-parallel pretraining without unrolling the full sequence.

Claude artifacts can persist data now. The leaked Fable 5 prompt explains how.
5 min read

Claude artifacts can persist data now. The…

Anthropic's Claude Fable 5 system prompt leak reveals window.storage, a key-value API for artifacts that remember data between chats. Here is what builders can actually do with it.

Claude Fable 5 spent 1.4M tokens designing a humanoid robot. What shipped?
4 min read

Claude Fable 5 spent 1.4M tokens designing…

Jake Fitzgerald's viral demo used two hours and 1.4 million tokens to generate CAD-ready humanoid robot designs, kinematics, and animations. I broke down what is real versus render hype.

Claude Fable 5 built a playable Minecraft clone from one prompt. I checked the repo.
5 min read

Claude Fable 5 built a playable Minecraft…

Developers are shipping browser Minecraft clones with Claude Fable 5 in 20 to 40 minutes for roughly $12 to $30. The interesting part is not the game. It is the systems design the model held in one context.

MoneyPrinterTurbo turns one keyword into a short video pipeline you can self-host
4 min read

MoneyPrinterTurbo turns one keyword into a short…

The open-source MoneyPrinterTurbo repo chains LLM scripts, TTS, stock footage, and FFmpeg into finished 9:16 or 16:9 videos. Here is the architecture worth copying even if you never post on TikTok.

Cognition raised $1B because Devin now writes 89% of its own code
4 min read

Cognition raised $1B because Devin now writes…

Cognition closed a $1B Series D at a $26B valuation with $492M run-rate revenue. The clearest proof point is internal: 89% of Cognition's committed code now comes from Devin. Here's what that means if you ship software for a living.