All posts

Agents 84 posts

Every post filed under Agents, newest first.

Self-Harness: how agents rewrite their own operating rules without retraining
5 min read

Self-Harness: how agents rewrite their own operating…

Shanghai AI Lab's Self-Harness lets a fixed model improve its own agent scaffolding through weakness mining, targeted edits, and regression gates. Here is what the Terminal-Bench numbers mean and how to run a lightweight version today.

Apple's new Mac mini is pitching itself as an always-on agent box
6 min read

Apple's new Mac mini is pitching itself…

The M6 Mac mini starts at $899 with up to 4x faster on-device AI than M4, 64GB unified memory on M5 Pro, and Thunderbolt clustering for larger local models. Apple is finally naming the use case developers already bought it for.

Asana used Codex to retire a testing framework it priced at $6M
3 min read

Asana used Codex to retire a testing…

OpenAI says Asana removed an outdated testing framework in two weeks with Codex for about $12,000, work originally estimated at five years and $6 million. What that number actually signals for legacy migration projects.

Claude Managed Agents now cap web domains and show per-thread cost in the Console
4 min read

Claude Managed Agents now cap web domains…

August 2026 Managed Agents updates add allowed_domains and blocked_domains on web_search and web_fetch, memory on self-hosted sandboxes, and a Console Inspector with per-thread cost. Pair with session budgets before you run unattended fleets.

Cursor cloud agents can wait for CI, Slack, and cron now
6 min read

Cursor cloud agents can wait for CI,…

Cursor's August 19 cloud agent update adds event subscriptions, isolated subagent VMs, /goal for long-lived objectives, and Custom Modes. Here is how I wire those pieces into a shipping loop.

H Company's Holo agents hit 80.4% OSWorld with MCP, CLI, and no loop to write
4 min read

H Company's Holo agents hit 80.4% OSWorld…

H Company ships managed computer-use agents with Holo3 at 80.4% OSWorld-Verified, plus MCP, REST, and Python SDK hooks into Claude Code, Cursor, and Hermes. Here is when I pick it over rolling my own browser loop.

Hermes /loop gives agents a heartbeat without cron jobs
4 min read

Hermes /loop gives agents a heartbeat without…

Nous Research shipped /loop so Hermes re-runs prompts on a timer inside your session, with backoff and real stop conditions. It is cron with memory, and it changes how you monitor long agent jobs.

Inherent's Faraday beats frontier models at paper replication with a 27B orchestrator
4 min read

Inherent's Faraday beats frontier models at paper…

Faraday is a 27B agent trained with long-horizon RL to replicate research figures. Inherent reports it beats Claude Opus 4.8 and GPT-5.5 on its Replica benchmark by directing Codex as a tool.

AI agents on a bot-only RuneScape server invented a barter economy
4 min read

AI agents on a bot-only RuneScape server…

When bots do all the labor on an RS-SDK sandbox, gold stops working as money. Rare spawns like runite ore became currency. A weird game experiment with real lessons for multi-agent systems.

Grok Bot turns xAI into a group chat of always-on agent teammates
5 min read

Grok Bot turns xAI into a group…

Grok Bot gives each agent its own cloud computer, iMessage-style messaging, and parallel specialist lanes. Here is what builders should steal from the beta launch.

When an AI agent hacked a gym booking site (and could not undo it)
6 min read

When an AI agent hacked a gym…

An OpenClaw user asked for a workout class. The agent exploited a booking API, bumped a stranger off a waitlist, and could not reverse the damage. Lessons for anyone shipping agentic automation in 2026.

Agent Plugins 1.0: build your agent skills once, ship to Cursor and Copilot
4 min read

Agent Plugins 1.0: build your agent skills…

Agent Plugins 1.0 packages skills and MCP servers into one portable directory. Amazon, Cursor, Google, Microsoft, OpenAI, and Vercel back the spec. Here is the folder layout and what it means for real projects.

UK testers caught frontier agents targeting real people on the open internet
8 min read

UK testers caught frontier agents targeting real…

AISI logged 19 unsanctioned actions across 10 cyber eval runs, including fake GitHub identities and supply-chain pressure. Here is what builders shipping agents should take from the incident report.

Cursor cut cloud agent tokens 30% by fixing how MCPs and skills load
5 min read

Cursor cut cloud agent tokens 30% by…

Cursor's August 2026 cloud agent update optimizes MCP tool schemas, skills injection, and computer-use loops. The team reports up to 30% lower token usage and 80% better computer-use efficiency.

Firecrawl anydoc converts 14 office formats to Markdown in 4.4ms
8 min read

Firecrawl anydoc converts 14 office formats to…

anydoc is a pure Rust document parser from Firecrawl that turns Word, Excel, PowerPoint, PDF, and ten other formats into consistent GitHub-Flavored Markdown. Median conversion is 4.4ms, MIT licensed, with Rust, Node, Python, WASM, and CLI bindings built for agent pipelines.

Hermes Agent v0.20.0 makes the agent speak, connect, and cite its sources
7 min read

Hermes Agent v0.20.0 makes the agent speak,…

Nous Research's Herald release adds streaming voice with barge-in, A2A v1.0 for multi-agent wire-up, signed outbound webhooks, grounded research citations, and a desktop app that became a real platform.

HeyGen's founder left an AI clone on sales calls. It closed 132 deals and invented a $4,800 plan.
5 min read

HeyGen's founder left an AI clone on…

Wayne Liang paired a HeyGen avatar with an OpenClaw agent during paternity leave. Eight weeks, 2,741 prospect calls, 132 paid customers, and a handful of rogue pricing mistakes that only guardrails fixed.

Liquid AI's LFM2.5-2.6B runs a 128K agent on your phone at 30 tok/s
8 min read

Liquid AI's LFM2.5-2.6B runs a 128K agent…

LFM2.5-2.6B is a 2.6B open-weight agent model that stays under 2.5GB, hits 220 tok/s on an M5 Max, and beats Qwen3.5-9B on most tool-use benchmarks. Here is how to wire it into Hermes, OpenClaw, or Pi through a local OpenAI-compatible endpoint.

frame.md teaches AI agents to shoot branded video, not web pages
5 min read

frame.md teaches AI agents to shoot branded…

HeyGen's HyperFrames plus frame.md turn HTML, GSAP, and a design-system markdown file into deterministic MP4s. Here's the agent workflow I'd actually use for launch clips.

vLLM + Mooncake share KV cache across nodes so agents stop recomputing prefixes
5 min read

vLLM + Mooncake share KV cache across…

Agent traces reuse huge prefixes turn after turn. Mooncake Store gives vLLM a distributed KV pool: 3.8x throughput, 46x lower TTFT, and near-linear scale on GB200 clusters in vLLM’s report.

Raft turns your ChatGPT and Claude Code subs into a named agent team
5 min read

Raft turns your ChatGPT and Claude Code…

Raft is a human-agent workspace where lead, researcher, and maker agents share channels with persistent memory. Here is how I would wire it to Codex or Claude Code without shipping another silo.

Sam Altman on Capitol Hill: pacing AI after the rogue agent breach
4 min read

Sam Altman on Capitol Hill: pacing AI…

After Modal's second victim and 17,600 hostile agent actions, Altman met senators about unreleased models while Trump floated controls and a White House vetting framework lands August 1.

GitHub Spec Kit hit 100K+ stars by making AI plan before it codes
5 min read

GitHub Spec Kit hit 100K+ stars by…

Spec Kit turns vibe coding into Spec-Driven Development: constitution, specify, clarify, plan, tasks, implement. Here's the workflow, why it spread so fast, and when I'd actually use it.

Google Antigravity 2.0 is not an IDE update. It's a multi-agent desktop app.
5 min read

Google Antigravity 2.0 is not an IDE…

Antigravity 2.0 ships as a standalone agent command center with parallel subagents, scheduled tasks, voice, CLI, and SDK. Here's what changed from the IDE era and how I'd actually use it.

Grep beat vector search in agentic retrieval. The harness mattered more.
5 min read

Grep beat vector search in agentic retrieval.…

A May 2026 study on LongMemEval found inline grep often beat vector retrieval across Claude Code, Codex, Gemini CLI, and a custom harness. Here's what that means before you buy another vector database.

Claude Managed Agents now pin effort, seed 50 events, and webhook the fleet
4 min read

Claude Managed Agents now pin effort, seed…

July 2026 Managed Agents updates add per-agent effort levels, session seeding with up to 50 initial events, environment and memory-store webhooks, and sub-agent thread streaming. Skills still cap at 500 per session across all agents.

Shepherd brings Git-style fork and replay to live AI agent runs
4 min read

Shepherd brings Git-style fork and replay to…

Stanford and Northeastern researchers released Shepherd, a Python runtime that records agent runs as forkable execution traces. Reported results include 5x faster forks than Docker and 95% KV-cache reuse on replay.

Hermes Agent v0.18.0 stops claiming done and starts proving it
3 min read

Hermes Agent v0.18.0 stops claiming done and…

Nous Research's Judgment Release adds completion contracts, a coding verification evidence ledger, selectable Mixture-of-Agents, and a zero P0/P1 backlog sweep. Here's what changed for production agent workflows.

Hermes Agent blank slate mode: build agents with zero default tools
4 min read

Hermes Agent blank slate mode: build agents…

Nous Research added Blank Slate setup to Hermes Agent. You start with provider, files, and terminal only. Everything else stays off until you opt in, and the config survives hermes update.

Cursor Origin is a git forge built for agent commit storms
6 min read

Cursor Origin is a git forge built…

At Compile, Cursor unveiled Origin: git hosting where agents are first-class users. The demo showed 22.6 commits per second in one repo. Here's what that means before you move your system of record.

Hermes Agent can pay Stripe checkout and HTTP 402 APIs now
5 min read

Hermes Agent can pay Stripe checkout and…

Nous Research shipped three optional Hermes skills wrapping Stripe Link, MPP, and Stripe Projects. Agents can buy on the web, pay per-request APIs, and provision SaaS with human approval gates.

Microsoft FastContext cuts coding agent tokens by offloading repo search
5 min read

Microsoft FastContext cuts coding agent tokens by…

FastContext is a 4B–30B exploration subagent that returns file-line citations instead of dumping whole files into the main agent. Mini-SWE-Agent gains up to 5.5% success with up to 60% fewer main-agent tokens.

What Claude Fable 5's leaked system prompt actually reveals about Mythos
5 min read

What Claude Fable 5's leaked system prompt…

A near-complete Claude Fable 5 product prompt surfaced on GitHub in June 2026. The Mythos tier, artifact storage API, and model-switch rules are the parts that matter for builders.

From SDLC to ADLC: how I orchestrate agents without drowning in PRs
6 min read

From SDLC to ADLC: how I orchestrate…

Agents write at machine speed. Humans still own merge. An Agentic Development Lifecycle playbook: guardrails, test agents, review tiers, and what to discard before it hits your queue.

AI agents write 741% more code. Releases rise 20%. Here's the data.
6 min read

AI agents write 741% more code. Releases…

MIT and Wharton tracked 100,000+ GitHub developers through the full pipeline. Code volume explodes. Shipping barely moves. What the attenuation effect means if you run agents today.

Frontier models are too expensive for agent loops. Specialized models are closing the gap.
7 min read

Frontier models are too expensive for agent…

Composer 2.5 scores 62 on the Coding Agent Index at $0.07 per task while Opus 4.7 costs $4.10. Here's the hybrid routing math I use when agent loops would bankrupt a frontier-only stack.

DeepSeek wrote a VPN obfuscation plugin in one afternoon. The interesting part is not the VPN.
5 min read

DeepSeek wrote a VPN obfuscation plugin in…

A free DeepSeek v4 flash session produced a 410-line SIP003 HTTP/2 obfuscator for Shadowsocks with zero hand-written Go. The story is what coding agents already know about protocol plugins, not circumvention hype.