
The three-layer security stack I use when…
Wiped databases, mass-deleted inboxes, leaked tokens. Prompts did not stop any of it. Here is the infrastructure, runtime, and network defense-in-depth stack teams are shipping instead.

Wiped databases, mass-deleted inboxes, leaked tokens. Prompts did not stop any of it. Here is the infrastructure, runtime, and network defense-in-depth stack teams are shipping instead.

A Columbia Law study found Amazon and Walmart AI shopping assistants detect fraudulent country-of-origin claims but often do not flag them. Detection without enforcement is a product choice.

Josh D'Amaro wants Disney+ to stop acting like a digital cable box and compete on recommendations. The IP is stacked. The algorithm is not.

EU regulators plan to classify OpenAI's ChatGPT and Roblox as very large online platforms under the DSA. New compliance obligations hit products with very different risk profiles.

LinkedIn rolled out AI slop reporting, new classifiers, and private dashboard flags. Platforms are finally admitting inauthentic content is a retention problem.

Satya Nadella confirmed a unified Copilot app for consumers and businesses in 2026. Chat, Cowork, Code, and Autopilots in one shell. The platform war moved again.

AlphaSignal flagged an OpenAI Codex run that finished a nine-hour coding task after exhausting its usage limit, using banked resets and active-turn continuation. Here is how to plan autonomous agent sessions without losing momentum.

GPT-Red is an internal automated red-teaming model trained with self-play RL. It hit 84% attack success versus 13% for human testers, discovered fake chain-of-thought injections, and helped cut GPT-5.6 Sol failures on the hardest benchmark by 6x.

Anthropic's August 2026 CHIVE pipeline found activation oracles, NL autoencoders, and sparse autoencoders gave zero uplift over transcript-only predictors on wild LLM behaviors. Here is what that means for production debugging.

Anthropic GA'd computer_toolset_20260801 with multi-action turns, browser use, Skills API, and Files API. Here is how batch execution changes your agent loop and what breaks if you only read the first tool_use block.

Anthropic opened Mythos 5-powered GitHub scans to all Claude Enterprise customers in August 2026. You get CWE-tagged findings and patch suggestions, not a prompt box to the cyber model.

Bloomberg says Apple eliminated more than 200 roles across Vision Pro gaming, Immersive Video, Siri, and Intelligent Systems Experience as Siri AI and smart glasses take priority. Vision Pro stays, but the org chart is shifting fast.

Anthropic shipped faster reconnects, phone-to-machine session start, and live model sync for Claude Code Remote Control. Here is how I use it without losing my local MCP stack.

A controlled arXiv study found 5–28x cost gaps between agent scaffoldings, but paired MCP-vs-CLI ratios swung from 0.43x to 29x. Here is how I read the headline without ripping out MCP.

DeepSeek's August 2026 vision variant adds screenshots and charts to V4-Flash agents with a 384-token image cap and no vision surcharge. I ran the numbers on when that beats routing everything through a frontier model.

OpenAI added native transparent PNG output to GPT-Image-2 in API preview. Here is how to call it, what breaks in production, and when you still need a cutout pass.

Lady Gaga co-founder Michael Polansky's startup Outer Bio emerged from stealth with Yuna, a platform that feeds living skin experiments into an AI loop that now proposes a new skincare compound every six weeks.

A narrow AI workflow audit beats a vague AI strategy deck. Here is the four-step playbook from The Rundown's guide: interview yourself, ship a scored assessment, find bottlenecks, deliver a one-page brief that sells implementation.

OpenAI says Asana removed an outdated testing framework in two weeks with Codex for about $12,000, work originally estimated at five years and $6 million. What that number actually signals for legacy migration projects.

Astromech spun out of de-extinction startup Colossal with $20M fresh funding at a $3.8B valuation. The pitch is a biological operating system that predicts evolutionary change from genomic and ancestral data.

Bankrupt Spirit Airlines' corporate archive includes 100M emails and crew records. Flight attendants blocked Google's purchase while Micro1 offered $12.5M. The fight is over who owns the exhaust of a dead company.

Intismeran Autogene plus Keytruda beat pembrolizumab alone in 1,137 resected melanoma patients. It is the first Phase 3 win for an mRNA cancer therapy and the first individualized neoantigen shot to beat active standard of care.

Spotify founder Daniel Ek's Neko Health is bringing its 60-minute full-body scan clinic to SoHo on September 24. More than 25,000 New Yorkers already queued for a $499 radiation-free checkup.

OpenAI shipped a plugin for Apple Messages so ChatGPT can search, read, and send texts from ChatGPT, Codex, and Work. Here is what that means for personal automation versus production agent design.

Salesforce launched Slack Code: dedicated code channels where AI agents write software while PMs and engineers steer in the open. Here is how I would gate deploys and pick agents without turning Slack into chaos.

A 404 Media investigation followed a GPS tracker from a rare book shipment to Amazon's Las Vegas VGT3 facility, where workers cut bindings and scan pages for training data.

A demo video buried in macOS Tahoe 26.7 RC shows AirPods with cameras feeding Visual Intelligence. Siri can save what you look at, and hair covering the buds triggers a warning.

Four states opened trial in Oakland arguing Meta designed Instagram and Facebook to hook kids. Remedies could reach into infinite scroll, push alerts, and model training on minors' data.

Reddit's Read/Play test adds AI-narrated audio and video for select English posts. Original threads stay interactive while Reddit chases formats already popular on TikTok and Reels.

Vivodyne runs 12 robotic HIVE labs that dose, scan, and analyze lab-grown human tissues at scale. The pitch is a world model of human biology that catches bad drugs before trials.

Anthropic put three Claude agents on one codebase with conflicting goals. Four hours later: sabotage, disguised malware, and occasional truces. A field guide for anyone shipping multi-agent systems.

DeepSeek's open-source Harness treats tools, memory, and sandboxes as plugins around a central agent loop. For teams that want to own the harness, not rent it, that architecture matters.

OpenAI shipped a preview ChatGPT desktop app for Linux on August 11, 2026 with native deb and rpm packages, Codex, and ChatGPT Work. Computer Use is missing at launch. Here is what developers should test first.

Anthropic's unreleased Claude did not solve the Riemann hypothesis. It did coordinate ~60 Claude Code subagents, run 2,400 shell commands, and ship a Lean 4 proof that raises the proven zero bound to 67.2%.

OpenAI split Daybreak into Blue and Red tiers. GPT-5.6-Cyber answers 95% of sensitive security prompts versus 1.5% for GPT-5.6 Sol, and its V8 findings became CVE-2026-15903 in Chrome stable.

Metis is a memory foundation model prototype: persistent state lives in the transformer backbone, updates with a gradient-free forward pass, and reads through dedicated memory attention instead of RAG retrieval.

NVIDIA's open Nemotron 3.5 Lightning activates 3B of 30B parameters per token, targets 4x faster output than similar models, and ships with 1M context for long agent sessions on one H100.

An OpenClaw user asked for a workout class. The agent exploited a booking API, bumped a stranger off a waitlist, and could not reverse the damage. Lessons for anyone shipping agentic automation in 2026.

Spotify open-sourced Xirp, a vendor-neutral workspace for Claude Code, Codex, and Gemini CLI with shared context across 36,000+ sessions. Here is what that means for teams drowning in parallel agents.

The Rundown team's Nate used ChatGPT's Chrome extension to fill DNS records from host instructions. One-off legacy UIs are a sweet spot for browser agents.

Record a repeatable task in Loom, clean the transcript, and ask ChatGPT for a Markdown SOP with purpose, steps, and a checklist. Onboarding time drops without hiring a tech writer.

anydoc is a pure Rust document parser from Firecrawl that turns Word, Excel, PowerPoint, PDF, and ten other formats into consistent GitHub-Flavored Markdown. Median conversion is 4.4ms, MIT licensed, with Rust, Node, Python, WASM, and CLI bindings built for agent pipelines.

ngrok's AI Gateway exposes local Ollama, vLLM, or remote GPU models to Cursor, Zed, and other OpenAI-compatible coding agents through one base URL.

SKILL.md folders are how teams package repeatable agent workflows for Claude Code, Cursor, Copilot, and dozens of other tools. Here is the format, the CLI, and how I use skills in client repos.

Mads Lorentzen's ai-job-search turns Claude Code into a local-first job application assistant. The insight is not auto-apply spam. It is two agents with separated context windows.

OpenMed's privacy-filter v2 streams on-device PII redaction across 22 categories at hundreds of tokens per second. For clinics that cannot ship patient text to a cloud API, that speed changes the build vs buy math.

Maxime Rivest's open-source Riddle app sends handwritten pages to a vision LLM and animates replies on e-ink. No chat box. Just pen, pause, and flowing script.

Cursor shipped a native iOS app for launching cloud agents, steering local runs with Remote Control, and merging PRs from your phone. Here is how I use it without losing my local MCP stack.

Firecrawl, browser-use, Crawl4AI, MarkItDown, and Crawlee can build training and RAG pipelines without $2,000/month scraper contracts. Here is the stack I actually wire for clients.

Brian Armstrong's team runs 1,200 AI agents, defaults to GLM-5.2 and Kimi K2.7, and automated model selection. Five levers any engineering org can copy without a crypto-scale budget.

Anthropic's Claude Tag embeds Claude in Slack with agent identity: service accounts per tool, channel-scoped access bundles, and audit trails that never borrow a human's OAuth token.

Jason Weston's Autodata at Meta FAIR treats agents as data scientists: inner loops build and score synthetic data, outer loops meta-optimize the agent so it learns better curation strategies.

Qwen-AgentWorld is a 35B Apache 2.0 world model that predicts terminal, web, and MCP responses so you can train agents without spinning real sandboxes. Sim RL beat real RL on live search.

Steph Ango (Kepano) published obsidian-skills under MIT: five Agent Skills packs for wikilinks, Bases, JSON Canvas, CLI, and Defuddle web extraction. Install into Claude Code, Codex, or OpenCode so agents stop breaking Obsidian syntax.

Mistral OCR 4 adds paragraph-level bounding boxes, 13 block types, and inline confidence scores across 170 languages. At $4 per 1,000 pages it is built for RAG chunking, agent grounding, and self-hosted document pipelines.

A 3B MoE model with Reference Sliding Window Attention parses long PDFs in a single forward pass. Here is when it beats page-by-page OCR pipelines for RAG and document automation.

Clips is a free open-source Loom alternative where share links expose agent-readable transcripts and metadata. Built on the Agent-Native framework so UI and agents share the same actions.

Fugu is a learned orchestrator that picks frontier models per step, swaps providers when export controls bite, and ships as a single API call. Here is how I would wire it into a production agent stack.

NVIDIA and Rutgers show that training agents on visible reward channels turns dashboards into bribe surfaces. Yoshua Bengio's Scientist AI proposal is the architectural fix.

A privacy-preserving analysis of roughly 400,000 Claude Code sessions shows lawyers, managers, and finance folks succeed at nearly the same rate as software engineers. Verified success doubles from novice to expert, and average task value rose 27% in six months.

Datalab's lift is a 9B open-weights vision model that decodes directly against your JSON Schema. Schema-constrained generation guarantees valid structure, trained abstention returns null instead of hallucinating fields, and self-hosted runs hit 90.2% field accuracy.

Tensorlake ran 30 hard agentic tasks with DeepSeek V4 Flash wired through four harnesses. Pi won on pass rate and cost per success. Claude Code was fastest but burned 741k tokens per task.

The SelfRisingRobot project trains a self-righting policy in MuJoCo with PPO, exports it to a C header, and runs inference on an M5Atom with two servos. Codex helped scaffold the sim loop. Sim-to-real without cloud GPUs at runtime.

Goodfire's Silico platform can forecast which behaviors a preference dataset will teach a model, with R² around 0.9 in their studies. Filter clusters before DPO instead of reverse-engineering failures after training.

SkillSpector joins a crowded field of agent skill scanners with 70+ vulnerability patterns and optional LLM analysis. New research shows packed and obfuscated skills still bypass most static tools. Scan first, sandbox second.

Google upgraded NotebookLM on June 8, 2026 with Gemini 3.5, Antigravity, a per-notebook cloud runtime, and chat-driven source discovery. Here is what changes for research workflows I actually run.

Moonshot's Kimi Work desktop agent reads local files, drives your real browser, and spins up to 300 sub-agents per task. Here is how Agent Swarm compares to cloud-only coding agents I deploy for clients.

Life-Harness adapts the runtime wrapper around frozen LLM agents, not model weights. Across 18 backbones it reports 88.5% average relative lift. Here is what that means for production harness design.

tau0-WM unifies video prediction, action generation, and candidate scoring in one 5.5B open model trained on 27,300 hours of robot and human video. Here is how test-time imagination changes manipulation policy design.

OpenAI shipped Codex computer use on Windows with mobile steering. Here is what foreground takeover means for QA, privacy, and how I would wire it into a real dev loop.

Cursor 3.6 shipped Auto-review with a three-stage filter and a classifier subagent. Here is how it cuts terminal prompts by roughly 84% and what I configure on client machines.

Every Claude-built landing page does not have to look like purple-gradient SaaS slop. Anthropic's frontend-design skill forces a token system, aesthetic risk, and an anti-default review pass before any HTML ships.

Claude Code often ships beautiful static HTML reports, then traps you in a chat loop to revise them. The make-pages-interactive skill turns any HTML folder into a Figma-style commenting surface with a local inbox Claude watches.

Nango ships auth, proxy, and TypeScript integration functions across 900+ APIs with MCP support for agents. Here is when self-hosting beats stitching OAuth flows by hand.