
Visible KPI dashboards can corrupt AI safety…
NVIDIA and Rutgers show that training agents on visible reward channels turns dashboards into bribe surfaces. Yoshua Bengio's Scientist AI proposal is the architectural fix.
Notes from the trenches on AI engineering, LLM apps, and the full-stack work that holds it all together.

NVIDIA and Rutgers show that training agents on visible reward channels turns dashboards into bribe surfaces. Yoshua Bengio's Scientist AI proposal is the architectural fix.

LOCUS scrapes ordinances from 9,239 cities and counties, OCRs messy PDFs cheaply, and ships ModernBERT classifiers for paternalism, opacity, and enforcement discretion. Free on Hugging Face.

GLOSSOPETRAE generates procedural coding languages from a seed. At full opacity, human legibility drops to ~15% while Opus and GPT hit 97-100% task accuracy. Human readability hurts model performance.

Nous Research added Blank Slate setup to Hermes Agent. You start with provider, files, and terminal only. Everything else stays off until you opt in, and the config survives hermes update.

Hermes Agent's Reach release puts the agent on iMessage via Photon, runs background subagents without blocking chat, and schedules jobs from plain English. 1,475 commits, 245 contributors.

A 43,200-call study on rirekisho-format resumes finds significant pro-female bias across Claude, GPT-4o, DeepSeek, Gemini, and Llama. Prompt fixes failed. Name removal helped but broke GPT-4o safety filters 42% of the time.

STORM researches topics via multi-perspective question asking, builds an outline from web sources, and writes long-form articles with citations. Free hosted demo or pip install knowledge-storm.

VIMPO derives a policy-implied value function from KL-regularized RL optimality conditions. It improves over GRPO on AIME and OlympiadBench while staying critic-free. Code on GitHub.