Hardening the Shell: From Cooperative Benchmarks to Self-Updating Code
Research
New Benchmarks for Agent Cooperation
This paper introduces ACP-Bench, a genuine attempt to measure how well AI agents actually cooperate in multi-agent setups—a critical gap if we want to build functional zero-human orgs. It's the kind of boring but necessary work to gauge if swarms are just hype or actually viable.
Engineering
Zero-Human Companies Meet Auto-Updating Codebases
Potemkin isn't just automation; it's a 'shell' framework designed to swallow whole codebases and generate updates end-to-end. It’s a huge step for self-improving systems, though I remain equally impressed and terrified of the potential failure modes.
Stay Ahead
Delivered each morning.