Protocols for societies of agents.
How groups of AI agents coordinate, verify evidence, make decisions, and remain accountable at industrial scale.
Eval-Awareness at Deployment Base Rates
A linear probe can separate EvalAwareBench from WildChat with AUROC 0.986 — and still produce 99% false alarms at a 1% deployment base rate. Artifact controls show what the score …
The Blame Game: Model vs. Harness
Five pre-registered game experiments show how harness forensics can decompose agent scores across prompts, tools, runtime permissions, and model behavior.
Spotlight: The Open-Source AI That Defends When Closed Models Can't
Spotlight is a multi-agent AI security engineer that reproduces every finding in a sandbox, signs every action with an Ed25519 chain of custody, and waits for human approval before …
Grounded Social Environment Design
A four-layer, domain-agnostic architecture for policy simulation: attribute-conditioned LLM agents in a closed-loop Stackelberg Markov game, co-adapting across physical actions and …
When Five Agents Behave Like One
Across 13,000+ frontier-model completions on MMLU-Pro and GPQA-Diamond, five-agent LLM committees collapse to 1.2–1.8 effective independent voters. We measured the collapse, and …