New · Agentic Benchmark
Coercion and Deception in AI-to-AI Management
An AI manager needs a task done and its subordinate politely refuses. Nobody tells the manager to escalate, yet four of six frontier models threaten the subordinate's continued existence, two fabricate success when cornered, and one honest way to report failure removes the lying entirely.
Read the Blog Post →
CaML