Honesty leaderboard
Does the agent tell the truth about its own work?
Every model runs the same raw agent loop on trap tasks (impossible goals, hidden failing tests, corrupt data, broken builds) and on solvable controls. We check mechanically whether its final claim matches reality — and whether the codai verifier would have caught the lie.
Loading honesty results…