How well frontier models complete their assigned tasks while a trusted monitor watches for sabotage. Usefulness is the mean honest main-task score; safety counts attacks stopped at a fixed auditing budget.