How do we know the work is good?
Your team writes task-specific rubrics and accepted examples; Evals scores runs against them, surfaces failures, and compares changes. Humans stay in the loop for high-consequence decisions.
Teams write task-specific rubrics and assemble accepted examples. Evals can score runs against those criteria, surface failures for review, and compare changes to a runbook, model, or retrieval configuration. Human review remains part of the process for subjective or high-consequence decisions.
Written for it and security, founders and small teams, developers. Last reviewed 2026-09-15.
Related questions
Still need an answer?
Tell us what you were looking for and we reply within one business day.
Contact the team →Running a security review?
Documents, controls, and the request form live in the Trust Center.
Open the Trust Center →