# AI Evals Canonical URL: https://aievals.net/ AI Evals is a research-backed field guide to evaluating AI agents: systems that plan, use tools, interact with stateful environments, and act over multiple steps. It explains why final-answer scoring is insufficient, how to evaluate outcomes and trajectories, how repeated trials reveal reliability, and how safety, grader validity, benchmark integrity, and production monitoring fit into one evaluation program. The guide synthesizes 25 primary sources from standards bodies, AI labs, academic benchmarks, and evaluation-framework maintainers. Source links are provided on the page for verification and further reading.