Lecture 4

AI Evaluations

Monday, October 5, 2026

Technical Foundations

This lecture focuses on the role that evaluations play in AI governance. We explore how model evaluations are used to inform access policies, reporting requirements, and system release decisions. We'll also discuss technical challenges such as the absence of shared standards for benchmark development, concerns about (construct) validity, and data contamination in test sets. We'll further examine recent proposals for evaluation frameworks, safe harbor mechanisms for auditing, and how red-teaming practices can supplement formal evaluation protocols.


All lectures · Schedule

Edit this page.

Licensed under CC BY-NC-SA 4.0.