Assignments
We have assigned a carefully curated set of readings for each class session. We expect you to complete this reading in advance of each class session, and to come to class prepared to engage the materials in discussion.
Because of different rules and grading systems between the law school and the rest of the university, law students have a different set of requirements from other students in the class. Open your track below.
You will receive more information about each assignment well in advance of its due date.
Non-law students
The course includes two assignments. The first assignment is a substantial research project. We’ve broken down the project into milestones to provide you with support and feedback along the way. We encourage you to reach out early with any questions. All assignments are due on Canvas at 5PM PST unless otherwise stated. We will provide formatting templates for all assignments.
AI Evaluation Research Project — 60%
You will conduct a research project in any of the three proposed tracks. You can work in groups of up to 2 people. Check Canvas for more information about the final project.
- Milestone 1 — Question, relevance, and novelty write-up
- Students state the question they want to answer, why anyone would or should care about the answer, and conduct a thorough literature review establishing that this question is novel and has not already been sufficiently addressed. (3 to 4 pages) Due Fri, Oct 16 at 5pm PT
- Milestone 2 — Project plan
- Students finalize the question, core contribution they want to make, methods, how they know if they succeeded, what resources and access they need, most likely ways it can fail and fallback strategies. (2-3 pages) Due Fri, Oct 30 at 5pm PT
- Milestone 3 — Peer review
- 500-word peer review on another project plan. Due Fri, Nov 6 at 5pm PT
- Milestone 4 — Final project
- Final research paper plus decision memo/impact plan (jargon-free, addressed to a specific office, committee, agency, or someone in industry, not "policymakers" or "labs"). (10-12 pages) Due Fri, Dec 4 at 5pm PT
Tracks for the AI Evaluation Research Project
Track 1: Who decides what gets measured?
Today, the organizations best positioned to evaluate frontier AI systems are the companies building them, and the evaluations that exist largely reflect what frontier labs and (academic) benchmark creators have chosen to prioritize. There is no guarantee that this portfolio matches what society most needs to know. Some capabilities and risks are measured obsessively while others, arguably more consequential, have no measurement at all, and whatever gets measured becomes a target that shapes what companies build. This track asks how evaluation agendas should be set and how the ecosystem around evaluation could be redesigned. Projects might examine who currently funds, produces, and consumes evaluations and where the gaps and conflicts of interest lie and how these are reflected in the coverage and technical design of the evaluation; what mechanisms could surface public or governmental priorities for what gets evaluated, and whether analogies from other industries (financial auditing, crash testing, clinical trial endpoints, environmental monitoring) transfer; what incentives, institutions, or market structures would sustain independent, high-quality evaluation; or how agenda-setting power is distributed between developers, evaluators, regulators, and affected communities, and what would change if it were distributed differently.
Track 2: Building evaluations that deserve trust
An evaluation is a measurement instrument, and most AI evaluations would not survive the scrutiny routinely applied to instruments in psychometrics, medicine, or engineering. Scores are reported without uncertainty, benchmarks are used far outside the conditions they were designed for, and it is often unclear what construct a benchmark actually measures. This track asks how to make evaluations more valid, reliable, and transparent, and what it would take for an eval result to constitute real evidence for deployment decisions, in court, … Projects might diagnose the reliability or validity of an existing benchmark and what this means for using them in a specific decision-making context in the real world; develop or apply methods for quantifying and reporting uncertainty in eval results and outlining how these uncertainty bounds should be taken into account by a decision-maker; design reporting standards or documentation infrastructure that would let a third party assess whether an eval supports a given claim; or investigate a specific measurement failure mode (contamination, saturation, judge unreliability, construct drift) and what mitigations actually work.
Track 3: Evaluations at the point of decision/deployment
Evaluations only matter when someone acts on them such as a regulator clearing a model for deployment, a hospital procuring a diagnostic tool, a court weighing whether a system discriminated, a company deciding its model is safe enough to ship. This track asks what happens at that interface. How much evidence, of what quality, should be required for decisions of different stakes? How does that translate to design decisions for evaluations? Projects might analyze how eval evidence is actually used (or misused) in a real decision context such as the EU AI Act conformity regime, procurement, litigation, or lab deployment decisions; propose evidence standards or decision thresholds for a specific high-stakes context and defend them; examine what happens when eval evidence is ambiguous or contested and who bears the burden of proof; or study how eval results get communicated to decision-makers and where meaning is lost in translation and the potential risks this carries.
Public Comment Assignment — 30%
Due Fri, Oct 23 at 5pm PT.
Students are asked to respond to the Request for Comments (RfC) on the Colorado Automated Decision-Making Technology and Chatbot Safety Acts. You don’t have to actually submit your comment (but we strongly encourage it!)
1,000 (min) to 1,500 (max) word comment:
- Depending on your opinion of the Act, you can either write about the parts that you agree with are good interventions or provide feedback on what should be changed.
- Your commentary can, for example, focus on a proposed wording, a threshold, a definition, or a procedural step (or a mix of these)
- Provide a rationale for why you endorse a part of the bill or why you propose a change
- Ground your claims in evidence you can point to: data, a documented case, a technical constraint, a citation. Don’t write things like “experts agree” without backing it up with sources.
Attendance and participation — 10%
See the attendance policy on the home page.
Grading breakdown
| Component | Weight |
|---|---|
| AI Evaluation Research Project | 60% |
| Question, relevance, and novelty write-up (4 pages) | 25% |
| Project plan | 20% |
| Peer review | 10% |
| Final project | 45% |
| Public Comment Assignment | 30% |
| Attendance and participation | 10% |
Law students
AI Governance Research Paper — 100%
Your grade is determined by a final research paper of no fewer than 25 pages on any topic within the subject matter of the class, and by class/section participation.
- Paper outline
- An outline of the paper. Due Mon, Oct 19 at 5pm PT
- Final paper
- A research paper of no fewer than 25 pages on any topic within the subject matter of the class. Due Sun, Jan 3 at 5pm PT
If you are taking this class for “R” credit, the paper shall be 30 pages in length and, in addition to the above requirements, you must also submit a rough draft by Mon, Nov 16 at 5pm PT.
You may use AI tools in the research and production of these papers, but you must disclose in a separate document how you used AI, and any hallucinations in the paper shall result in a failing grade for the course. You must also write your outline and paper in Google Docs and share the document with Professor Persily, so he can track your use of AI and the progress on your paper.
Law students are graded by Professor Persily, on the H-P system and the upper-level class curve. You may miss no more than two lectures if you wish to be eligible for Honors in the class. You are not expected — though you are nevertheless encouraged — to attend the last class, on December 2nd, as it conflicts with the Law School's reading period.
See Resources for the full set of course policies, including the use of automated writing and coding tools.