Fellowship Syllabus
Strategy and Forecasting
As frontier models grow more capable and the cost of frontier-level performance continues to fall, transformative AI is approaching quickly. Navigating it well requires a coherent picture of where AGI is headed, one that integrates technical progress with the political dynamics of an AGI arms race. This reading and discussion group is dedicated to building shared understanding around the crucial questions of transformative AI. Fellows engage in forecasting informed by scaling laws, threat modeling, and a tabletop wargame exercise, culminating in each fellow writing their own AGI takeoff scenario.
Week 01
Benchmarking Progress and Forecasting
How do we track AI capability progress, and how good are we at predicting it? Why is forecasting AI valuable? Fellows open the quarter with calibration exercises on concrete AI milestones and are introduced to the scenario-writing framework that threads through the program.
- Measuring AI Ability to Complete Long Tasks
METR's benchmark tracking the length of tasks frontier AI agents can complete — how we're measuring capabilities.
- What Will AI Look Like in 2030?
Epoch AI on the limitations of scaling and what continued progress looks like.
- How Well Does RL Scale?
Toby Ord examines the scaling behavior of reinforcement learning and what it implies for capability forecasts.
- Forecasting the Economic Effects of AI
The Forecasting Research Institute's survey of expert and superforecaster views on AI's economic impacts.
Week 02
Narrative Forecasts and Timelines
How do we reason about discontinuous or high-stakes futures? What makes a scenario analytically useful rather than merely aesthetically compelling? Fellows dissect existing narrative forecasts, rate specific claims on internal consistency and empirical grounding, and identify the assumptions that seed their own scenario drafts.
- AI 2040
The AI Futures Project's narrative forecast of AI development through 2040, built around a verification plan involving compute governance.
- Broad Timelines
Toby Ord's thoughts on how to reason about AGI timelines.
Week 03
Recursive Self-Improvement and Takeoff
What mechanisms could produce fast takeoff? What are the cruxes between a slow, detectable transition and a fast, catastrophic one? Fellows map their top empirical cruxes about takeoff speed and draft the opening world-state of their own scenarios.
- When AI Builds Itself
Anthropic's argument, using public benchmarks and previously unreported internal data, that AI is already accelerating AI development.
- Where I Agree and Disagree with Eliezer
Paul Christiano's point-by-point response to Eliezer Yudkowsky, mapping the key disagreements about takeoff and doom.
Week 04
Planning For Safety
What's the threat model? What solution classes exist, and what are their failure modes? Fellows map specific failure modes, such as deceptive alignment, onto candidate plans and examine how hard misalignment is to measure in the first place.
- Plans A, B, C, and D for Misalignment Risk
Different levels of government intervention suggest different plans and timelines for AI safety.
- Corrigibility
The foundational MIRI paper on building agents that tolerate correction and shutdown rather than resisting them.
- Responsible Scaling Policy
Anthropic's internal governance policy for scaling frontier models safely.
- A Pragmatic Vision for Interpretability
Neel Nanda on what mechanistic interpretability can realistically contribute to safety.
- Autonomy Evaluation Resources
METR on the hardness of measuring misalignment: safety evals, benchmark saturation, and eval-awareness.
Week 05
Governing Frontier Development
What are the realistic levers for reducing risk at the policy level, and where do they conflict? Fellows stress-test policy recommendations by arguing against them from the perspectives of specific actors, and add the inciting decision point to their scenarios.
- Superintelligence Strategy (MAIM)
Hendrycks, Schmidt, and Wang's proposal for a deterrence regime — Mutual Assured AI Malfunction — modeled loosely on nuclear deterrence.
- Policy on the AI Exponential
Dario Amodei argues that exponential AI progress has outpaced the policy process, and proposes concrete responses.
- AI Scenarios 2030
The UK government's scenario-planning exercise helping policymakers prepare for the future of AI.
- Computing Power and the Governance of Artificial Intelligence
GovAI argues compute is uniquely governable among AI inputs and surveys how governments can use it to monitor and steer AI development.
Week 06
Envisioning Positive Worlds
What are we working toward, not just against? Fellows present their own scenarios — or structured critiques of existing ones — with feedback on the key claim, the best objection, and what evidence would change their view.
- Introducing Better Futures
Forethought's essay series on making the future go well, not merely avoiding catastrophe.
- How to Make the Future Better
Concrete actions to improve the long-run future conditional on surviving the transition to advanced AI.
Week 07
Tabletop Exercise
A longer capstone session. The tabletop exercise simulates the development and geopolitical implications of advanced AI from late 2027 through 2028, and participants assume roles such as the President of the United States, foreign nations, or employees at a frontier AI lab.
- AI 2027
The narrative scenario the exercise is based on.
