AGI Safety and Alignment: Technical Approaches
AI Safety pro 2025-11-03 7 min read

The alignment problem in plain terms

An AGI system optimizing for a goal it was given imperfectly can cause serious harm — not from malice, but from precision. That is the alignment problem. This program treats it as an engineering challenge, not a science fiction scenario.

We cover reward modeling, constitutional AI, debate-based alignment, and interpretability tooling. You will run experiments with smaller-scale models to observe misalignment failure modes directly, which makes the theory considerably easier to retain.

Practical lab component

Four of the eight sessions include hands-on labs using open-source frameworks. You will instrument a reward model, observe reward hacking in a controlled environment, and test an interpretability probe on a transformer layer. These are not toy exercises — they reflect methods used in current alignment research.

Realistic expectations

Alignment research is genuinely hard and unsolved. This program does not promise you will leave with answers. It gives you the vocabulary, tooling, and mental models to contribute meaningfully to the conversation — whether in a research role, a policy context, or as an engineer building safety-critical AI systems.

AGI Safety and Alignment: Technical Approaches

About this material

The alignment problem in plain terms

An AGI system optimizing for a goal it was given imperfectly can cause serious harm — not from malice, but from precision. That is the alignment problem. This program treats it as an engineering challenge, not a science fiction scenario.

We cover reward modeling, constitutional AI, debate-based alignment, and interpretability tooling. You will run experiments with smaller-scale models to observe misalignment failure modes directly, which makes the theory considerably easier to retain.

Practical lab component

Four of the eight sessions include hands-on labs using open-source frameworks. You will instrument a reward model, observe reward hacking in a controlled environment, and test an interpretability probe on a transformer layer. These are not toy exercises — they reflect methods used in current alignment research.

Realistic expectations

Alignment research is genuinely hard and unsolved. This program does not promise you will leave with answers. It gives you the vocabulary, tooling, and mental models to contribute meaningfully to the conversation — whether in a research role, a policy context, or as an engineer building safety-critical AI systems.

Program structure

What gets covered and in what order — no filler, no repetition.

Program Outline

Module 1
Failure modes of reward specification — Goodhart's Law in practice
Module 2
Reinforcement learning from human feedback (RLHF): mechanics and limitations
Module 3
Constitutional AI and rule-based constraint systems
Module 4
Lab: reward hacking demonstration with a gridworld agent
Module 5
Interpretability methods: probing, activation patching, causal tracing
Module 6
Lab: building a simple linear probe on a language model
Module 7
Scalable oversight and debate as alignment strategies
Module 8
Capstone: written alignment proposal for a specified AGI use case

Ready to start?

Seats fill up quickly — once the cohort closes, the next opening is months away.

9 seats remaining