Rohin Shah - Berkeley
webPersonal homepage of Rohin Shah, who leads the AGI Safety & Alignment team at Google DeepMind and previously authored the influential Alignment Newsletter, making him a key figure in the AI safety research community.
Metadata
Importance: 55/100homepage
Summary
Rohin Shah's personal homepage summarizes his role leading the AGI Safety & Alignment team at Google DeepMind and his prior PhD work at UC Berkeley's CHAI on value learning. His research spans amplified oversight, debate, interpretability via sparse autoencoders, deployment monitoring, and dangerous capability evaluations. He also previously authored the widely-read Alignment Newsletter.
Key Points
- •Leads the AGI Safety & Alignment team at Google DeepMind, focusing on research and policy implementation for powerful AI systems.
- •PhD research at UC Berkeley's CHAI focused on value learning: building AI that can infer and assist with human goals.
- •Research interests include empirical debate, sparse autoencoders for interpretability, deployment monitoring, and dangerous capability evaluations.
- •Authored the Alignment Newsletter, a major resource for tracking AI safety research (now on indefinite hiatus).
- •Maintains a public FAQ on careers in AI alignment, serving as a resource for aspiring researchers.
Cached Content Preview
HTTP 200Fetched Sep 14, 20262 KB
Hi, I’m Rohin! I lead the AGI Safety & Alignment team at Google DeepMind , where we prepare for the development of powerful AI systems, through both research and policy implementation . I completed my PhD at the Center for Human-Compatible AI at UC Berkeley, where I worked on building AI systems that can learn to assist a human user, even if they don’t initially know what the user wants. I used to write up paper summaries in the Alignment Newsletter , though the newsletter is unfortunately on indefinite hiatus now. In my free time, I enjoy puzzles, board games, and karaoke. You can email me at rohinmshah@gmail.com, though if you want to ask me about careers in AI alignment, you should read my FAQ first. Research → Alignment Newsletter → Research My research focuses on AI safety: techniques that ensure that AI systems do what their developers intend. Amplified oversight leverages AI capabilities to evaluate AI outputs. I’m particularly excited about empirical work on debate . Since we build AI systems through machine learning, we don’t understand how they work internally. Interpretability research such as sparse autoencoders aims to bridge this gap. Monitoring AI systems after they are deployed broadly can defend against cases where AI systems appear safe during testing but cause problems “in the wild”. Dangerous capability evaluations like these can provide an early warning for risks, allowing us to put appropriate mitigations in place. Papers → Latest Articles January 3, 2021 FAQ: Advice for AI alignment researchers January 8, 2017 Teaching from Simple Abstractions December 20, 2016 Thoughts on the “Meta Trap”
Resource ID:
3fe7bb7b357ec56d | Stable ID: sid_EqHoxVsuUQ