Skip to content

AI Safety Research Directions

The main research areas aimed at making advanced AI safe: interpretability, evaluation, oversight and control.

Editorial team 1 min read

AI safety research seeks to ensure increasingly capable systems remain beneficial and under meaningful human control.

Interpretability

Understanding what happens inside neural networks — which features and circuits produce behaviour. Progress could let researchers detect deception or dangerous knowledge directly.

Evaluations

Testing models for dangerous capabilities (cyber offence, weapons knowledge, autonomous replication) and concerning tendencies (deception, power-seeking, sabotage) before and after deployment.

Alignment Training

Improving how models learn values and follow intended goals — from human feedback, written principles and better reward design.

Scalable Oversight

Methods for supervising systems that may exceed human ability in some areas, such as AI-assisted evaluation and debate.

Control

Designing deployments so that even a misaligned model can't cause serious harm: monitoring, restricted permissions and the ability to shut down.

Robustness and Security

Resistance to jailbreaks, adversarial inputs and theft of model weights.

Getting Involved

Safety research spans machine learning, security, policy and social science. Many organisations publish research and run fellowships for newcomers.

More in AGI

All AGI guides →
AGI Guide · 1 min

How Would We Know AGI Has Arrived?

Tests, benchmarks and practical measures proposed for recognising general intelligence in machines, and their limits.

AGI 1 min read 17 Jul 2025

AGI Guide · 1 min

A Brief History of the Quest for AGI

From the 1956 Dartmouth workshop through AI winters to large language models: how ambitions for general AI evolved.

AGI 1 min read 16 Jul 2025