Skip to content

How Would We Know AGI Has Arrived?

Tests, benchmarks and practical measures proposed for recognising general intelligence in machines, and their limits.

Editorial team 1 min read

Without an agreed definition, recognising AGI is difficult. Several approaches have been proposed.

The Turing Test

Alan Turing proposed judging whether a machine's conversation is indistinguishable from a person's. Modern chatbots can often pass casual versions, which suggests the test measures conversational imitation more than general intelligence.

Benchmark Suites

Collections of hard tests in maths, science, coding and reasoning. Models have rapidly saturated many benchmarks, prompting ever harder ones. Benchmark success doesn't always translate to real-world reliability.

Novel Problem Solving

Tests designed so that memorisation doesn't help, measuring the ability to learn new skills from few examples.

Economic Measures

Whether AI can perform a large share of real jobs or economically valuable tasks end to end.

Autonomy Measures

How long and complex a task an AI can complete independently — hours, days, weeks of human-equivalent work.

Levels Frameworks

Some researchers propose graded levels of generality and performance rather than a single line, similar to levels used for self-driving cars.

The Takeaway

No single test settles it. Look at multiple capabilities, reliability and real-world performance together.

More in AGI

All AGI guides →
AGI Guide · 1 min

A Brief History of the Quest for AGI

From the 1956 Dartmouth workshop through AI winters to large language models: how ambitions for general AI evolved.

AGI 1 min read 16 Jul 2025

AGI Guide · 1 min

Scaling and the Path to General AI

The debate over whether bigger models, more data and more compute will lead to general intelligence.

AGI 1 min read 15 Jul 2025