AGI Index
Updated weekly, or upon major news · last updated August 25, 2026
45.1
/ 100
+3.7 since last update
Current band
Supervised agents
Multi-step work executed end to end, reliably enough to delegate under review.
0
Narrow tools
20
Broad assistants
40
Supervised agents
60
Autonomous colleagues
80
General intelligence
Solid needle is the composite you are currently looking at. The two dashed marks are where the index stood at the previous update (41.4) and where our published weighting puts it (45.1).
The model
Scores are our reading of the published evidence. Weights are judgement about what general means. Drag either and the composite recomputes — here, in the meter above, and in the band you land in.
45.1
Published index · Supervised agents
Move any slider to build your own
Language & reasoning
9.4 of 45.1 pts
78
Wt
12.0%
Multi-step deduction, ambiguity handling, and holding an argument together over long context.
Software engineering
6.8 of 45.1 pts
68
Wt
10.0%
Resolving real issues in real repositories, including the parts that are not writing code.
Mathematics & formal proof
5.8 of 45.1 pts
72
Wt
8.0%
Competition mathematics, proof construction, and machine-checkable formal reasoning.
Long-horizon autonomy
5.1 of 45.1 pts
34
Wt
15.0%
How long a system can pursue an objective before a human has to step in.
Multimodal & world modelling
5.0 of 45.1 pts
55
Wt
9.0%
Perceiving and predicting a physical, consistent world across video, audio, and space.
Novel scientific discovery
3.1 of 45.1 pts
31
Wt
10.0%
Producing findings that are new to the world, not just new to the model.
Reliability & calibration
3.0 of 45.1 pts
38
Wt
8.0%
Knowing what it does not know, failing loudly, and behaving the same way twice.
Continual learning & memory
2.9 of 45.1 pts
22
Wt
13.0%
Getting durably better from experience after training, the way a new hire does.
Physical embodiment
2.1 of 45.1 pts
26
Wt
8.0%
Acting competently in the physical world, including manipulation and unstructured environments.
Economic substitution
2.0 of 45.1 pts
29
Wt
7.0%
The share of real, paid work actually performed end to end by machines.
Rows are ordered by how many points each dimension contributes to the running total, so they reorder as you reweight. The tick on each score bar is the previous update’s value. Weights are normalised, so they do not have to add up to 100%.
Bands
Each band describes a different division of labour between a person and the system. The marker follows whatever composite you have built.
0–20
Narrow tools
Superhuman inside one task, useless one inch outside it.
The human’s role
A person picks the tool and interprets the output. Nothing transfers between tasks.
20–40
Broad assistants
One model covers many tasks at draft quality.
The human’s role
A person reads every output before it is used. The model drafts; the human decides.
40–60
Supervised agents
Multi-step work executed end to end, reliably enough to delegate under review.
The human’s role
A person defines a bounded task and checks the result. The model executes the steps in between.
60–80
Autonomous colleagues
Week-long objectives pursued without step-by-step direction.
The human’s role
A person sets an objective and reviews progress. The model plans, notices its own mistakes, and corrects.
80–100
General intelligence
Matches a capable professional across essentially any cognitive task, including ones nobody trained for.
The human’s role
The division of labour stops being a useful question. Capability is no longer the limiting factor.
Evidence
Every published result behind a dimension score, including the ones that pushed the index down. Filter by dimension to see what a single score rests on.
Effect
Dimension
Showing 12 of 12
Aug 2025
▲
Moved it up
Google DeepMind
Real-time interactive environments generated from a prompt, holding physical and visual consistency over minutes. World models are the missing training ground for embodied agents, and they got materially better.
Jul 2025
▲
Moved it up
Google DeepMind / OpenAI
Two labs independently reached gold-medal scores on the same contest, working in natural language within the official time limit. Formal, closed-domain reasoning is close to solved; the open-ended version is not.
May 2025
▲
Moved it up
Anthropic
Sustained autonomous engineering sessions moved from demo to shipped capability, with customers reporting hours of unattended work. This is the dimension where the gap between benchmark and billable output is narrowest.
Apr 2025
◆
Reframed it
Stanford HAI
The best single source for the base rates behind this page: benchmark movement, cost-per-token collapse, the narrowing gap between open and closed models, and the fact that responsible-AI evaluation is not keeping pace with capability.
Mar 2025
◆
Reframed it
ARC Prize Foundation
The counterweight to the entry above. A successor benchmark, still easy for humans, dropped frontier scores back toward the floor — a reminder that saturating a test is not the same as acquiring the capability it was built to measure.
Mar 2025
▲
Moved it up
METR
The single most load-bearing measurement on this page. The length of task a model finishes at a 50% success rate has been doubling roughly every seven months — which is the strongest reason to think the autonomy score keeps climbing, and the clearest evidence of how low it still is.
Feb 2025
▼
Held it back
Anthropic
Measured usage, not projected usage. AI is concentrated in a narrow slice of occupations and skews toward augmenting a person rather than completing the task alone — which is why economic substitution stays the lowest-scoring dimension we track.
Jan 2025
◆
Reframed it
CAIS & Scale AI
Commissioned because the previous generation of academic benchmarks had been saturated. Frontier scores started in single digits and have climbed steeply since — worth watching precisely because it was designed to be the last hard exam.
Dec 2024
▲
Moved it up
ARC Prize Foundation
A benchmark built explicitly to resist memorisation went from near-zero to human-comparable in a single model generation, at very high inference cost. Test-time compute became a real axis of capability, not a research curiosity.
Nov 2024
◆
Reframed it
Epoch AI
Problems written by working mathematicians and never published, so they cannot be in the training data. Progress here is the cleanest available proxy for genuine mathematical reasoning as opposed to recall.
Oct 2024
▲
Moved it up
Physical Intelligence
A single generalist policy driving multiple robot platforms across dextrous tasks. Robotics is now on the foundation-model curve — a few years behind language, and bottlenecked on data rather than on architecture.
Oct 2024
▲
Moved it up
Anthropic
Opened up every task that has a GUI but no API. Early accuracy on realistic desktop tasks was far below human, which is exactly why reliability — not capability — is the dimension gating agent deployment.
Links go to the original publisher. Tags are clickable — they filter this list to the other results bearing on the same dimension.
Method
Ten dimensions, one anchor, weights that are stated rather than implied. The table below tracks whatever weighting you have set above.
One anchor, ten dimensions
Every dimension is scored 0-100 against the same question: could this replace a competent human professional at this class of work, unsupervised? The composite is the weighted mean.
Weighted toward the bottlenecks
Autonomy, continual learning and discovery carry the heaviest weights, because they are what separates a very good assistant from a general intelligence. Load the benchmark-weighted preset above to see the same scores read about nine points higher.
Anchored to published results
Every score cites a published result, and every result that moved a score is listed below with a link. Vendor claims without a reproducible artefact are not scored.
Benchmarks are not capabilities
Saturating a test measures the test as much as the model. Where a benchmark has been beaten and a successor immediately reset the score, we treat that as evidence about measurement, not progress.
Published weights
What each dimension contributes
Long-horizon autonomy
15.0%
Continual learning & memory
13.0%
Language & reasoning
12.0%
Software engineering
10.0%
Novel scientific discovery
10.0%
Multimodal & world modelling
9.0%
Mathematics & formal proof
8.0%
Physical embodiment
8.0%
Reliability & calibration
8.0%
Economic substitution
7.0%
Total
100.0%
Caveats
This is an editorial index: a reading of public evidence at a point in time, not a benchmark, a forecast, or investment advice. Definitions of AGI differ enough that any single number is a summary of an argument rather than a measurement — which is why the scores and weights here are adjustable rather than fixed. If your version of the number differs from ours, that is the page working correctly.