The timeline

Selected milestones in how people use AI

Selected publications and public milestones, not dates of invention. Select a product for the event, date and evidence.

Foundational publication Product / architecture milestone
Level
20192021202320252027
7Always-on agents
5Agent loops
3Workflows
2Added context
1Direct prompting

0 · Conventional software Rules, search and conventional software have no single arrival date.

Lines connect selected dates, not development time. Sources checked through Sep 20, 2026.

Across all levels · changes in the models

Reasoning models · Sep 2024

Models trained to think before they answer. Not a new level: a better engine for every level, and the one OpenAI credits for its first agent. Source ↗

Typed decision models (Jev) · Sep 2026

TypeSafe introduced Jev in early access on September 15, 2026: a model that returns typed decisions with probabilities rather than generated text. “System One” is TypeSafe’s name for its proposed model category. Source ↗

Selected milestone

7 · Always-on agentsGrok BotPaid beta released

An example of agents that keep working beyond an interactive session. The product also describes a team of agents: one product can illustrate more than one level.

Marked event: Aug 11, 2026 · Introducing Grok Bot ↗

Foundational publication: Generative Agents ↗ · Apr 7, 2023. This is a reference point, not a claim to the first use of the idea.

6 · Teams of AgentsClaude ResearchMulti-agent architecture documented

Anthropic explicitly describes a lead agent delegating to parallel subagents on June 13. Research launched on April 15, but that announcement does not establish whether the same architecture was in use then. June is the documentation date, not an asserted deployment date.

Marked event: Jun 13, 2025 · How we built our multi-agent research system ↗

Foundational publication: CAMEL ↗ · Mar 31, 2023. This is a reference point, not a claim to the first use of the idea.

5 · Agent loopsOperatorBrowser-use agent preview · US Pro

The agent observes a browser, acts, and decides what to do next. This is a browser-use milestone; agents and autonomous loops existed earlier.

Marked event: Jan 23, 2025 · Introducing Operator ↗

Foundational publication: ReAct ↗ · Oct 6, 2022. This is a reference point, not a claim to the first use of the idea.

4 · Tool useChatGPT pluginsLimited alpha announced · waitlist

A visible example of a model choosing a tool. This dates the limited rollout, not broad availability or the first implementation of tool use.

Marked event: Mar 23, 2023 · ChatGPT plugins ↗

Foundational publication: MRKL Systems ↗ · May 1, 2022. This is a reference point, not a claim to the first use of the idea.

Later access: rollout to Plus subscribers begins, May 12, 2023.

3 · WorkflowsOpenAI steps in ZapierArchived availability evidence

A model call inside a workflow whose steps are set in advance. December 9 is an archived-page capture, not a confirmed launch date. Microsoft 365 Copilot remains a related workplace-AI milestone, rather than the defining workflow example.

Marked event: Available by Dec 9, 2022 · OpenAI Integrations ↗

Foundational publication: AI Chains ↗ · Oct 4, 2021. This is a reference point, not a claim to the first use of the idea.

2 · Added contextPerplexitySearch-and-answer product launched

The early Perplexity Ask combined retrieved search results with a generated answer and citations. That is the specific Level 2 pattern illustrated here, not a classification of every later Perplexity feature. Cofounder Aravind Srinivas dates the launch to December 7, 2022 in a retrospective interview; this is founder testimony, not a contemporaneous launch announcement.

Marked event: Dec 7, 2022 · Aravind Srinivas on the launch date · Stripe Sessions 2024 ↗

Early behavior: Founder interview describing the first Perplexity Ask ↗.

Foundational publication: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks ↗ · May 22, 2020. This is a reference point, not a claim to the first use of the idea.

1 · Direct promptingChatGPTResearch preview released

A recognizable public example of prompting a language model. Earlier products existed; this is a selected milestone, not the invention of prompting.

Marked event: Nov 30, 2022 · ChatGPT ↗

Foundational publication: Better Language Models and Their Implications (GPT-2) ↗ · Feb 14, 2019. This is a reference point, not a claim to the first use of the idea.

How fast

What the models could do over the same years

Every model Epoch AI has scored, by release date, with the best score to date drawn over them and this site's level dates underneath.

Epoch Capabilities Index by model release date, 2023 to 09/20/2026266 models scored by Epoch AI, plotted by release date. A stepped line follows the best score to date, from 109.94 to 166.31. Markers under the axis show selected milestones for this site's levels, and a vertical line marks the first reasoning model.4060801001201401601802023202420252026ECI scoreFirst reasoning model · Sep 2024LLaMA-13B, Meta AI · 100.2 · 02/24/2023LLaMA-33B, Meta AI · 107.14 · 02/24/2023LLaMA-65B, Meta AI · 109.94 · 02/24/2023LLaMA-7B, Meta AI · 96.25 · 02/24/2023GPT-4 (Mar 2023), OpenAI · 125.89 · 03/14/2023Cerebras-GPT-13B, Cerebras Systems · 82.57 · 03/20/2023Dolly 2.0-12b, Databricks · 89.11 · 04/11/2023vicuna-13b-v1.1 · 94.19 · 04/12/2023stablelm-tuned-alpha-7b · 54.64 · 04/19/2023Falcon-7B, Technology Innovation Institute · 94.63 · 04/24/2023RedPajama-INCITE-7B-Base · 89.87 · 05/04/2023MPT-7B, MosaicML · 94.11 · 05/05/2023PaLM 2-L · 114.88 · 05/17/2023PaLM 2-M · 108.04 · 05/17/2023PaLM 2-S · 105.88 · 05/17/2023Falcon-40B, Technology Innovation Institute · 104.13 · 05/25/2023Baichuan1-7B, Baichuan · 89.92 · 06/01/2023open_llama_7b · 91.12 · 06/07/2023GPT-3.5 Turbo (Jun 2023), OpenAI · 113.19 · 06/13/2023GPT-4 (Jun 2023), OpenAI · 123.1 · 06/13/2023MPT-30B, MosaicML · 100.28 · 06/22/2023chatglm2-6b · 98.53 · 06/24/2023XGen-7B, Salesforce · 92.9 · 06/27/2023internlm-7b · 102.53 · 07/05/2023Claude 2, Anthropic · 120.09 · 07/11/2023Llama 2-13B, Meta AI · 105.88 · 07/18/2023Llama 2-34B, Meta AI · 104.9 · 07/18/2023Llama 2-70B, Meta AI · 113.63 · 07/18/2023Llama 2-7B, Meta AI · 98.66 · 07/18/2023Stable Beluga 2, Stability AI · 117 · 07/20/2023Claude Instant, Anthropic · 120.2 · 08/09/2023Baichuan2-13B, Baichuan · 102.86 · 09/06/2023Falcon-180B, Technology Innovation Institute · 111.94 · 09/06/2023Phi-1.5, Microsoft · 90.99 · 09/11/2023internlm-20b · 111.89 · 09/18/2023Baichuan 2-7B, Baichuan · 95.88 · 09/20/2023Qwen-14B, Alibaba · 112.85 · 09/24/2023Mistral 7B v0.1, Mistral AI · 112.02 · 09/27/2023Qwen-7B, Alibaba · 106.55 · 09/28/2023DeepSeek Coder 1.3B, DeepSeek,Peking University · 62.54 · 11/02/2023DeepSeek Coder 33B, DeepSeek,Peking University · 95.87 · 11/02/2023DeepSeek Coder 6.7B, DeepSeek,Peking University · 89.05 · 11/02/2023Yi-34B, 01.AI · 117.3 · 11/02/2023GPT-3.5 Turbo (Nov 2023), OpenAI · 118.49 · 11/06/2023Claude 2.1, Anthropic · 119.22 · 11/21/2023Yi 6B, 01.AI · 104.46 · 11/22/2023DeepSeek LLM 67B, DeepSeek · 110.56 · 11/29/2023Qwen-1_8B · 92.43 · 11/30/2023Mixtral 8x7B, Mistral AI · 118.4 · 12/11/2023Phi-2, Microsoft · 107.69 · 12/12/2023Gemini 1.0 Pro, Google DeepMind · 116.97 · 12/13/2023GPT-3.5 Turbo (Jan 2024), OpenAI · 115.61 · 01/25/2024GPT-4 Turbo (Nov 2023), OpenAI · 126.45 · 01/25/2024StarCoder 2 15B, Hugging Face,ServiceNow,NVIDIA,BigCode · 104.7 · 02/20/2024StarCoder 2 7B, Hugging Face,ServiceNow,NVIDIA,BigCode · 93.06 · 02/20/2024Gemma 2B, Google DeepMind · 93.72 · 02/21/2024Gemma 7B, Google DeepMind · 111.79 · 02/21/2024StarCoder 2 3B, Hugging Face,ServiceNow,NVIDIA,BigCode · 88.18 · 02/22/2024Mistral Large, Mistral AI · 122 · 02/26/2024Nemotron-4 15B, NVIDIA · 107.45 · 02/26/2024Claude 3 Opus, Anthropic · 126.91 · 02/29/2024Claude 3 Sonnet, Anthropic · 120.68 · 02/29/2024Yi-9B · 107.35 · 03/01/2024Claude 3 Haiku, Anthropic · 118.3 · 03/07/2024GPT-4 Turbo (Apr 2024), OpenAI · 127.25 · 04/09/2024CodeQwen1.5-7B · 94.38 · 04/15/2024Mixtral 8x22B, Mistral AI · 122 · 04/17/2024Llama 3-70B, Meta AI · 122.91 · 04/18/2024Llama 3-8B, Meta AI · 116.35 · 04/18/2024phi-3-medium 14B, Microsoft · 121.19 · 04/23/2024phi-3-mini 3.8B, Microsoft · 117.25 · 04/23/2024phi-3-small 7.4B, Microsoft · 121.8 · 04/23/2024DeepSeek-V2 (MoE-236B, May 2024), DeepSeek · 124.77 · 05/07/2024Falcon 2 11B, Technology Innovation Institute · 109.32 · 05/09/2024GPT-4o (May 2024), OpenAI · 128.97 · 05/13/2024Gemini 1.5 Pro (May 2024), Google DeepMind · 126.9 · 05/14/2024Gemini 1.5 Flash (May 2024), Google DeepMind · 122.57 · 05/23/2024Mistral 7B v0.3, Mistral AI · 108.78 · 05/27/2024Qwen2-72B, Alibaba · 125.27 · 06/07/2024DeepSeek-Coder-V2-Lite-Base · 108.64 · 06/13/2024Claude 3.5 Sonnet, Anthropic · 130 · 06/20/2024Gemma 2 27B, Google DeepMind · 122.05 · 06/24/2024Gemma 2 9B, Google DeepMind · 119.79 · 06/24/2024GPT-4o mini, OpenAI · 126.56 · 07/18/2024Mistral NeMo, Mistral AI · 118.64 · 07/18/2024Llama 3.1-405B, Meta AI · 128.75 · 07/23/2024Llama 3.1-70B, Meta AI · 125.91 · 07/23/2024Llama 3.1-8B, Meta AI · 116.5 · 07/23/2024Mistral Large 2 (Jul 2024), Mistral AI · 127.54 · 07/24/2024GPT-4o (Aug 2024), OpenAI · 128.77 · 08/06/2024Command R+, Cohere,Cohere Labs (formerly Cohere for AI) · 119.25 · 08/30/2024o1-mini, OpenAI · 135.82 · 09/12/2024o1-preview, OpenAI · 134.78 · 09/12/2024Qwen2.5-32B, Alibaba · 128.52 · 09/17/2024Qwen2.5-Coder (1.5B), Alibaba · 102.66 · 09/18/2024Qwen2.5-Coder (7B), Alibaba · 112.97 · 09/18/2024Qwen2.5-Coder-0.5B · 87.56 · 09/18/2024Qwen2.5-Coder-14B · 116.28 · 09/18/2024Qwen2.5-Coder-32B, Alibaba · 119.42 · 09/18/2024Qwen2.5-Coder-3B · 107.48 · 09/18/2024Qwen2.5-72B, Alibaba · 129 · 09/19/2024Qwen2.5-7B, Alibaba · 118.44 · 09/19/2024Gemini 1.5 Flash (Sep 2024), Google DeepMind · 129.36 · 09/24/2024Gemini 1.5 Pro (Sept 2024), Google DeepMind · 131.73 · 09/24/2024Llama 3.2 1B, Meta AI · 102.43 · 09/24/2024Llama 3.2 90B, Meta AI · 125.5 · 09/24/2024Ministral 3B, Mistral AI · 118.06 · 10/16/2024Claude 3.5 Haiku, Anthropic · 127.15 · 10/22/2024Claude 3.5 Sonnet (October 2024), Anthropic · 133.55 · 10/22/2024Mistral Large 2 (Nov 2024), Mistral AI · 128.52 · 11/18/2024GPT-4o (Nov 2024), OpenAI · 128.81 · 11/20/2024INTELLECT-1, Prime Intellect,Hugging Face,Arcee AI · 100.49 · 11/29/2024Amazon Nova Pro, Amazon · 123.77 · 12/03/2024Llama 3.3 70B, Meta AI · 127.32 · 12/06/2024Gemini 2.0 Flash (Dec 2024), Google DeepMind,Google · 134.71 · 12/11/2024Grok-2 (Dec 2024), xAI · 130.48 · 12/12/2024Phi-4, Microsoft Research · 130.42 · 12/12/2024o1, OpenAI · 141.86 · 12/17/2024DeepSeek-V3, DeepSeek · 132.35 · 12/26/2024DeepSeek-R1, DeepSeek · 138.97 · 01/20/2025DeepSeek-R1-Distill-Qwen-14B, DeepSeek · 135.43 · 01/20/2025DeepSeek-R1-Distill-Qwen-32B, DeepSeek · 137.42 · 01/20/2025Gemini 2.0 Flash Thinking (Jan 2025), Google DeepMind,Google · 135.37 · 01/21/2025Qwen2.5-Max, Alibaba · 132.53 · 01/25/2025Mistral Small 3, Mistral AI · 127.07 · 01/30/2025o3-mini, OpenAI · 140.35 · 01/31/2025Gemini 2.0 Flash (Feb 2025), Google DeepMind,Google · 134.69 · 02/05/2025Gemini 2.0 Pro, Google DeepMind · 135.06 · 02/05/2025Claude 3.7 Sonnet, Anthropic · 141.16 · 02/24/2025GPT-4.5, OpenAI · 136.75 · 02/27/2025QwQ-32B, Alibaba · 137.6 · 03/05/2025Gemma 3 12B, Google DeepMind · 123.46 · 03/12/2025Gemma 3 27B, Google DeepMind · 130.03 · 03/12/2025Gemma 3 4B, Google DeepMind · 115.97 · 03/12/2025Mistral Small 3.1, Mistral AI · 127.48 · 03/17/2025DeepSeek-V3 (Mar 2025), DeepSeek · 135.95 · 03/24/2025Gemini 2.5 Pro (Mar 2025), Google DeepMind · 144.16 · 03/31/2025Llama 4 Scout, Meta AI · 129.64 · 04/05/2025Llama 4 Maverick, Meta AI · 132.2 · 04/06/2025Grok 3, xAI · 138.3 · 04/09/2025GPT-4.1, OpenAI · 136.8 · 04/14/2025GPT-4.1 mini, OpenAI · 135.02 · 04/14/2025GPT-4.1 nano, OpenAI · 129.63 · 04/14/2025o3, OpenAI · 146.91 · 04/16/2025o4-mini, OpenAI · 145.65 · 04/16/2025Gemini 2.5 Flash (Apr 2025), Google DeepMind · 139.97 · 04/17/2025Qwen3-235B-A22B, Alibaba · 139.35 · 04/28/2025Qwen3-8B, Alibaba · 136.18 · 04/28/2025Qwen3-14B, Alibaba · 138.24 · 04/29/2025Qwen3-30B-A3B, Alibaba · 136.19 · 04/29/2025Qwen3-32B, Alibaba · 138.51 · 04/29/2025Gemini 2.5 Pro (May 2025), Google DeepMind · 142.47 · 05/06/2025Mistral Medium 3, Mistral AI · 134.07 · 05/07/2025Gemini 2.5 Flash (May 2025), Google DeepMind · 141.54 · 05/20/2025Claude Opus 4, Anthropic · 142.68 · 05/22/2025Claude Sonnet 4, Anthropic · 141.69 · 05/22/2025DeepSeek-R1 (May 2025), DeepSeek · 141.29 · 05/28/2025Gemini 2.5 Pro (Jun 2025), Google DeepMind · 145.26 · 06/05/2025Magistral Small 1.0, Mistral AI · 133.19 · 06/10/2025o3-pro, OpenAI · 147.45 · 06/10/2025Gemini 2.5 Flash (Jun 2025), Google DeepMind · 140.53 · 06/17/2025Gemini 2.5 Flash-Lite (Jun 2025), Google DeepMind · 133.93 · 06/17/2025Mistral Small 3.2, Mistral AI · 131.74 · 06/20/2025Grok-3 mini, xAI · 140.35 · 06/24/2025Grok 4, xAI · 146.45 · 07/09/2025Kimi K2 (Jul 2025), Moonshot · 140.11 · 07/12/2025Qwen3-235B-A22B-Instruct (Jul 2025), Alibaba · 138.92 · 07/25/2025Qwen3-235B-A22B-Thinking (Jul 2025), Alibaba · 143.88 · 07/25/2025Qwen3-30B-A3B-Instruct (Jul 2025), Alibaba · 137.42 · 07/29/2025Qwen3-30B-A3B-Thinking (Jul 2025), Alibaba · 139.64 · 07/30/2025Claude Opus 4.1, Anthropic · 144.11 · 08/05/2025gpt-oss-120b, OpenAI · 140.1 · 08/05/2025gpt-oss-20b, OpenAI · 137.8 · 08/05/2025GPT-5, OpenAI · 150 · 08/07/2025GPT-5 mini, OpenAI · 145.52 · 08/07/2025GPT-5 nano, OpenAI · 139.39 · 08/07/2025DeepSeek-V3.1, DeepSeek · 139.92 · 08/21/2025Magistral Small 1.2, Mistral AI · 131.41 · 09/18/2025Grok 4 Fast, xAI · 144.21 · 09/19/2025Qwen3-Max, Alibaba · 142.43 · 09/24/2025Gemini 2.5 Flash (Sep 2025), Google DeepMind · 142.99 · 09/25/2025Claude Sonnet 4.5, Anthropic · 146.84 · 09/29/2025DeepSeek-V3.2-Exp, DeepSeek · 145.02 · 09/29/2025GLM-4.6, Z.ai (Zhipu AI),Tsinghua University · 140.78 · 09/30/2025GPT-5 Pro, OpenAI · 150.29 · 10/07/2025Claude Haiku 4.5, Anthropic · 142.37 · 10/15/2025Kimi K2 Thinking, Moonshot · 145.76 · 11/06/2025GPT-5.1, OpenAI · 149.64 · 11/13/2025Gemini 3 Pro, Google DeepMind · 153 · 11/18/2025Claude Opus 4.5, Anthropic · 150.11 · 11/24/2025DeepSeek-V3.2, DeepSeek · 146.19 · 12/01/2025GPT-5.2, OpenAI · 153.51 · 12/11/2025GPT-5.2 Pro, OpenAI · 155.38 · 12/11/2025Gemini 3 Flash, Google DeepMind · 151.84 · 12/17/2025GLM-4.7, Z.ai (Zhipu AI) · 143.46 · 12/22/2025Kimi K2.5, Moonshot · 148.01 · 01/27/2026Claude Opus 4.6, Anthropic · 155.34 · 02/05/2026GPT-5.3 Codex, OpenAI · 156.58 · 02/05/2026GLM-5, Z.ai (Zhipu AI) · 145.85 · 02/11/2026MiniMax-M2.5, MiniMax · 146.51 · 02/12/2026Qwen3.5 397B-A17B, Alibaba · 146.96 · 02/13/2026Qwen 3.5 Plus (hosted 397B-A17B), Alibaba · 146.73 · 02/16/2026Claude Sonnet 4.6, Anthropic · 152.25 · 02/17/2026Grok 4.20, xAI · 152.03 · 02/17/2026Gemini 3.1 Pro, Google DeepMind · 155 · 02/19/2026Qwen3.5-35B-A3B, Alibaba · 142.53 · 02/24/2026Qwen3.5-9B, Alibaba · 139.44 · 02/24/2026Qwen 3.5 Flash (hosted 35B-A3B), Alibaba · 144 · 02/25/2026Gemini 3.1 Flash-Lite, Google · 144.49 · 03/03/2026GPT-5.4, OpenAI · 156.86 · 03/05/2026GPT-5.4 Pro, OpenAI · 158.95 · 03/05/2026GPT-5.4 Mini, OpenAI · 148.96 · 03/17/2026GPT-5.4 Nano, OpenAI · 145.85 · 03/17/2026MiniMax-M2.7, MiniMax · 145.8 · 03/18/2026Qwen 3.6 Plus, Alibaba · 147.67 · 03/31/2026Gemma 4 26B A4B, Google DeepMind · 141.86 · 04/02/2026Gemma 4 31B IT, Google DeepMind · 142.66 · 04/02/2026GLM-5.1, Z.ai (Zhipu AI) · 149.71 · 04/07/2026Muse Spark, Meta AI · 152.12 · 04/08/2026Qwen 3.6 35B-A3B, Alibaba · 143.86 · 04/14/2026Claude Opus 4.7, Anthropic · 156.34 · 04/16/2026Grok 4.3 Beta, xAI · 149.17 · 04/17/2026Kimi K2.6, Moonshot · 151 · 04/20/2026Qwen 3.6 Max (Preview), Alibaba · 149.27 · 04/20/2026Qwen3.6 27B, Alibaba · 146.49 · 04/22/2026GPT-5.5, OpenAI · 159.12 · 04/23/2026GPT-5.5 Pro, OpenAI · 162.25 · 04/23/2026DeepSeek-V4-Flash, DeepSeek · 146.11 · 04/24/2026DeepSeek-V4-Pro, DeepSeek · 149.09 · 04/24/2026Qwen 3.6 Flash, Alibaba · 143.27 · 04/27/2026Mistral Medium 3.5, Mistral AI · 141.28 · 04/28/2026GPT-5.5 Instant, OpenAI · 142.5 · 05/05/2026Gemini 3.5 Flash, Google DeepMind · 154.69 · 05/19/2026Qwen3.7-Max, Alibaba · 153.73 · 05/19/2026Claude Opus 4.8, Anthropic · 158.3 · 05/28/2026MiniMax-M3, MiniMax · 146.47 · 06/01/2026Qwen3.7-Plus, Alibaba · 147.41 · 06/02/2026Nemotron 3 Ultra, NVIDIA · 146.27 · 06/04/2026Claude Fable 5, Anthropic · 163.27 · 06/09/2026Kimi K2.7 Code, Moonshot · 150.07 · 06/12/2026GLM-5.2, Z.ai (Zhipu AI) · 151.86 · 06/16/2026Claude Sonnet 5, Anthropic · 156.22 · 06/30/2026Grok 4.5, xAI · 153.92 · 07/08/2026GPT-5.6 Luna, OpenAI · 156.31 · 07/09/2026GPT-5.6 Sol, OpenAI · 161.81 · 07/09/2026GPT-5.6 Terra, OpenAI · 159.14 · 07/09/2026Muse Spark 1.1, Meta AI · 154.62 · 07/09/2026Inkling, Thinking Machines · 148.64 · 07/15/2026Inkling-Small, Thinking Machines · 150.15 · 07/15/2026Kimi K3, Moonshot · 157.63 · 07/16/2026Gemini 3.5 Flash-Lite, Google DeepMind · 145.13 · 07/21/2026Gemini 3.6 Flash, Google DeepMind · 154.33 · 07/21/2026Claude Opus 5, Anthropic · 162.3 · 07/24/2026Qwen3.7 Flash, Alibaba · 144.63 · 07/27/2026DeepSeek V4 Flash 0731, DeepSeek · 154.49 · 07/31/2026Qwen 3.8 Max, Alibaba · 156.62 · 08/02/2026Muse Spark 1.2, Meta AI · 155.48 · 08/05/2026Grok 4.6, xAI · 156.35 · 08/12/2026DeepSeek V4 Pro 0813, DeepSeek · 155.47 · 08/13/2026Gemini 3.7 Flash, Google DeepMind · 157.44 · 08/13/2026GLM-5.3, Z.ai (Zhipu AI) · 155.25 · 08/14/2026GLM-5.3-Flash, Z.ai (Zhipu AI) · 151.45 · 08/20/2026Claude Fable 5.1, Anthropic · 164.47 · 09/01/2026Qwen3.8 Max (0902), Alibaba · 155.22 · 09/01/2026Gemini 3.8 Flash, Google DeepMind · 156.54 · 09/02/2026GPT-6 Astra, OpenAI · 166.31 · 09/03/2026New high: LLaMA-65B, Meta AI · 109.94 · 02/24/2023New high: GPT-4 (Mar 2023), OpenAI · 125.89 · 03/14/2023New high: GPT-4 Turbo (Nov 2023), OpenAI · 126.45 · 01/25/2024New high: Claude 3 Opus, Anthropic · 126.91 · 02/29/2024New high: GPT-4 Turbo (Apr 2024), OpenAI · 127.25 · 04/09/2024New high: GPT-4o (May 2024), OpenAI · 128.97 · 05/13/2024New high: Claude 3.5 Sonnet, Anthropic · 130 · 06/20/2024New high: o1-mini, OpenAI · 135.82 · 09/12/2024New high: o1, OpenAI · 141.86 · 12/17/2024New high: Gemini 2.5 Pro (Mar 2025), Google DeepMind · 144.16 · 03/31/2025New high: o3, OpenAI · 146.91 · 04/16/2025New high: o3-pro, OpenAI · 147.45 · 06/10/2025New high: GPT-5, OpenAI · 150 · 08/07/2025New high: GPT-5 Pro, OpenAI · 150.29 · 10/07/2025New high: Gemini 3 Pro, Google DeepMind · 153 · 11/18/2025New high: GPT-5.2 Pro, OpenAI · 155.38 · 12/11/2025New high: GPT-5.3 Codex, OpenAI · 156.58 · 02/05/2026New high: GPT-5.4 Pro, OpenAI · 158.95 · 03/05/2026New high: GPT-5.5 Pro, OpenAI · 162.25 · 04/23/2026New high: Claude Fable 5, Anthropic · 163.27 · 06/09/2026New high: Claude Fable 5.1, Anthropic · 164.47 · 09/01/2026New high: GPT-6 Astra, OpenAI · 166.31 · 09/03/2026LLaMA-65BClaude 3 Opuso1-miniGemini 2.5 ProGPT-5 ProFable 5GPT-6 Astra1Level 1, Direct prompting: ChatGPT, Nov 30, 20222Level 2, Added context: Perplexity, Dec 7, 20223Level 3, Workflows: OpenAI steps in Zapier, Dec 9, 20224Level 4, Tool use: ChatGPT plugins, Mar 23, 20235Level 5, Agent loops: OpenAI Operator, Jan 23, 20256Level 6, Teams of Agents: How we built our multi-agent research system, Jun 13, 20257Level 7, Always-on agents: Grok Bot, Aug 11, 2026Selected milestones by levelEpoch Capabilities Index by model release date, 2023 to 09/20/2026266 models scored by Epoch AI, plotted by release date. A stepped line follows the best score to date, from 109.94 to 166.31. Markers under the axis show selected milestones for this site's levels, and a vertical line marks the first reasoning model.4060801001201401601802023202420252026ECI scoreReasoning · Sep 2024LLaMA-13B, Meta AI · 100.2 · 02/24/2023LLaMA-33B, Meta AI · 107.14 · 02/24/2023LLaMA-65B, Meta AI · 109.94 · 02/24/2023LLaMA-7B, Meta AI · 96.25 · 02/24/2023GPT-4 (Mar 2023), OpenAI · 125.89 · 03/14/2023Cerebras-GPT-13B, Cerebras Systems · 82.57 · 03/20/2023Dolly 2.0-12b, Databricks · 89.11 · 04/11/2023vicuna-13b-v1.1 · 94.19 · 04/12/2023stablelm-tuned-alpha-7b · 54.64 · 04/19/2023Falcon-7B, Technology Innovation Institute · 94.63 · 04/24/2023RedPajama-INCITE-7B-Base · 89.87 · 05/04/2023MPT-7B, MosaicML · 94.11 · 05/05/2023PaLM 2-L · 114.88 · 05/17/2023PaLM 2-M · 108.04 · 05/17/2023PaLM 2-S · 105.88 · 05/17/2023Falcon-40B, Technology Innovation Institute · 104.13 · 05/25/2023Baichuan1-7B, Baichuan · 89.92 · 06/01/2023open_llama_7b · 91.12 · 06/07/2023GPT-3.5 Turbo (Jun 2023), OpenAI · 113.19 · 06/13/2023GPT-4 (Jun 2023), OpenAI · 123.1 · 06/13/2023MPT-30B, MosaicML · 100.28 · 06/22/2023chatglm2-6b · 98.53 · 06/24/2023XGen-7B, Salesforce · 92.9 · 06/27/2023internlm-7b · 102.53 · 07/05/2023Claude 2, Anthropic · 120.09 · 07/11/2023Llama 2-13B, Meta AI · 105.88 · 07/18/2023Llama 2-34B, Meta AI · 104.9 · 07/18/2023Llama 2-70B, Meta AI · 113.63 · 07/18/2023Llama 2-7B, Meta AI · 98.66 · 07/18/2023Stable Beluga 2, Stability AI · 117 · 07/20/2023Claude Instant, Anthropic · 120.2 · 08/09/2023Baichuan2-13B, Baichuan · 102.86 · 09/06/2023Falcon-180B, Technology Innovation Institute · 111.94 · 09/06/2023Phi-1.5, Microsoft · 90.99 · 09/11/2023internlm-20b · 111.89 · 09/18/2023Baichuan 2-7B, Baichuan · 95.88 · 09/20/2023Qwen-14B, Alibaba · 112.85 · 09/24/2023Mistral 7B v0.1, Mistral AI · 112.02 · 09/27/2023Qwen-7B, Alibaba · 106.55 · 09/28/2023DeepSeek Coder 1.3B, DeepSeek,Peking University · 62.54 · 11/02/2023DeepSeek Coder 33B, DeepSeek,Peking University · 95.87 · 11/02/2023DeepSeek Coder 6.7B, DeepSeek,Peking University · 89.05 · 11/02/2023Yi-34B, 01.AI · 117.3 · 11/02/2023GPT-3.5 Turbo (Nov 2023), OpenAI · 118.49 · 11/06/2023Claude 2.1, Anthropic · 119.22 · 11/21/2023Yi 6B, 01.AI · 104.46 · 11/22/2023DeepSeek LLM 67B, DeepSeek · 110.56 · 11/29/2023Qwen-1_8B · 92.43 · 11/30/2023Mixtral 8x7B, Mistral AI · 118.4 · 12/11/2023Phi-2, Microsoft · 107.69 · 12/12/2023Gemini 1.0 Pro, Google DeepMind · 116.97 · 12/13/2023GPT-3.5 Turbo (Jan 2024), OpenAI · 115.61 · 01/25/2024GPT-4 Turbo (Nov 2023), OpenAI · 126.45 · 01/25/2024StarCoder 2 15B, Hugging Face,ServiceNow,NVIDIA,BigCode · 104.7 · 02/20/2024StarCoder 2 7B, Hugging Face,ServiceNow,NVIDIA,BigCode · 93.06 · 02/20/2024Gemma 2B, Google DeepMind · 93.72 · 02/21/2024Gemma 7B, Google DeepMind · 111.79 · 02/21/2024StarCoder 2 3B, Hugging Face,ServiceNow,NVIDIA,BigCode · 88.18 · 02/22/2024Mistral Large, Mistral AI · 122 · 02/26/2024Nemotron-4 15B, NVIDIA · 107.45 · 02/26/2024Claude 3 Opus, Anthropic · 126.91 · 02/29/2024Claude 3 Sonnet, Anthropic · 120.68 · 02/29/2024Yi-9B · 107.35 · 03/01/2024Claude 3 Haiku, Anthropic · 118.3 · 03/07/2024GPT-4 Turbo (Apr 2024), OpenAI · 127.25 · 04/09/2024CodeQwen1.5-7B · 94.38 · 04/15/2024Mixtral 8x22B, Mistral AI · 122 · 04/17/2024Llama 3-70B, Meta AI · 122.91 · 04/18/2024Llama 3-8B, Meta AI · 116.35 · 04/18/2024phi-3-medium 14B, Microsoft · 121.19 · 04/23/2024phi-3-mini 3.8B, Microsoft · 117.25 · 04/23/2024phi-3-small 7.4B, Microsoft · 121.8 · 04/23/2024DeepSeek-V2 (MoE-236B, May 2024), DeepSeek · 124.77 · 05/07/2024Falcon 2 11B, Technology Innovation Institute · 109.32 · 05/09/2024GPT-4o (May 2024), OpenAI · 128.97 · 05/13/2024Gemini 1.5 Pro (May 2024), Google DeepMind · 126.9 · 05/14/2024Gemini 1.5 Flash (May 2024), Google DeepMind · 122.57 · 05/23/2024Mistral 7B v0.3, Mistral AI · 108.78 · 05/27/2024Qwen2-72B, Alibaba · 125.27 · 06/07/2024DeepSeek-Coder-V2-Lite-Base · 108.64 · 06/13/2024Claude 3.5 Sonnet, Anthropic · 130 · 06/20/2024Gemma 2 27B, Google DeepMind · 122.05 · 06/24/2024Gemma 2 9B, Google DeepMind · 119.79 · 06/24/2024GPT-4o mini, OpenAI · 126.56 · 07/18/2024Mistral NeMo, Mistral AI · 118.64 · 07/18/2024Llama 3.1-405B, Meta AI · 128.75 · 07/23/2024Llama 3.1-70B, Meta AI · 125.91 · 07/23/2024Llama 3.1-8B, Meta AI · 116.5 · 07/23/2024Mistral Large 2 (Jul 2024), Mistral AI · 127.54 · 07/24/2024GPT-4o (Aug 2024), OpenAI · 128.77 · 08/06/2024Command R+, Cohere,Cohere Labs (formerly Cohere for AI) · 119.25 · 08/30/2024o1-mini, OpenAI · 135.82 · 09/12/2024o1-preview, OpenAI · 134.78 · 09/12/2024Qwen2.5-32B, Alibaba · 128.52 · 09/17/2024Qwen2.5-Coder (1.5B), Alibaba · 102.66 · 09/18/2024Qwen2.5-Coder (7B), Alibaba · 112.97 · 09/18/2024Qwen2.5-Coder-0.5B · 87.56 · 09/18/2024Qwen2.5-Coder-14B · 116.28 · 09/18/2024Qwen2.5-Coder-32B, Alibaba · 119.42 · 09/18/2024Qwen2.5-Coder-3B · 107.48 · 09/18/2024Qwen2.5-72B, Alibaba · 129 · 09/19/2024Qwen2.5-7B, Alibaba · 118.44 · 09/19/2024Gemini 1.5 Flash (Sep 2024), Google DeepMind · 129.36 · 09/24/2024Gemini 1.5 Pro (Sept 2024), Google DeepMind · 131.73 · 09/24/2024Llama 3.2 1B, Meta AI · 102.43 · 09/24/2024Llama 3.2 90B, Meta AI · 125.5 · 09/24/2024Ministral 3B, Mistral AI · 118.06 · 10/16/2024Claude 3.5 Haiku, Anthropic · 127.15 · 10/22/2024Claude 3.5 Sonnet (October 2024), Anthropic · 133.55 · 10/22/2024Mistral Large 2 (Nov 2024), Mistral AI · 128.52 · 11/18/2024GPT-4o (Nov 2024), OpenAI · 128.81 · 11/20/2024INTELLECT-1, Prime Intellect,Hugging Face,Arcee AI · 100.49 · 11/29/2024Amazon Nova Pro, Amazon · 123.77 · 12/03/2024Llama 3.3 70B, Meta AI · 127.32 · 12/06/2024Gemini 2.0 Flash (Dec 2024), Google DeepMind,Google · 134.71 · 12/11/2024Grok-2 (Dec 2024), xAI · 130.48 · 12/12/2024Phi-4, Microsoft Research · 130.42 · 12/12/2024o1, OpenAI · 141.86 · 12/17/2024DeepSeek-V3, DeepSeek · 132.35 · 12/26/2024DeepSeek-R1, DeepSeek · 138.97 · 01/20/2025DeepSeek-R1-Distill-Qwen-14B, DeepSeek · 135.43 · 01/20/2025DeepSeek-R1-Distill-Qwen-32B, DeepSeek · 137.42 · 01/20/2025Gemini 2.0 Flash Thinking (Jan 2025), Google DeepMind,Google · 135.37 · 01/21/2025Qwen2.5-Max, Alibaba · 132.53 · 01/25/2025Mistral Small 3, Mistral AI · 127.07 · 01/30/2025o3-mini, OpenAI · 140.35 · 01/31/2025Gemini 2.0 Flash (Feb 2025), Google DeepMind,Google · 134.69 · 02/05/2025Gemini 2.0 Pro, Google DeepMind · 135.06 · 02/05/2025Claude 3.7 Sonnet, Anthropic · 141.16 · 02/24/2025GPT-4.5, OpenAI · 136.75 · 02/27/2025QwQ-32B, Alibaba · 137.6 · 03/05/2025Gemma 3 12B, Google DeepMind · 123.46 · 03/12/2025Gemma 3 27B, Google DeepMind · 130.03 · 03/12/2025Gemma 3 4B, Google DeepMind · 115.97 · 03/12/2025Mistral Small 3.1, Mistral AI · 127.48 · 03/17/2025DeepSeek-V3 (Mar 2025), DeepSeek · 135.95 · 03/24/2025Gemini 2.5 Pro (Mar 2025), Google DeepMind · 144.16 · 03/31/2025Llama 4 Scout, Meta AI · 129.64 · 04/05/2025Llama 4 Maverick, Meta AI · 132.2 · 04/06/2025Grok 3, xAI · 138.3 · 04/09/2025GPT-4.1, OpenAI · 136.8 · 04/14/2025GPT-4.1 mini, OpenAI · 135.02 · 04/14/2025GPT-4.1 nano, OpenAI · 129.63 · 04/14/2025o3, OpenAI · 146.91 · 04/16/2025o4-mini, OpenAI · 145.65 · 04/16/2025Gemini 2.5 Flash (Apr 2025), Google DeepMind · 139.97 · 04/17/2025Qwen3-235B-A22B, Alibaba · 139.35 · 04/28/2025Qwen3-8B, Alibaba · 136.18 · 04/28/2025Qwen3-14B, Alibaba · 138.24 · 04/29/2025Qwen3-30B-A3B, Alibaba · 136.19 · 04/29/2025Qwen3-32B, Alibaba · 138.51 · 04/29/2025Gemini 2.5 Pro (May 2025), Google DeepMind · 142.47 · 05/06/2025Mistral Medium 3, Mistral AI · 134.07 · 05/07/2025Gemini 2.5 Flash (May 2025), Google DeepMind · 141.54 · 05/20/2025Claude Opus 4, Anthropic · 142.68 · 05/22/2025Claude Sonnet 4, Anthropic · 141.69 · 05/22/2025DeepSeek-R1 (May 2025), DeepSeek · 141.29 · 05/28/2025Gemini 2.5 Pro (Jun 2025), Google DeepMind · 145.26 · 06/05/2025Magistral Small 1.0, Mistral AI · 133.19 · 06/10/2025o3-pro, OpenAI · 147.45 · 06/10/2025Gemini 2.5 Flash (Jun 2025), Google DeepMind · 140.53 · 06/17/2025Gemini 2.5 Flash-Lite (Jun 2025), Google DeepMind · 133.93 · 06/17/2025Mistral Small 3.2, Mistral AI · 131.74 · 06/20/2025Grok-3 mini, xAI · 140.35 · 06/24/2025Grok 4, xAI · 146.45 · 07/09/2025Kimi K2 (Jul 2025), Moonshot · 140.11 · 07/12/2025Qwen3-235B-A22B-Instruct (Jul 2025), Alibaba · 138.92 · 07/25/2025Qwen3-235B-A22B-Thinking (Jul 2025), Alibaba · 143.88 · 07/25/2025Qwen3-30B-A3B-Instruct (Jul 2025), Alibaba · 137.42 · 07/29/2025Qwen3-30B-A3B-Thinking (Jul 2025), Alibaba · 139.64 · 07/30/2025Claude Opus 4.1, Anthropic · 144.11 · 08/05/2025gpt-oss-120b, OpenAI · 140.1 · 08/05/2025gpt-oss-20b, OpenAI · 137.8 · 08/05/2025GPT-5, OpenAI · 150 · 08/07/2025GPT-5 mini, OpenAI · 145.52 · 08/07/2025GPT-5 nano, OpenAI · 139.39 · 08/07/2025DeepSeek-V3.1, DeepSeek · 139.92 · 08/21/2025Magistral Small 1.2, Mistral AI · 131.41 · 09/18/2025Grok 4 Fast, xAI · 144.21 · 09/19/2025Qwen3-Max, Alibaba · 142.43 · 09/24/2025Gemini 2.5 Flash (Sep 2025), Google DeepMind · 142.99 · 09/25/2025Claude Sonnet 4.5, Anthropic · 146.84 · 09/29/2025DeepSeek-V3.2-Exp, DeepSeek · 145.02 · 09/29/2025GLM-4.6, Z.ai (Zhipu AI),Tsinghua University · 140.78 · 09/30/2025GPT-5 Pro, OpenAI · 150.29 · 10/07/2025Claude Haiku 4.5, Anthropic · 142.37 · 10/15/2025Kimi K2 Thinking, Moonshot · 145.76 · 11/06/2025GPT-5.1, OpenAI · 149.64 · 11/13/2025Gemini 3 Pro, Google DeepMind · 153 · 11/18/2025Claude Opus 4.5, Anthropic · 150.11 · 11/24/2025DeepSeek-V3.2, DeepSeek · 146.19 · 12/01/2025GPT-5.2, OpenAI · 153.51 · 12/11/2025GPT-5.2 Pro, OpenAI · 155.38 · 12/11/2025Gemini 3 Flash, Google DeepMind · 151.84 · 12/17/2025GLM-4.7, Z.ai (Zhipu AI) · 143.46 · 12/22/2025Kimi K2.5, Moonshot · 148.01 · 01/27/2026Claude Opus 4.6, Anthropic · 155.34 · 02/05/2026GPT-5.3 Codex, OpenAI · 156.58 · 02/05/2026GLM-5, Z.ai (Zhipu AI) · 145.85 · 02/11/2026MiniMax-M2.5, MiniMax · 146.51 · 02/12/2026Qwen3.5 397B-A17B, Alibaba · 146.96 · 02/13/2026Qwen 3.5 Plus (hosted 397B-A17B), Alibaba · 146.73 · 02/16/2026Claude Sonnet 4.6, Anthropic · 152.25 · 02/17/2026Grok 4.20, xAI · 152.03 · 02/17/2026Gemini 3.1 Pro, Google DeepMind · 155 · 02/19/2026Qwen3.5-35B-A3B, Alibaba · 142.53 · 02/24/2026Qwen3.5-9B, Alibaba · 139.44 · 02/24/2026Qwen 3.5 Flash (hosted 35B-A3B), Alibaba · 144 · 02/25/2026Gemini 3.1 Flash-Lite, Google · 144.49 · 03/03/2026GPT-5.4, OpenAI · 156.86 · 03/05/2026GPT-5.4 Pro, OpenAI · 158.95 · 03/05/2026GPT-5.4 Mini, OpenAI · 148.96 · 03/17/2026GPT-5.4 Nano, OpenAI · 145.85 · 03/17/2026MiniMax-M2.7, MiniMax · 145.8 · 03/18/2026Qwen 3.6 Plus, Alibaba · 147.67 · 03/31/2026Gemma 4 26B A4B, Google DeepMind · 141.86 · 04/02/2026Gemma 4 31B IT, Google DeepMind · 142.66 · 04/02/2026GLM-5.1, Z.ai (Zhipu AI) · 149.71 · 04/07/2026Muse Spark, Meta AI · 152.12 · 04/08/2026Qwen 3.6 35B-A3B, Alibaba · 143.86 · 04/14/2026Claude Opus 4.7, Anthropic · 156.34 · 04/16/2026Grok 4.3 Beta, xAI · 149.17 · 04/17/2026Kimi K2.6, Moonshot · 151 · 04/20/2026Qwen 3.6 Max (Preview), Alibaba · 149.27 · 04/20/2026Qwen3.6 27B, Alibaba · 146.49 · 04/22/2026GPT-5.5, OpenAI · 159.12 · 04/23/2026GPT-5.5 Pro, OpenAI · 162.25 · 04/23/2026DeepSeek-V4-Flash, DeepSeek · 146.11 · 04/24/2026DeepSeek-V4-Pro, DeepSeek · 149.09 · 04/24/2026Qwen 3.6 Flash, Alibaba · 143.27 · 04/27/2026Mistral Medium 3.5, Mistral AI · 141.28 · 04/28/2026GPT-5.5 Instant, OpenAI · 142.5 · 05/05/2026Gemini 3.5 Flash, Google DeepMind · 154.69 · 05/19/2026Qwen3.7-Max, Alibaba · 153.73 · 05/19/2026Claude Opus 4.8, Anthropic · 158.3 · 05/28/2026MiniMax-M3, MiniMax · 146.47 · 06/01/2026Qwen3.7-Plus, Alibaba · 147.41 · 06/02/2026Nemotron 3 Ultra, NVIDIA · 146.27 · 06/04/2026Claude Fable 5, Anthropic · 163.27 · 06/09/2026Kimi K2.7 Code, Moonshot · 150.07 · 06/12/2026GLM-5.2, Z.ai (Zhipu AI) · 151.86 · 06/16/2026Claude Sonnet 5, Anthropic · 156.22 · 06/30/2026Grok 4.5, xAI · 153.92 · 07/08/2026GPT-5.6 Luna, OpenAI · 156.31 · 07/09/2026GPT-5.6 Sol, OpenAI · 161.81 · 07/09/2026GPT-5.6 Terra, OpenAI · 159.14 · 07/09/2026Muse Spark 1.1, Meta AI · 154.62 · 07/09/2026Inkling, Thinking Machines · 148.64 · 07/15/2026Inkling-Small, Thinking Machines · 150.15 · 07/15/2026Kimi K3, Moonshot · 157.63 · 07/16/2026Gemini 3.5 Flash-Lite, Google DeepMind · 145.13 · 07/21/2026Gemini 3.6 Flash, Google DeepMind · 154.33 · 07/21/2026Claude Opus 5, Anthropic · 162.3 · 07/24/2026Qwen3.7 Flash, Alibaba · 144.63 · 07/27/2026DeepSeek V4 Flash 0731, DeepSeek · 154.49 · 07/31/2026Qwen 3.8 Max, Alibaba · 156.62 · 08/02/2026Muse Spark 1.2, Meta AI · 155.48 · 08/05/2026Grok 4.6, xAI · 156.35 · 08/12/2026DeepSeek V4 Pro 0813, DeepSeek · 155.47 · 08/13/2026Gemini 3.7 Flash, Google DeepMind · 157.44 · 08/13/2026GLM-5.3, Z.ai (Zhipu AI) · 155.25 · 08/14/2026GLM-5.3-Flash, Z.ai (Zhipu AI) · 151.45 · 08/20/2026Claude Fable 5.1, Anthropic · 164.47 · 09/01/2026Qwen3.8 Max (0902), Alibaba · 155.22 · 09/01/2026Gemini 3.8 Flash, Google DeepMind · 156.54 · 09/02/2026GPT-6 Astra, OpenAI · 166.31 · 09/03/2026New high: LLaMA-65B, Meta AI · 109.94 · 02/24/2023New high: GPT-4 (Mar 2023), OpenAI · 125.89 · 03/14/2023New high: GPT-4 Turbo (Nov 2023), OpenAI · 126.45 · 01/25/2024New high: Claude 3 Opus, Anthropic · 126.91 · 02/29/2024New high: GPT-4 Turbo (Apr 2024), OpenAI · 127.25 · 04/09/2024New high: GPT-4o (May 2024), OpenAI · 128.97 · 05/13/2024New high: Claude 3.5 Sonnet, Anthropic · 130 · 06/20/2024New high: o1-mini, OpenAI · 135.82 · 09/12/2024New high: o1, OpenAI · 141.86 · 12/17/2024New high: Gemini 2.5 Pro (Mar 2025), Google DeepMind · 144.16 · 03/31/2025New high: o3, OpenAI · 146.91 · 04/16/2025New high: o3-pro, OpenAI · 147.45 · 06/10/2025New high: GPT-5, OpenAI · 150 · 08/07/2025New high: GPT-5 Pro, OpenAI · 150.29 · 10/07/2025New high: Gemini 3 Pro, Google DeepMind · 153 · 11/18/2025New high: GPT-5.2 Pro, OpenAI · 155.38 · 12/11/2025New high: GPT-5.3 Codex, OpenAI · 156.58 · 02/05/2026New high: GPT-5.4 Pro, OpenAI · 158.95 · 03/05/2026New high: GPT-5.5 Pro, OpenAI · 162.25 · 04/23/2026New high: Claude Fable 5, Anthropic · 163.27 · 06/09/2026New high: Claude Fable 5.1, Anthropic · 164.47 · 09/01/2026New high: GPT-6 Astra, OpenAI · 166.31 · 09/03/2026LLaMA-65BClaude 3.5 SonnetFable 5GPT-6 Astra1Level 1, Direct prompting: ChatGPT, Nov 30, 20222Level 2, Added context: Perplexity, Dec 7, 20223Level 3, Workflows: OpenAI steps in Zapier, Dec 9, 20224Level 4, Tool use: ChatGPT plugins, Mar 23, 20235Level 5, Agent loops: OpenAI Operator, Jan 23, 20256Level 6, Teams of Agents: How we built our multi-agent research system, Jun 13, 20257Level 7, Always-on agents: Grok Bot, Aug 11, 2026Selected milestones by level
  • A model, closed weights
  • A model, open weights
  • Best score to date
  • A selected milestone for a level

Anthropic · ECI 163.27 · Highest score Jun 9, 2026–Aug 31, 2026

Epoch AI’s Capabilities Index combines scores from many different AI benchmarks into a single “general capability” scale, allowing comparisons between models even over timespans long enough for single benchmarks to reach saturation. Epoch adds that “Absolute ECI values are meaningless by themselves, but meaningful comparisons can be made between models.” So read the shape, not the numbers. Record periods are reconstructed from this snapshot’s scores and release dates, not historical leaderboard snapshots. Highest ECI does not mean best at every task.

The shape has a bend in it. From GPT-4 on Mar 14, 2023 to the day before the first reasoning model, 18 months, the best score rose 4.1 points, from 125.89 to 130. In the 24 months since, it has risen 36.3, to 166.31 (GPT-6 Astra, Sep 3, 2026). The selected milestones for levels 1 to 4 fall before the bend; the selected milestones for levels 5 to 7 fall after it. This reflects the editorial selection, not the first emergence of those patterns. Both figures are this site’s subtraction over Epoch’s scores. The index starts in early 2023, so it says nothing about what came before.

The 22 models that set a new high
ReleasedModelMakerScore
02/24/2023LLaMA-65BMeta AI109.94
03/14/2023GPT-4 (Mar 2023)OpenAI125.89
01/25/2024GPT-4 Turbo (Nov 2023)OpenAI126.45
02/29/2024Claude 3 OpusAnthropic126.91
04/09/2024GPT-4 Turbo (Apr 2024)OpenAI127.25
05/13/2024GPT-4o (May 2024)OpenAI128.97
06/20/2024Claude 3.5 SonnetAnthropic130
09/12/2024o1-miniOpenAI135.82
12/17/2024o1OpenAI141.86
03/31/2025Gemini 2.5 Pro (Mar 2025)Google DeepMind144.16
04/16/2025o3OpenAI146.91
06/10/2025o3-proOpenAI147.45
08/07/2025GPT-5OpenAI150
10/07/2025GPT-5 ProOpenAI150.29
11/18/2025Gemini 3 ProGoogle DeepMind153
12/11/2025GPT-5.2 ProOpenAI155.38
02/05/2026GPT-5.3 CodexOpenAI156.58
03/05/2026GPT-5.4 ProOpenAI158.95
04/23/2026GPT-5.5 ProOpenAI162.25
06/09/2026Claude Fable 5Anthropic163.27
09/01/2026Claude Fable 5.1Anthropic164.47
09/03/2026GPT-6 AstraOpenAI166.31

Data: Epoch Capabilities Index, Epoch AI, reused under CC BY 4.0, retrieved 09/18/2026; 266 models. Scores, release dates and makers are Epoch’s, unchanged apart from rounding. The level markers and the reasoning line are this site’s. Best score on the day of the snapshot: 166.31.

Published measures

What outside researchers have measured

Quoted exactly, with the scope the source itself gives. The site does not compute or extend these.

“Our original time horizon dataset, released in March 2025, showed a smooth trend with the frontier time-horizon doubling around every 7 months over the period 2019 to 2025.”

Time Horizon 1.1 · METR · 01/29/2026

What it measures: The 50%-reliability time horizon. METR's March 2025 post states the finding this way: "The length of tasks (measured by how long they take human professionals) that generalist frontier model agents can complete autonomously with 50% reliability has been doubling approximately every 7 months for the last 6 years." The length is a human time, not the model's; the tasks are METR's own suite of software and research tasks, not work in general.

Scope: METR's original (TH1) dataset: frontier models released 2019 to 2025, scored on METR's own task suite. The quotation is METR's January 2026 restatement of that earlier finding. METR's March 2025 post now carries this banner: "⚠️ Some of the text and figures in this post are out of date. The interactive chart below is kept up to date, but the static figures and some claims in the text (e.g. the doubling time) reflect the state of the data at the time of original publication." A doubling time measured over a past period is a description of that period. It is not a forecast, and this site does not extend it forward.

“The post-2023 doubling-time is 131 days under TH1.1, compared to 165 days under TH1, meaning progress is estimated to be 20% more rapid under TH1.1.”

Time Horizon 1.1 · METR · 01/29/2026

What it measures: The same 50%-reliability time horizon, re-estimated under METR's revised TH1.1 methodology and restricted to the period since 2023. METR says of the suite: "We increased our suite from 170 to 228 tasks." Evaluation also moved to Inspect, which METR describes as "a widely-adopted open-source framework for AI evaluations developed by the UK AI Security Institute".

Scope: Models released since 2023 only. METR says "We have re-estimated the effective time horizons for 14 models", of the 33 that had estimates under TH1, excluding the rest for "the model no longer being publicly available, (ii) the model requiring significant changes to the tool-calling scaffold (e.g. for GPT-2, GPT-3, and GPT-3.5), or (iii) because the model was far from the capability frontier at the time of release, so unlikely to change the estimated trend." METR also reports that two models scored significantly higher under its previous harness than under Inspect, so the numbers are sensitive to the scaffold. This is a measurement of METR's task suite, not of real-world software work, and not a prediction of anything.

Every milestone

Grouped by year, newest first

Newest year first, newest date first within a year: the direction a reader checking "what's new" would scan. This list is the chart's full text alternative: every milestone above appears here once, with its date, maker, one-sentence description, source and links.

2026

Sep 16, 2026

Cowork merges into ClaudeLevel 7

Anthropic

Anthropic's announcement folds a separate always-on product back into the chat app: "hand over a report due at noon, and Claude takes it from there, even after you've closed your laptop", and "You can check progress from your phone on the way to the office." Cowork's capabilities are now "available from any conversation, with the context, skills, and connectors you already have."

Sep 15, 2026

Jev, a typed decision model (early access)Level 1

TypeSafe AI

TypeSafe's founder, Diogo Almeida, announces "a new class of frontier models built to make fast, structured decisions that software can use directly": "unstructured state in, typed probabilistic decisions out." Jev writes no text. Its possible outputs are defined in advance, every answer carries a probability, and all of it comes back in one parallel pass, not word by word. TypeSafe quotes "70ms-500ms" a call and "$0.042 / MTok" of input. Its own list of uses is this site's level 3: "classify, route, score, extract, or branch where hand-written logic is too brittle."

Sep 10, 2026

DeepSeek-V4.1-FlashLevel 1

DeepSeek

DeepSeek's release note of September 10, 2026 says "V4.1-Flash is now live on the DeepSeek API with native multimodal support", describing a mixture-of-experts model with 552B parameters and 8B active for input, 16B for output: a frontier-class open-weight release, one call at a time.

Sep 8, 2026

MuseLevel 7

Meta

Meta's announcement says "Muse is a personal AI agent. It doesn't just answer questions, it actually does the work", and that "A separate Sentinel agent runs on that same machine, kept apart from Muse at the system level. Nothing Muse does reaches the internet unless the Sentinel approves it, and it asks the person for permission when needed." A checking agent shipped inside a consumer product, not only proposed in a paper.

Aug 11, 2026

Grok BotLevel 7

SpaceXAI

xAI's announcement says "Grok Bot is your team of always-on agents" and, of those agents, "They have their own computer, work inside tools and apps like you do, and keep working 24/7." It says Grok Bot "is in beta and available today" to named SuperGrok and Cursor subscriber tiers on desktop and iOS.

Aug 10, 2026

Flowise sunsetLevel 3

Flowise

Flowise's sunset notice says "we've decided to wind down our operations for Flowise", with a feature freeze on July 29, 2026 and the repository archived on August 10, 2026, adding that "Flowise source code will still remain on Github and the Apache 2.0 licensed code is yours to keep building on." A visual workflow builder closing while the pattern it built stayed in use.

Jul 16, 2026

NotebookLM becomes Gemini NotebookLevel 2

Google

Google's announcement says: "We're renaming NotebookLM to Gemini Notebook. It's the same standalone product, now doing more across the Google ecosystem and updated with a secure cloud computer." The product whose whole premise is answering only from the documents you gave it is still the plainest consumer example of level 2.

Jul 9, 2026

ChatGPT WorkLevel 7

OpenAI

OpenAI's product page says “Powered by GPT‑5.6, ChatGPT Work brings together context from your team’s tools to turn scattered notes, drafts, and ideas into finished work — and keeps projects moving while you stay in control”, and that it “gathers context, plans the approach, and takes action across your tools, files, and desktop apps”.

Not verified today: openai.com/chatgpt-work/ returns 403 to this site's fetcher, matching content/landscape.json's own note. The Internet Archive's capture of July 14, 2026 opens and shows the wording quoted here, but the page carries no date, so July 9, 2026 is still unverified and this milestone is not used as a marked date.
Jul 7, 2026

Claude Cowork runs without your computerLevel 7

Anthropic

Anthropic's dated release notes record, on July 7, 2026: "Cowork runs your sessions remotely (in beta), so your sessions and files are saved to your Claude account and go where you go, on any device. Work continues when you close your laptop, and scheduled tasks run with no device online." This is the entry that makes Cowork an always-on product; the January release ran on the person's own machine.

Jun 30, 2026

Gemini Spark in subscribers' handsLevel 7

Google

Google's post of June 30, 2026 says "Gemini Spark for macOS is available in Beta to Google AI Ultra subscribers aged 18 and over, starting in the US." The May announcement had described Spark as "a 24/7 personal AI agent" that takes "recurring tasks or triggers" and "keeps working in the background even when you close your laptop or lock your phone."

May 19, 2026

Gemini CLI becomes Antigravity CLILevel 5

Google

Google's post says "we're unifying our efforts into Google Antigravity, our premier agent-first development platform", and that "On June 18, 2026, Gemini CLI and Gemini Code Assist IDE extensions will stop serving requests for Google AI Pro and Ultra, as well as those using it free of charge using Gemini Code Assist for individuals." Enterprise licenses keep Gemini CLI.

May 19, 2026

Gemini SparkLevel 7

Google

Google's announcement introduces "Gemini Spark, a 24/7 personal AI agent that helps you navigate your digital life", "deeply integrated with the Workspace tools you rely on daily, like Gmail, Docs, Slides and more", and says that "because it is a cloud-based agent, Spark keeps working in the background even when you close your laptop or lock your phone."

Apr 16, 2026

π0.7Level 7

Physical Intelligence

Physical Intelligence's post calls π0.7 "a steerable generalist model that can perform dexterous tasks across robots, scenes, and skills", and reports "compositional generalization, recombining skills from various tasks to solve new problems", including a robot folding laundry with no laundry-folding data in its training.

Jan 12, 2026

Claude CoworkLevel 5

Anthropic

Anthropic's dated release notes record, on January 12, 2026, “Cowork research preview on Claude Desktop (macOS only) for Max plans”, and say Cowork “brings Claude Code’s agentic capabilities to the Claude desktop app for knowledge work beyond coding”. Pro plans followed on January 16, 2026 and the same notes record “Claude Cowork is now generally available on macOS and Windows through the Claude Desktop app” on April 9, 2026.

2025

Nov 24, 2025

OpenClawLevel 7

OpenClaw Foundation

GitHub's record of openclaw/openclaw gives a creation date of November 24, 2025. Its README says “OpenClaw is an open-source AI assistant that runs on your own computer and meets you in the channels you already use”, with “One Gateway” running it “as a personal assistant on a laptop or as a shared team deployment”, and carries an MIT license badge.

Oct 16, 2025

Agent SkillsLevel 5

Anthropic

Anthropic's announcement says "Skills are folders that include instructions, scripts, and resources that Claude can load when needed", and that "Claude will only access a skill when it's relevant to the task at hand": the agent deciding which of its own instructions to read.

Jul 22, 2025

Hermes AgentLevel 7

Nous Research

GitHub's record of NousResearch/hermes-agent gives a creation date of July 22, 2025 and an MIT license. The README calls it “The self-improving AI agent built by Nous Research” and says “Run it on a $5 VPS, a GPU cluster, or serverless infrastructure that costs nearly nothing when idle”; the project's own site (hermes-agent.nousresearch.com) says “Tasks that need an active agent will not run while it is stopped; hosted agents are managed separately in Nous Portal.” That is standing work that runs because the agent is running, on hardware the developer keeps up.

Jul 17, 2025

ChatGPT agentLevel 5

OpenAI

OpenAI's post of July 17, 2025 says “ChatGPT can now do work for you using its own computer, handling complex tasks from start to finish”, with “a visual browser that interacts with the web through a graphical-user interface, a text-based browser for simpler reasoning-based web queries, a terminal, and direct API access”, and that “Starting today, Pro, Plus, and Team users can activate ChatGPT’s new agentic capabilities”. The person still starts every task, which is what keeps this at level 5 rather than level 7.

Jul 9, 2025

Grok 4 HeavyLevel 6

SpaceXAI

xAI's announcement says: "We have made further progress on parallel test-time compute, which allows Grok to consider multiple hypotheses at once. We call this model Grok 4 Heavy", sold through "a new SuperGrok Heavy tier". Several runs on one question, bought as a tier.

Jun 13, 2025

How we built our multi-agent research systemLevel 6

Anthropic

Anthropic's engineering post says of a feature customers were already using: "Our Research system uses a multi-agent architecture with an orchestrator-worker pattern, where a lead agent coordinates the process while delegating to specialized subagents that operate in parallel." The subagents work "with their own context windows" before condensing what they found for the lead agent.

Apr 15, 2025

Claude ResearchLevel 5

Anthropic

Anthropic's announcement says Claude "operates agentically, conducting multiple searches that build on each other while determining exactly what to investigate next", across "both your internal work context and the web". It launched "in early beta for Max, Team, and Enterprise plans in the United States, Japan, and Brazil."

Apr 15, 2025

Claude Research reaches customersLevel 5

Anthropic

Anthropic's launch post of April 15, 2025 says Claude “operates agentically, conducting multiple searches that build on each other while determining exactly what to investigate next”, and that “Research is now available in early beta for Max, Team, and Enterprise plans in the United States, Japan, and Brazil.” Anthropic's engineering post two months later says of the same feature: “Our Research system uses a multi-agent architecture with an orchestrator-worker pattern, where a lead agent coordinates the process while delegating to specialized subagents that operate in parallel.” The launch post itself does not say several agents run on one request.

Apr 9, 2025

Agent2Agent ProtocolLevel 6

Google

Google's announcement says "The A2A protocol will allow AI agents to communicate with each other, securely exchange information, and coordinate actions on top of various enterprise platforms or applications", and says more than 50 technology partners contributed to it.

Mar 12, 2025

Gemini RoboticsLevel 7

Google DeepMind

Google DeepMind's announcement introduces "Gemini Robotics, an advanced vision-language-action (VLA) model that was built on Gemini 2.0 with the addition of physical actions as a new output modality for the purpose of directly controlling robots", alongside Gemini Robotics-ER for spatial reasoning. The registry's Gemini Robotics entry is the later Gemini Robotics 2.

Mar 11, 2025

Responses APILevel 4

OpenAI

OpenAI's API changelog for March 11, 2025 records "the Responses API, a new API for creating and using agents and tools" and, the same day, "a set of built-in tools" for it: web search, file search and computer use. Remote MCP servers and a code interpreter were added on May 20, 2025.

Mar 6, 2025

ManusLevel 7

Manus

Manus describes itself, on its own site, as building "general AI agents as the Action Engine for life" and as "building the hands for AI to do" rather than the reasoning underneath: an agent a person points at a goal and lets run.

Not verified today: Manus states no launch date on any page of its own, confirmed again on September 18, 2026, and neither does the Internet Archive's earliest capture of manus.im (March 11, 2025), which does show the product live and saying “Manus is a general AI agent that bridges minds and actions: it doesn’t just think, it delivers results.” March 6, 2025 is the date consistently reported at the time; this site could not confirm it, so the date stays unverified and Manus is not used as a marked date.
Feb 24, 2025

Claude CodeLevel 5

Anthropic

Anthropic's announcement says Claude Code, released "as a limited research preview" alongside Claude 3.7 Sonnet, "enables developers to delegate substantial engineering tasks to Claude directly from their terminal", describing it as "an active collaborator that can search and read code, edit files, write and run tests, commit and push code to GitHub, and use command line tools".

Feb 2, 2025

ChatGPT deep researchLevel 5

OpenAI

OpenAI's post of February 2, 2025 describes deep research as “An agent that uses reasoning to synthesize large amounts of online information and complete multi-step research tasks for you”, which “conducts multi-step research on the internet for complex tasks” and returns a report with citations. The same page says: “Available to Pro users today, Plus and Team next.”

Jan 23, 2025

OperatorLevel 5

OpenAI

OpenAI's post introduces "A research preview of an agent that can use its own browser to perform tasks for you." The person describes a task; the model looks at screenshots of a browser running in the cloud and clicks and types until the job is done. It hands back when it should: "Operator is trained to proactively ask the user to take over for tasks that require login, payment details, or when solving CAPTCHAs." The post says "Available to Pro users in the U.S."

Jan 22, 2025

DeepSeek-R1Level 1

DeepSeek-AI

The paper reports that "the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories", and that the training brings out "advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation." A second lab, in the open, reaching what o1 had shown four months earlier.

2024

Dec 19, 2024

Building Effective AgentsLevel 3

Anthropic

Anthropic's engineering post names and diagrams five workflow patterns (prompt chaining, routing, parallelization, orchestrator-workers and evaluator-optimizer) and advises: "Start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when simpler solutions fall short."

Dec 14, 2024

interrupt: human approval in LangGraphLevel 3

LangChain

LangChain's post introduces interrupt, which will "pause execution of the graph, mark the thread you are running as interrupted, and put whatever you passed as an input to interrupt into the persistence layer." The pattern it names first: "Pause the graph before a critical step, such as an API call, to review and approve the action. If the action is rejected, you can prevent the graph from executing the step".

Dec 11, 2024

Deep Research in GeminiLevel 5

Google

Google's announcement says Deep Research creates "a multi-step research plan for you to either revise or approve", then works "browsing the web the way you do: searching, finding interesting pieces of information and then starting a new search based on what it's learned", ending in "a comprehensive report of the key findings". It launched that day on desktop and mobile web for Gemini Advanced subscribers.

Dec 10, 2024

Devin, generally availableLevel 5

Cognition

Cognition's post of December 10, 2024, headed “Devin is now generally available”, says “Today we’re making Devin generally available starting at $500 a month for engineering teams”. The nine months between this and Devin's announcement are the gap between a demonstration and something a customer could buy.

Nov 25, 2024

Model Context ProtocolLevel 4

Anthropic

Anthropic's announcement says "The Model Context Protocol is an open standard that enables developers to build secure, two-way connections between their data sources and AI-powered tools", and that it provides "a universal, open standard for connecting AI systems with data sources, replacing fragmented integrations with a single protocol."

Oct 31, 2024

π0Level 7

Physical Intelligence

Physical Intelligence's post says "We've developed a general-purpose robot foundation model that we call π0 (pi-zero)" that "spans images, text, and actions and acquires physical intelligence by training on embodied experience from robots, learning to directly output low-level motor commands", trained on open-source data plus dexterous tasks collected "across 8 distinct robots".

Oct 24, 2024

Claude analysis toolLevel 4

Anthropic

Anthropic's announcement describes a tool that lets Claude "write and run JavaScript code directly in Claude.ai" to "process data, conduct analysis, and produce real-time insights", available to "all Claude.ai users in feature preview": code execution reaching people who write none.

Oct 22, 2024

Computer useLevel 4

Anthropic

Anthropic's announcement describes Claude working a computer "by looking at a screen, moving a cursor, clicking buttons, and typing text", and says of the capability: "At this stage, it is still experimental—at times cumbersome and error-prone." It adds that "we encourage developers to begin exploration with low-risk tasks."

Oct 1, 2024

Prompt cachingLevel 2

OpenAI

OpenAI's API changelog entry for October 1, 2024 reads: "Prompt caching: Discounts and faster processing times on recently seen input tokens." Anthropic's own API release notes record prompt caching leaving beta on the Claude API on December 17, 2024. Reusing a long shared prefix is what makes a large fixed context affordable to send on every request.

Sep 12, 2024

OpenAI o1, the first reasoning modelLevel 1

OpenAI

OpenAI's post of September 12, 2024 opens: "We are introducing OpenAI o1, a new large language model trained with reinforcement learning to perform complex reasoning. o1 thinks before it answers—it can produce a long internal chain of thought before responding to the user." It adds that the model "learns to recognize and correct its mistakes" and that its performance improves "with more time spent thinking (test-time compute)". The extra work happens inside one request and response, which is why this sits at level 1 and not higher. OpenAI's API changelog for the same day records the release of o1-preview and o1-mini.

Sep 5, 2024

Replit AgentLevel 5

Replit

Replit's post says “Last week, we launched Replit Agent, our AI system that can create and deploy applications”, and that “It configures your development environment, installs dependencies, and executes code”: the model choosing each next step and stopping when the app runs. Replit adds: “The agent is available today in early access to all Replit Core subscribers” and “it should be treated as ‘alpha’ software.” A subscription, not an invitation.

Jun 18, 2024

GensparkLevel 6

Genspark

A search product that builds a page of results for each question. Its launch post, dated Jun 18, 2024, describes a "multi-agent framework" and a "team of specialized AI agents". It is the earliest product this site found whose own launch page says several agents share one job, ten months before Claude Research. The post never says how the agents divide the work.

May 13, 2024

GPT-4oLevel 1

OpenAI

OpenAI's API changelog entry for May 13, 2024 reads: "Released GPT-4o in the API. GPT-4o is our fastest and most affordable flagship model." The changelog does not describe how its modalities are combined; this site cites it only for the release date and OpenAI's own description.

Mar 12, 2024

DevinLevel 5

Cognition

Cognition's announcement, headed "Introducing Devin, the first AI software engineer", says it equipped Devin "with common developer tools including the shell, code editor, and browser within a sandboxed compute environment—everything a human would need to do their work." On this date it was not something a customer could buy: "Devin is currently in early access as we ramp up capacity."

Feb 15, 2024

Gemini 1.5 ProLevel 2

Google

Google's announcement says "We can now run up to 1 million tokens in production", and that at this date "a limited group of developers and enterprise customers can try it with a context window of up to 1 million tokens". A private preview, not a general release, and the standard window was 128,000 tokens.

Feb 13, 2024

GraphRAGLevel 2

Microsoft Research

Microsoft Research's post describes an approach in which "The LLM processes the entire private dataset, creating references to all entities and relationships within the source data, which are then used to create an LLM-generated knowledge graph", which is then clustered and pre-summarized so questions can be answered across many documents rather than from the nearest few passages.

Feb 13, 2024

ChatGPT memoryLevel 2

OpenAI

OpenAI's post of February 13, 2024 says “We’re testing the ability for ChatGPT to remember things you discuss to make future chats more helpful”, that a user “can explicitly tell it to remember something, ask it what it remembers, and tell it to forget conversationally or through settings”, and that memory can be turned off entirely.

Jan 17, 2024

LangGraphLevel 3

LangChain

LangChain's launch post says "LangGraph is module built on top of LangChain to better enable creation of cyclical graphs, often needed for agent runtimes", and that until then "we've lacked a method for easily introducing cycles into these chains." A developer describes the application as a state machine of nodes and edges.

2023

Dec 19, 2023

AI Builder GPT Prompts in Power AutomateLevel 3

Microsoft

Microsoft's Power Platform blog says "GPT Prompts with Prompt Builder, a new feature of AI Builder, is now generally available!" and that it lets a person "add content processing and content generation capabilities to Power Automate". An ordinary customer drops a model call into a flow whose next step is still chosen by the flow, not the model.

Nov 9, 2023

Semantic RouterLevel 3

Aurelio Labs

PyPI shows version 0.0.1 of semantic-router uploaded on November 9, 2023. Its project page calls it "a superfast decision-making layer for your LLMs and agents" that routes requests "using semantic meaning" rather than waiting on a model's generation: code picks the branch.

Sep 19, 2023

Bard ExtensionsLevel 4

Google

Google's announcement says Extensions let Bard "find and show you relevant information from the Google tools you use every day — like Gmail, Docs, Drive, Google Maps, YouTube, and Google Flights and hotels — even when the information you need is across multiple apps and services." The model decides which of those to reach for mid-conversation; Google's code makes the call. The product is now Gemini's connected apps.

Aug 18, 2023

AutoGenLevel 6

Microsoft

GitHub's record of the microsoft/autogen repository gives a creation date of August 18, 2023; the first pyautogen release reached PyPI a week later, on August 25, 2023. The README calls it "a framework for creating multi-agent AI applications that can act autonomously or work alongside humans".

Jun 13, 2023

Function callingLevel 4

OpenAI

OpenAI's post of June 13, 2023 says “Developers can now describe functions to gpt-4-0613 and gpt-3.5-turbo-0613, and have the model intelligently choose to output a JSON object containing arguments to call those functions”, through “new API parameters in our /v1/chat/completions endpoint, functions and function_call”. The model chooses the call; the developer's code makes it.

May 23, 2023

Improving Factuality and Reasoning in Language Models through Multiagent DebateLevel 6

Du et al., MIT and Google Brain

The paper has "multiple language model instances propose and debate their individual responses and reasoning processes over multiple rounds to arrive at a common final answer", and reports that this "improves the factual validity of generated content, reducing fallacious answers and hallucinations that contemporary models are prone to".

May 12, 2023

ChatGPT plugins open to Plus subscribersLevel 4

OpenAI

OpenAI's ChatGPT release notes carry an entry headed “Web browsing and Plugins are now rolling out in beta (May 12)”, which says “If you are a ChatGPT Plus user, enjoy early access to experimental new features” through a beta panel “which is rolling out to all Plus users over the course of the next week”, and describes plugins as “a new version of ChatGPT that knows when and how to use third-party plugins that you enable”. No waitlist and no invitation: a subscription was enough.

May 11, 2023

100K context windowsLevel 2

Anthropic

Anthropic's announcement says: "We've expanded Claude's context window from 9K to 100K tokens, corresponding to around 75,000 words!" Enough room to paste the material in rather than retrieve from it.

May 4, 2023

the new Bing, open to everyoneLevel 2

Microsoft

Microsoft's announcement says "the new Bing is now in Open Preview and no longer has a waitlist": anyone with a Microsoft account could now ask a question and get an answer written from pages retrieved for it, with links to those pages.

Apr 9, 2023

AgentGPTLevel 5

Reworkd

A free web page, live by April 9, 2023: "Assemble, configure, and deploy autonomous AI Agents in your browser. Create an agent by adding a name / goal, and hitting deploy!" It wrote itself a task list, worked through it and added tasks from the results. Part of the AutoGPT wave of spring 2023, which put a looping agent in front of anyone with a browser, and which few people remember as a product.

Apr 7, 2023

Generative AgentsLevel 7

Park et al., Stanford University and Google Research

The paper puts "a small town of twenty five agents" in a sandbox. Its architecture, in the paper's words, is one that agents use to "store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior": agents that decide when to act, not only what to answer.

Mar 31, 2023

CAMELLevel 6

Li, Hammoud, Itani, Khizbullin and Ghanem

Two model agents, one given the role of the person with a task and one the role of the assistant, work the task out between them with no human in the conversation. The abstract proposes "a novel communicative agent framework named role-playing" and reports "comprehensive studies on instruction-following cooperation in multi-agent settings." Eight weeks before the multi-agent debate paper.

Mar 23, 2023

ChatGPT pluginsLevel 4

OpenAI

OpenAI's post of March 23, 2023 says “We’ve implemented initial support for plugins in ChatGPT”, tools that “help ChatGPT access up-to-date information, run computations, or use third-party services”, and that “we're also hosting two plugins ourselves, a web browser and code interpreter”. The model picks which plugin to call mid-conversation. This was an invitation: “Today, we will begin extending plugin alpha access to users and developers from our waitlist.”

Mar 16, 2023

Microsoft 365 Copilot unveiledLevel 3

Microsoft

Microsoft's announcement says Copilot "is more than OpenAI’s ChatGPT embedded into Microsoft 365. It’s a sophisticated processing and orchestration engine working behind the scenes to combine the power of LLMs, including GPT-4, with the Microsoft 365 apps and your business data in the Microsoft Graph". Software runs the steps (fetch the person's files and mail, build the prompt, call the model, check the result, write into Word or Outlook) and the model fills them in. Shown that day: a first draft in Word from your own files, a deck in PowerPoint from a document, trend analysis in Excel, thread summaries and draft replies in Outlook, and live meeting summaries in Teams.

Mar 16, 2023

AutoGPTLevel 5

Significant Gravitas

GitHub's record of the Significant-Gravitas/AutoGPT repository gives a creation date of March 16, 2023. The project describes itself as "The open-source platform for AI agents" and says it "lets you build, deploy, and run AI agents that carry out complete workflows": a goal-seeking loop a developer could run instead of writing one.

Mar 15, 2023

GPT-4 Technical ReportLevel 1

OpenAI

OpenAI's technical report describes "a large-scale, multimodal model which can accept image and text inputs and produce text outputs", and reports that it "exhibits human-level performance on various professional and academic benchmarks". That is OpenAI's own evaluation of its own model.

Feb 9, 2023

ToolformerLevel 4

Schick et al., Meta AI

The paper describes a model trained in a self-supervised way, from a handful of examples per API, to "decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction": the model choosing the action, not code choosing it for the model.

Feb 7, 2023

Bing Chat launches (the new Bing)Level 2

Microsoft

Microsoft's announcement says "Bing reviews results from across the web to find and summarize the answer you're looking for" and that "The new Bing also cites all its sources". On this date it was a limited preview on desktop with a waitlist, so it is not the level's availability date.

Jan 6, 2023

NeevaAILevel 2

Neeva

Neeva's post, bylined "The Neeva Team on 01/06/23", introduces a written answer at the top of its search results with citations embedded in the text, open at once to account holders in the United States. A month before the new Bing. Neeva closed its search engine later in 2023.

2022

Dec 9, 2022

OpenAI steps in ZapierLevel 3

Zapier

Zapier's own OpenAI integration page offers an action that “Sends a prompt to OpenAI and generate a response” inside a multi-step Zap, and says “Zapier lets you connect OpenAI with thousands of the most popular apps, so you can automate your work and have more time for what matters most—no code required.” The flow decides what runs next; one of its steps calls a model. The page carries no launch date, so the date marked here is the earliest capture of it this site could read: the feature was live by then, and may well have shipped earlier.

Dec 7, 2022

Perplexity AskLevel 2

Perplexity

The early Perplexity Ask generated answers from retrieved search results with citations. Cofounder Aravind Srinivas retrospectively dates its launch to December 7, 2022 at Stripe Sessions 2024.

Nov 30, 2022

ChatGPTLevel 1

OpenAI

OpenAI's announcement, dated November 30, 2022, says “We’ve trained a model called ChatGPT which interacts in a conversational way” and “During the research preview, usage of ChatGPT is free. Try it now at chat.openai.com.” A chat product anyone could open in a browser, with no API key, no code and no invitation.

Nov 2, 2022

LlamaIndexLevel 2

Jerry Liu

GitHub's record of the repository now published as run-llama/llama_index gives a creation date of November 2, 2022, and the first release of its package on PyPI, then named gpt-index, is dated November 22, 2022 with the summary “Building an index of GPT summaries.” The first widely used open library built for this level's job: index a set of documents, retrieve from it, and hand what comes back to a language model.

Oct 25, 2022

LangChainLevel 3

Harrison Chase

PyPI's release history for the langchain package shows version 0.0.1 uploaded on October 25, 2022, with 0.0.2 the next day: the library a developer could install instead of writing chaining and model-swapping plumbing themselves.

Oct 6, 2022

ReActLevel 5

Yao et al., Princeton University and Google Research

The paper interleaves reasoning traces with actions, so that "reasoning traces help the model induce, track, and update action plans as well as handle exceptions, while actions allow it to interface with external sources, such as knowledge bases or environments, to gather additional information." The model reads each result and picks the next move, which is what makes the loop the model's rather than the code's.

May 1, 2022

MRKL SystemsLevel 4

Karpas et al., AI21 Labs

The paper describes a language model surrounded by tools (a calculator, a currency converter, a database call) and "a router that routes every incoming natural language input to a module that can best respond to the input". The router is itself a small neural network. A model chooses the tool and code runs it: the whole paper is about that one idea, nine months before Toolformer.

Mar 4, 2022

InstructGPTLevel 1

OpenAI

OpenAI fine-tuned GPT-3 on human demonstrations and human rankings of its outputs. The paper reports that "outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters", in human evaluations on OpenAI's own prompt distribution.

Jan 28, 2022

Chain-of-Thought PromptingLevel 1

Wei et al., Google

The paper reports that "generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning", and that it "improves performance on a range of arithmetic, commonsense, and symbolic reasoning tasks". Nothing about the model changes; only the prompt does.

Jan 25, 2022

Embeddings in the OpenAI APILevel 2

OpenAI

OpenAI's post of January 25, 2022 says “We are introducing embeddings, a new endpoint in the OpenAI API that makes it easy to perform natural language and code tasks like semantic search, clustering, topic modeling, and classification”, and that “Embeddings that are numerically similar are also semantically similar.” Vectors as a service: a developer no longer had to train or host an embedding model.

2021

Oct 4, 2021

AI ChainsLevel 3

Wu, Terry and Cai

The paper proposes "Chaining LLM steps together, where the output of one step becomes the input for the next, thus aggregating the gains per step", and is titled for what that buys: transparent and controllable human-AI interaction, because the intermediate results exist as steps a person can see.

2020

Jun 11, 2020

The OpenAI APILevel 1

OpenAI

OpenAI's announcement of June 11, 2020 says “We’re releasing an API for accessing new AI models developed by OpenAI”, providing “a general-purpose ‘text in, text out’ interface”, and that “Today the API runs models with weights from the GPT-3 family”. Access was not open: “we are launching today in a private beta rather than general availability”, with a waitlist.

May 28, 2020

GPT-3: Language Models are Few-Shot LearnersLevel 1

OpenAI

OpenAI's paper describes "an autoregressive language model with 175 billion parameters" applied "without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction". That is the one-call pattern that prompt engineering works within.

May 22, 2020

Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksLevel 2

Lewis et al., Meta AI (FAIR), UCL, NYU

The paper introduces RAG, in which "the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia", "accessed with a pre-trained neural retriever": the model answers from passages retrieved for the question rather than only from its own weights.

2019

Dec 5, 2019

AI Dungeon 2Level 1

Nick Walton (later Latitude)

A text adventure built on GPT-2: the player types any action in plain English and the model writes what happens next. Its maker's own post is dated Thursday, December 5, 2019. One request, one response, open to anyone, three years before ChatGPT. At first it ran from a shared Google notebook, not an app.

Feb 14, 2019

Better Language Models and Their Implications (GPT-2)Level 1

OpenAI

OpenAI's post on GPT-2 reports the model doing reading comprehension, translation and summarization "without any fine-tuning of our models, simply by prompting the trained model in the right way". That is level 1 stated plainly: write the request, read the response, no training step in between.

2018

Oct 29, 2018

TransformersLevel 1

Hugging Face

GitHub's record of huggingface/transformers gives a creation date of October 29, 2018. The repository describes itself as "the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training": the open library a developer could load a pretrained language model with, years before anyone could buy a chat product.

2017

Jun 12, 2017

Attention Is All You NeedLevel 1

Vaswani et al., Google

Google researchers introduced the Transformer, an architecture built on attention alone, dispensing with the recurrence and convolutions earlier sequence models relied on. It is the architecture under the one-call language models the rest of this level names.

Feb 7, 2017

FAISSLevel 2

Facebook AI Research

GitHub's record of facebookresearch/faiss gives a creation date of February 7, 2017 and describes it as "A library for efficient similarity search and clustering of dense vectors". That is the open library a developer could build vector search over their own documents with, and it is still the substrate under most retrieval code.

2016

Mar 9, 2016

XGBoost: A Scalable Tree Boosting SystemLevel 0

Tianqi Chen and Carlos Guestrin

The paper describes a scalable tree boosting system "used widely by data scientists to achieve state-of-the-art results". Classical machine learning, no language model anywhere in it, still advancing in the middle of the deep-learning decade.

2010

Feb 2010

ElasticsearchLevel 0

Shay Banon

Elastic's own history post places Elasticsearch's first release, which it says "happened to be 0.4.0", in February 2010: keyword search an ordinary developer could run without writing a search engine.

Feb 1, 2010

scikit-learnLevel 0

Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort and Vincent Michel, INRIA

scikit-learn's own About page says these four "took leadership of the project and made the first public release, February the 1st 2010", packaging classical machine learning algorithms a developer could call without implementing them.

1966

Jan 1966

ELIZALevel 0

Joseph Weizenbaum, MIT

Weizenbaum's paper describes ELIZA, a program that scans input for keywords and applies decomposition and reassembly rules attached to them; the paper says "Keywords and their associated transformation rules constitute the SCRIPT for a particular class of conversation." Code picks the reply, with no model involved.

Method

How this page is dated and corrected

What the selected dates mean

Described: A selected foundational publication describing the level’s defining idea. This is a reference point, not a claim to the first invention or use. Paper dates use the arXiv v1 submission where applicable.

Buildable: The earliest release the site could verify of a widely used open framework, library or API that a developer could build the level with instead of writing it from scratch. A developer still has to build something; this is not a product an ordinary customer can use.

Selected milestone: A selected product or documented architecture illustrating the level. Each label identifies the event: release, preview, archived availability, or architecture documentation. These are not universal arrival dates. Documentation and archive dates must not be read as exact deployment or launch dates.

Correcting a date

Nothing on this page is final: it is what this site could verify as of 09/20/2026. A reader who finds an earlier verifiable source, or a wrong date, corrects it the way any page here is corrected. See the method page for how a level is defined and how the site changes. The full method behind this specific page's dates is in docs/TIMELINE-METHOD.md in the site's repository.