Concepts, examples, and tools for working with AI

Learn to work with AI, from a question to a workforce.

AI is changing how work gets done. Explore the main ways to use a language model, from a single chat message to teams of agents that run on their own, grouped into eight levels by how much the model decides for itself.

49 techniques34 recipes223 named models, products and toolsNames listed 09/19/2026
00Conventional software ↗Rules, search, and automation01Direct prompting ↗Ask for a response02Added context ↗Supply relevant information03Workflows ↗Software organizes the steps04Tool use ↗The model requests an action05Agent loops ↗Observe, decide, act, repeat06Teams of Agents ↗Agents coordinate work07Always-on agents ↗Resume across sessions and events
Patterns can combineHigher ≠ better

Where to start

Start with the basics

From a chatbot to an always-on agent

See how conversational AI expanded through stronger reasoning, retrieval, workflows, tools, agent loops, and systems that keep working across sessions.

See the technology evolve
  1. 1 Direct prompting
  2. 2 Added context
  3. 3 Workflows
  4. 4 Tool use
  5. 5 Agent loops
  6. 6 Teams of Agents
  7. 7 Always-on agents

Overlapping developments, explained through seven levels. Teams are optional.

Explore the connections

See how the concepts fit together

Find a concept, trace its prerequisites, and see what it can lead to. The full map connects 54 concepts across the eight levels and the topics that run through them.

Explore the map

Want to see a concept in practice? Explore worked examples.

A few connections from the full map
Worked example

The same question at two levels

One task, two designs. Compare coverage, effort, and reliability; choose the design that fits the task.

The question: “How long is the warranty on the DW-480, and what voids it?”

Level 2 · RAG

Search once, then answer

  1. Question
  2. Retrieve passages
  3. Draft answer

Your code sets the search and passes the retrieved passages to the model.

In this illustration

The answer gives a two-year warranty and two exclusions. A third exclusion is missing from the retrieved excerpts.

A predictable sequence, but its answer depends on the evidence retrieved. RAG can also use richer retrieval and validation; one search is this design’s choice.

Explore the RAG guided example →
Level 5 · Agentic RAG

Search, inspect, and follow up

  1. Search
  2. Inspect evidence
  3. Follow up or answer

The model chooses follow-up searches and which documents to open, within tool permissions and run limits.

In this illustration

Opening the full exclusions section reveals a third condition: service by an unlisted technician also voids the warranty.

More chances to fill a gap, with extra calls and opportunities for error. The model can still stop too early; citations and coverage need checking.

Explore the agentic RAG guided example →

Fictional, scripted comparison—not a benchmark or a claim about every document-chat product. The linked examples explore these techniques in other scenarios.

The timeline

Selected milestones in how people use AI

Selected publications and public milestones, not dates of invention. Select a product for the event, date and evidence.

Foundational publication Product / architecture milestone
Level
20192021202320252027
7Always-on agents
5Agent loops
3Workflows
2Added context
1Direct prompting

0 · Conventional software Rules, search and conventional software have no single arrival date.

Lines connect selected dates, not development time. Sources checked through Sep 20, 2026.

Across all levels · changes in the models

Reasoning models · Sep 2024

Models trained to think before they answer. Not a new level: a better engine for every level, and the one OpenAI credits for its first agent. Source ↗

Typed decision models (Jev) · Sep 2026

TypeSafe introduced Jev in early access on September 15, 2026: a model that returns typed decisions with probabilities rather than generated text. “System One” is TypeSafe’s name for its proposed model category. Source ↗

Selected milestone

7 · Always-on agentsGrok BotPaid beta released

An example of agents that keep working beyond an interactive session. The product also describes a team of agents: one product can illustrate more than one level.

Marked event: Aug 11, 2026 · Introducing Grok Bot ↗

Foundational publication: Generative Agents ↗ · Apr 7, 2023. This is a reference point, not a claim to the first use of the idea.

6 · Teams of AgentsClaude ResearchMulti-agent architecture documented

Anthropic explicitly describes a lead agent delegating to parallel subagents on June 13. Research launched on April 15, but that announcement does not establish whether the same architecture was in use then. June is the documentation date, not an asserted deployment date.

Marked event: Jun 13, 2025 · How we built our multi-agent research system ↗

Foundational publication: CAMEL ↗ · Mar 31, 2023. This is a reference point, not a claim to the first use of the idea.

5 · Agent loopsOperatorBrowser-use agent preview · US Pro

The agent observes a browser, acts, and decides what to do next. This is a browser-use milestone; agents and autonomous loops existed earlier.

Marked event: Jan 23, 2025 · Introducing Operator ↗

Foundational publication: ReAct ↗ · Oct 6, 2022. This is a reference point, not a claim to the first use of the idea.

4 · Tool useChatGPT pluginsLimited alpha announced · waitlist

A visible example of a model choosing a tool. This dates the limited rollout, not broad availability or the first implementation of tool use.

Marked event: Mar 23, 2023 · ChatGPT plugins ↗

Foundational publication: MRKL Systems ↗ · May 1, 2022. This is a reference point, not a claim to the first use of the idea.

Later access: rollout to Plus subscribers begins, May 12, 2023.

3 · WorkflowsOpenAI steps in ZapierArchived availability evidence

A model call inside a workflow whose steps are set in advance. December 9 is an archived-page capture, not a confirmed launch date. Microsoft 365 Copilot remains a related workplace-AI milestone, rather than the defining workflow example.

Marked event: Available by Dec 9, 2022 · OpenAI Integrations ↗

Foundational publication: AI Chains ↗ · Oct 4, 2021. This is a reference point, not a claim to the first use of the idea.

2 · Added contextPerplexitySearch-and-answer product launched

The early Perplexity Ask combined retrieved search results with a generated answer and citations. That is the specific Level 2 pattern illustrated here, not a classification of every later Perplexity feature. Cofounder Aravind Srinivas dates the launch to December 7, 2022 in a retrospective interview; this is founder testimony, not a contemporaneous launch announcement.

Marked event: Dec 7, 2022 · Aravind Srinivas on the launch date · Stripe Sessions 2024 ↗

Early behavior: Founder interview describing the first Perplexity Ask ↗.

Foundational publication: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks ↗ · May 22, 2020. This is a reference point, not a claim to the first use of the idea.

1 · Direct promptingChatGPTResearch preview released

A recognizable public example of prompting a language model. Earlier products existed; this is a selected milestone, not the invention of prompting.

Marked event: Nov 30, 2022 · ChatGPT ↗

Foundational publication: Better Language Models and Their Implications (GPT-2) ↗ · Feb 14, 2019. This is a reference point, not a claim to the first use of the idea.

Explore the full timeline and sources →

AI model capability

What the models could do over the same years

Epoch AI’s Epoch Capabilities Index, plotted by model release date, with the best score to date and this guide’s level milestones.

Epoch Capabilities Index by model release date, 2023 to 09/20/2026266 models scored by Epoch AI, plotted by release date. A stepped line follows the best score to date, from 109.94 to 166.31. Markers under the axis show selected milestones for this site's levels, and a vertical line marks the first reasoning model.4060801001201401601802023202420252026ECI scoreFirst reasoning model · Sep 2024LLaMA-13B, Meta AI · 100.2 · 02/24/2023LLaMA-33B, Meta AI · 107.14 · 02/24/2023LLaMA-65B, Meta AI · 109.94 · 02/24/2023LLaMA-7B, Meta AI · 96.25 · 02/24/2023GPT-4 (Mar 2023), OpenAI · 125.89 · 03/14/2023Cerebras-GPT-13B, Cerebras Systems · 82.57 · 03/20/2023Dolly 2.0-12b, Databricks · 89.11 · 04/11/2023vicuna-13b-v1.1 · 94.19 · 04/12/2023stablelm-tuned-alpha-7b · 54.64 · 04/19/2023Falcon-7B, Technology Innovation Institute · 94.63 · 04/24/2023RedPajama-INCITE-7B-Base · 89.87 · 05/04/2023MPT-7B, MosaicML · 94.11 · 05/05/2023PaLM 2-L · 114.88 · 05/17/2023PaLM 2-M · 108.04 · 05/17/2023PaLM 2-S · 105.88 · 05/17/2023Falcon-40B, Technology Innovation Institute · 104.13 · 05/25/2023Baichuan1-7B, Baichuan · 89.92 · 06/01/2023open_llama_7b · 91.12 · 06/07/2023GPT-3.5 Turbo (Jun 2023), OpenAI · 113.19 · 06/13/2023GPT-4 (Jun 2023), OpenAI · 123.1 · 06/13/2023MPT-30B, MosaicML · 100.28 · 06/22/2023chatglm2-6b · 98.53 · 06/24/2023XGen-7B, Salesforce · 92.9 · 06/27/2023internlm-7b · 102.53 · 07/05/2023Claude 2, Anthropic · 120.09 · 07/11/2023Llama 2-13B, Meta AI · 105.88 · 07/18/2023Llama 2-34B, Meta AI · 104.9 · 07/18/2023Llama 2-70B, Meta AI · 113.63 · 07/18/2023Llama 2-7B, Meta AI · 98.66 · 07/18/2023Stable Beluga 2, Stability AI · 117 · 07/20/2023Claude Instant, Anthropic · 120.2 · 08/09/2023Baichuan2-13B, Baichuan · 102.86 · 09/06/2023Falcon-180B, Technology Innovation Institute · 111.94 · 09/06/2023Phi-1.5, Microsoft · 90.99 · 09/11/2023internlm-20b · 111.89 · 09/18/2023Baichuan 2-7B, Baichuan · 95.88 · 09/20/2023Qwen-14B, Alibaba · 112.85 · 09/24/2023Mistral 7B v0.1, Mistral AI · 112.02 · 09/27/2023Qwen-7B, Alibaba · 106.55 · 09/28/2023DeepSeek Coder 1.3B, DeepSeek,Peking University · 62.54 · 11/02/2023DeepSeek Coder 33B, DeepSeek,Peking University · 95.87 · 11/02/2023DeepSeek Coder 6.7B, DeepSeek,Peking University · 89.05 · 11/02/2023Yi-34B, 01.AI · 117.3 · 11/02/2023GPT-3.5 Turbo (Nov 2023), OpenAI · 118.49 · 11/06/2023Claude 2.1, Anthropic · 119.22 · 11/21/2023Yi 6B, 01.AI · 104.46 · 11/22/2023DeepSeek LLM 67B, DeepSeek · 110.56 · 11/29/2023Qwen-1_8B · 92.43 · 11/30/2023Mixtral 8x7B, Mistral AI · 118.4 · 12/11/2023Phi-2, Microsoft · 107.69 · 12/12/2023Gemini 1.0 Pro, Google DeepMind · 116.97 · 12/13/2023GPT-3.5 Turbo (Jan 2024), OpenAI · 115.61 · 01/25/2024GPT-4 Turbo (Nov 2023), OpenAI · 126.45 · 01/25/2024StarCoder 2 15B, Hugging Face,ServiceNow,NVIDIA,BigCode · 104.7 · 02/20/2024StarCoder 2 7B, Hugging Face,ServiceNow,NVIDIA,BigCode · 93.06 · 02/20/2024Gemma 2B, Google DeepMind · 93.72 · 02/21/2024Gemma 7B, Google DeepMind · 111.79 · 02/21/2024StarCoder 2 3B, Hugging Face,ServiceNow,NVIDIA,BigCode · 88.18 · 02/22/2024Mistral Large, Mistral AI · 122 · 02/26/2024Nemotron-4 15B, NVIDIA · 107.45 · 02/26/2024Claude 3 Opus, Anthropic · 126.91 · 02/29/2024Claude 3 Sonnet, Anthropic · 120.68 · 02/29/2024Yi-9B · 107.35 · 03/01/2024Claude 3 Haiku, Anthropic · 118.3 · 03/07/2024GPT-4 Turbo (Apr 2024), OpenAI · 127.25 · 04/09/2024CodeQwen1.5-7B · 94.38 · 04/15/2024Mixtral 8x22B, Mistral AI · 122 · 04/17/2024Llama 3-70B, Meta AI · 122.91 · 04/18/2024Llama 3-8B, Meta AI · 116.35 · 04/18/2024phi-3-medium 14B, Microsoft · 121.19 · 04/23/2024phi-3-mini 3.8B, Microsoft · 117.25 · 04/23/2024phi-3-small 7.4B, Microsoft · 121.8 · 04/23/2024DeepSeek-V2 (MoE-236B, May 2024), DeepSeek · 124.77 · 05/07/2024Falcon 2 11B, Technology Innovation Institute · 109.32 · 05/09/2024GPT-4o (May 2024), OpenAI · 128.97 · 05/13/2024Gemini 1.5 Pro (May 2024), Google DeepMind · 126.9 · 05/14/2024Gemini 1.5 Flash (May 2024), Google DeepMind · 122.57 · 05/23/2024Mistral 7B v0.3, Mistral AI · 108.78 · 05/27/2024Qwen2-72B, Alibaba · 125.27 · 06/07/2024DeepSeek-Coder-V2-Lite-Base · 108.64 · 06/13/2024Claude 3.5 Sonnet, Anthropic · 130 · 06/20/2024Gemma 2 27B, Google DeepMind · 122.05 · 06/24/2024Gemma 2 9B, Google DeepMind · 119.79 · 06/24/2024GPT-4o mini, OpenAI · 126.56 · 07/18/2024Mistral NeMo, Mistral AI · 118.64 · 07/18/2024Llama 3.1-405B, Meta AI · 128.75 · 07/23/2024Llama 3.1-70B, Meta AI · 125.91 · 07/23/2024Llama 3.1-8B, Meta AI · 116.5 · 07/23/2024Mistral Large 2 (Jul 2024), Mistral AI · 127.54 · 07/24/2024GPT-4o (Aug 2024), OpenAI · 128.77 · 08/06/2024Command R+, Cohere,Cohere Labs (formerly Cohere for AI) · 119.25 · 08/30/2024o1-mini, OpenAI · 135.82 · 09/12/2024o1-preview, OpenAI · 134.78 · 09/12/2024Qwen2.5-32B, Alibaba · 128.52 · 09/17/2024Qwen2.5-Coder (1.5B), Alibaba · 102.66 · 09/18/2024Qwen2.5-Coder (7B), Alibaba · 112.97 · 09/18/2024Qwen2.5-Coder-0.5B · 87.56 · 09/18/2024Qwen2.5-Coder-14B · 116.28 · 09/18/2024Qwen2.5-Coder-32B, Alibaba · 119.42 · 09/18/2024Qwen2.5-Coder-3B · 107.48 · 09/18/2024Qwen2.5-72B, Alibaba · 129 · 09/19/2024Qwen2.5-7B, Alibaba · 118.44 · 09/19/2024Gemini 1.5 Flash (Sep 2024), Google DeepMind · 129.36 · 09/24/2024Gemini 1.5 Pro (Sept 2024), Google DeepMind · 131.73 · 09/24/2024Llama 3.2 1B, Meta AI · 102.43 · 09/24/2024Llama 3.2 90B, Meta AI · 125.5 · 09/24/2024Ministral 3B, Mistral AI · 118.06 · 10/16/2024Claude 3.5 Haiku, Anthropic · 127.15 · 10/22/2024Claude 3.5 Sonnet (October 2024), Anthropic · 133.55 · 10/22/2024Mistral Large 2 (Nov 2024), Mistral AI · 128.52 · 11/18/2024GPT-4o (Nov 2024), OpenAI · 128.81 · 11/20/2024INTELLECT-1, Prime Intellect,Hugging Face,Arcee AI · 100.49 · 11/29/2024Amazon Nova Pro, Amazon · 123.77 · 12/03/2024Llama 3.3 70B, Meta AI · 127.32 · 12/06/2024Gemini 2.0 Flash (Dec 2024), Google DeepMind,Google · 134.71 · 12/11/2024Grok-2 (Dec 2024), xAI · 130.48 · 12/12/2024Phi-4, Microsoft Research · 130.42 · 12/12/2024o1, OpenAI · 141.86 · 12/17/2024DeepSeek-V3, DeepSeek · 132.35 · 12/26/2024DeepSeek-R1, DeepSeek · 138.97 · 01/20/2025DeepSeek-R1-Distill-Qwen-14B, DeepSeek · 135.43 · 01/20/2025DeepSeek-R1-Distill-Qwen-32B, DeepSeek · 137.42 · 01/20/2025Gemini 2.0 Flash Thinking (Jan 2025), Google DeepMind,Google · 135.37 · 01/21/2025Qwen2.5-Max, Alibaba · 132.53 · 01/25/2025Mistral Small 3, Mistral AI · 127.07 · 01/30/2025o3-mini, OpenAI · 140.35 · 01/31/2025Gemini 2.0 Flash (Feb 2025), Google DeepMind,Google · 134.69 · 02/05/2025Gemini 2.0 Pro, Google DeepMind · 135.06 · 02/05/2025Claude 3.7 Sonnet, Anthropic · 141.16 · 02/24/2025GPT-4.5, OpenAI · 136.75 · 02/27/2025QwQ-32B, Alibaba · 137.6 · 03/05/2025Gemma 3 12B, Google DeepMind · 123.46 · 03/12/2025Gemma 3 27B, Google DeepMind · 130.03 · 03/12/2025Gemma 3 4B, Google DeepMind · 115.97 · 03/12/2025Mistral Small 3.1, Mistral AI · 127.48 · 03/17/2025DeepSeek-V3 (Mar 2025), DeepSeek · 135.95 · 03/24/2025Gemini 2.5 Pro (Mar 2025), Google DeepMind · 144.16 · 03/31/2025Llama 4 Scout, Meta AI · 129.64 · 04/05/2025Llama 4 Maverick, Meta AI · 132.2 · 04/06/2025Grok 3, xAI · 138.3 · 04/09/2025GPT-4.1, OpenAI · 136.8 · 04/14/2025GPT-4.1 mini, OpenAI · 135.02 · 04/14/2025GPT-4.1 nano, OpenAI · 129.63 · 04/14/2025o3, OpenAI · 146.91 · 04/16/2025o4-mini, OpenAI · 145.65 · 04/16/2025Gemini 2.5 Flash (Apr 2025), Google DeepMind · 139.97 · 04/17/2025Qwen3-235B-A22B, Alibaba · 139.35 · 04/28/2025Qwen3-8B, Alibaba · 136.18 · 04/28/2025Qwen3-14B, Alibaba · 138.24 · 04/29/2025Qwen3-30B-A3B, Alibaba · 136.19 · 04/29/2025Qwen3-32B, Alibaba · 138.51 · 04/29/2025Gemini 2.5 Pro (May 2025), Google DeepMind · 142.47 · 05/06/2025Mistral Medium 3, Mistral AI · 134.07 · 05/07/2025Gemini 2.5 Flash (May 2025), Google DeepMind · 141.54 · 05/20/2025Claude Opus 4, Anthropic · 142.68 · 05/22/2025Claude Sonnet 4, Anthropic · 141.69 · 05/22/2025DeepSeek-R1 (May 2025), DeepSeek · 141.29 · 05/28/2025Gemini 2.5 Pro (Jun 2025), Google DeepMind · 145.26 · 06/05/2025Magistral Small 1.0, Mistral AI · 133.19 · 06/10/2025o3-pro, OpenAI · 147.45 · 06/10/2025Gemini 2.5 Flash (Jun 2025), Google DeepMind · 140.53 · 06/17/2025Gemini 2.5 Flash-Lite (Jun 2025), Google DeepMind · 133.93 · 06/17/2025Mistral Small 3.2, Mistral AI · 131.74 · 06/20/2025Grok-3 mini, xAI · 140.35 · 06/24/2025Grok 4, xAI · 146.45 · 07/09/2025Kimi K2 (Jul 2025), Moonshot · 140.11 · 07/12/2025Qwen3-235B-A22B-Instruct (Jul 2025), Alibaba · 138.92 · 07/25/2025Qwen3-235B-A22B-Thinking (Jul 2025), Alibaba · 143.88 · 07/25/2025Qwen3-30B-A3B-Instruct (Jul 2025), Alibaba · 137.42 · 07/29/2025Qwen3-30B-A3B-Thinking (Jul 2025), Alibaba · 139.64 · 07/30/2025Claude Opus 4.1, Anthropic · 144.11 · 08/05/2025gpt-oss-120b, OpenAI · 140.1 · 08/05/2025gpt-oss-20b, OpenAI · 137.8 · 08/05/2025GPT-5, OpenAI · 150 · 08/07/2025GPT-5 mini, OpenAI · 145.52 · 08/07/2025GPT-5 nano, OpenAI · 139.39 · 08/07/2025DeepSeek-V3.1, DeepSeek · 139.92 · 08/21/2025Magistral Small 1.2, Mistral AI · 131.41 · 09/18/2025Grok 4 Fast, xAI · 144.21 · 09/19/2025Qwen3-Max, Alibaba · 142.43 · 09/24/2025Gemini 2.5 Flash (Sep 2025), Google DeepMind · 142.99 · 09/25/2025Claude Sonnet 4.5, Anthropic · 146.84 · 09/29/2025DeepSeek-V3.2-Exp, DeepSeek · 145.02 · 09/29/2025GLM-4.6, Z.ai (Zhipu AI),Tsinghua University · 140.78 · 09/30/2025GPT-5 Pro, OpenAI · 150.29 · 10/07/2025Claude Haiku 4.5, Anthropic · 142.37 · 10/15/2025Kimi K2 Thinking, Moonshot · 145.76 · 11/06/2025GPT-5.1, OpenAI · 149.64 · 11/13/2025Gemini 3 Pro, Google DeepMind · 153 · 11/18/2025Claude Opus 4.5, Anthropic · 150.11 · 11/24/2025DeepSeek-V3.2, DeepSeek · 146.19 · 12/01/2025GPT-5.2, OpenAI · 153.51 · 12/11/2025GPT-5.2 Pro, OpenAI · 155.38 · 12/11/2025Gemini 3 Flash, Google DeepMind · 151.84 · 12/17/2025GLM-4.7, Z.ai (Zhipu AI) · 143.46 · 12/22/2025Kimi K2.5, Moonshot · 148.01 · 01/27/2026Claude Opus 4.6, Anthropic · 155.34 · 02/05/2026GPT-5.3 Codex, OpenAI · 156.58 · 02/05/2026GLM-5, Z.ai (Zhipu AI) · 145.85 · 02/11/2026MiniMax-M2.5, MiniMax · 146.51 · 02/12/2026Qwen3.5 397B-A17B, Alibaba · 146.96 · 02/13/2026Qwen 3.5 Plus (hosted 397B-A17B), Alibaba · 146.73 · 02/16/2026Claude Sonnet 4.6, Anthropic · 152.25 · 02/17/2026Grok 4.20, xAI · 152.03 · 02/17/2026Gemini 3.1 Pro, Google DeepMind · 155 · 02/19/2026Qwen3.5-35B-A3B, Alibaba · 142.53 · 02/24/2026Qwen3.5-9B, Alibaba · 139.44 · 02/24/2026Qwen 3.5 Flash (hosted 35B-A3B), Alibaba · 144 · 02/25/2026Gemini 3.1 Flash-Lite, Google · 144.49 · 03/03/2026GPT-5.4, OpenAI · 156.86 · 03/05/2026GPT-5.4 Pro, OpenAI · 158.95 · 03/05/2026GPT-5.4 Mini, OpenAI · 148.96 · 03/17/2026GPT-5.4 Nano, OpenAI · 145.85 · 03/17/2026MiniMax-M2.7, MiniMax · 145.8 · 03/18/2026Qwen 3.6 Plus, Alibaba · 147.67 · 03/31/2026Gemma 4 26B A4B, Google DeepMind · 141.86 · 04/02/2026Gemma 4 31B IT, Google DeepMind · 142.66 · 04/02/2026GLM-5.1, Z.ai (Zhipu AI) · 149.71 · 04/07/2026Muse Spark, Meta AI · 152.12 · 04/08/2026Qwen 3.6 35B-A3B, Alibaba · 143.86 · 04/14/2026Claude Opus 4.7, Anthropic · 156.34 · 04/16/2026Grok 4.3 Beta, xAI · 149.17 · 04/17/2026Kimi K2.6, Moonshot · 151 · 04/20/2026Qwen 3.6 Max (Preview), Alibaba · 149.27 · 04/20/2026Qwen3.6 27B, Alibaba · 146.49 · 04/22/2026GPT-5.5, OpenAI · 159.12 · 04/23/2026GPT-5.5 Pro, OpenAI · 162.25 · 04/23/2026DeepSeek-V4-Flash, DeepSeek · 146.11 · 04/24/2026DeepSeek-V4-Pro, DeepSeek · 149.09 · 04/24/2026Qwen 3.6 Flash, Alibaba · 143.27 · 04/27/2026Mistral Medium 3.5, Mistral AI · 141.28 · 04/28/2026GPT-5.5 Instant, OpenAI · 142.5 · 05/05/2026Gemini 3.5 Flash, Google DeepMind · 154.69 · 05/19/2026Qwen3.7-Max, Alibaba · 153.73 · 05/19/2026Claude Opus 4.8, Anthropic · 158.3 · 05/28/2026MiniMax-M3, MiniMax · 146.47 · 06/01/2026Qwen3.7-Plus, Alibaba · 147.41 · 06/02/2026Nemotron 3 Ultra, NVIDIA · 146.27 · 06/04/2026Claude Fable 5, Anthropic · 163.27 · 06/09/2026Kimi K2.7 Code, Moonshot · 150.07 · 06/12/2026GLM-5.2, Z.ai (Zhipu AI) · 151.86 · 06/16/2026Claude Sonnet 5, Anthropic · 156.22 · 06/30/2026Grok 4.5, xAI · 153.92 · 07/08/2026GPT-5.6 Luna, OpenAI · 156.31 · 07/09/2026GPT-5.6 Sol, OpenAI · 161.81 · 07/09/2026GPT-5.6 Terra, OpenAI · 159.14 · 07/09/2026Muse Spark 1.1, Meta AI · 154.62 · 07/09/2026Inkling, Thinking Machines · 148.64 · 07/15/2026Inkling-Small, Thinking Machines · 150.15 · 07/15/2026Kimi K3, Moonshot · 157.63 · 07/16/2026Gemini 3.5 Flash-Lite, Google DeepMind · 145.13 · 07/21/2026Gemini 3.6 Flash, Google DeepMind · 154.33 · 07/21/2026Claude Opus 5, Anthropic · 162.3 · 07/24/2026Qwen3.7 Flash, Alibaba · 144.63 · 07/27/2026DeepSeek V4 Flash 0731, DeepSeek · 154.49 · 07/31/2026Qwen 3.8 Max, Alibaba · 156.62 · 08/02/2026Muse Spark 1.2, Meta AI · 155.48 · 08/05/2026Grok 4.6, xAI · 156.35 · 08/12/2026DeepSeek V4 Pro 0813, DeepSeek · 155.47 · 08/13/2026Gemini 3.7 Flash, Google DeepMind · 157.44 · 08/13/2026GLM-5.3, Z.ai (Zhipu AI) · 155.25 · 08/14/2026GLM-5.3-Flash, Z.ai (Zhipu AI) · 151.45 · 08/20/2026Claude Fable 5.1, Anthropic · 164.47 · 09/01/2026Qwen3.8 Max (0902), Alibaba · 155.22 · 09/01/2026Gemini 3.8 Flash, Google DeepMind · 156.54 · 09/02/2026GPT-6 Astra, OpenAI · 166.31 · 09/03/2026New high: LLaMA-65B, Meta AI · 109.94 · 02/24/2023New high: GPT-4 (Mar 2023), OpenAI · 125.89 · 03/14/2023New high: GPT-4 Turbo (Nov 2023), OpenAI · 126.45 · 01/25/2024New high: Claude 3 Opus, Anthropic · 126.91 · 02/29/2024New high: GPT-4 Turbo (Apr 2024), OpenAI · 127.25 · 04/09/2024New high: GPT-4o (May 2024), OpenAI · 128.97 · 05/13/2024New high: Claude 3.5 Sonnet, Anthropic · 130 · 06/20/2024New high: o1-mini, OpenAI · 135.82 · 09/12/2024New high: o1, OpenAI · 141.86 · 12/17/2024New high: Gemini 2.5 Pro (Mar 2025), Google DeepMind · 144.16 · 03/31/2025New high: o3, OpenAI · 146.91 · 04/16/2025New high: o3-pro, OpenAI · 147.45 · 06/10/2025New high: GPT-5, OpenAI · 150 · 08/07/2025New high: GPT-5 Pro, OpenAI · 150.29 · 10/07/2025New high: Gemini 3 Pro, Google DeepMind · 153 · 11/18/2025New high: GPT-5.2 Pro, OpenAI · 155.38 · 12/11/2025New high: GPT-5.3 Codex, OpenAI · 156.58 · 02/05/2026New high: GPT-5.4 Pro, OpenAI · 158.95 · 03/05/2026New high: GPT-5.5 Pro, OpenAI · 162.25 · 04/23/2026New high: Claude Fable 5, Anthropic · 163.27 · 06/09/2026New high: Claude Fable 5.1, Anthropic · 164.47 · 09/01/2026New high: GPT-6 Astra, OpenAI · 166.31 · 09/03/2026LLaMA-65BClaude 3 Opuso1-miniGemini 2.5 ProGPT-5 ProFable 5GPT-6 Astra1Level 1, Direct prompting: ChatGPT, Nov 30, 20222Level 2, Added context: Perplexity, Dec 7, 20223Level 3, Workflows: OpenAI steps in Zapier, Dec 9, 20224Level 4, Tool use: ChatGPT plugins, Mar 23, 20235Level 5, Agent loops: OpenAI Operator, Jan 23, 20256Level 6, Teams of Agents: How we built our multi-agent research system, Jun 13, 20257Level 7, Always-on agents: Grok Bot, Aug 11, 2026Selected milestones by levelEpoch Capabilities Index by model release date, 2023 to 09/20/2026266 models scored by Epoch AI, plotted by release date. A stepped line follows the best score to date, from 109.94 to 166.31. Markers under the axis show selected milestones for this site's levels, and a vertical line marks the first reasoning model.4060801001201401601802023202420252026ECI scoreReasoning · Sep 2024LLaMA-13B, Meta AI · 100.2 · 02/24/2023LLaMA-33B, Meta AI · 107.14 · 02/24/2023LLaMA-65B, Meta AI · 109.94 · 02/24/2023LLaMA-7B, Meta AI · 96.25 · 02/24/2023GPT-4 (Mar 2023), OpenAI · 125.89 · 03/14/2023Cerebras-GPT-13B, Cerebras Systems · 82.57 · 03/20/2023Dolly 2.0-12b, Databricks · 89.11 · 04/11/2023vicuna-13b-v1.1 · 94.19 · 04/12/2023stablelm-tuned-alpha-7b · 54.64 · 04/19/2023Falcon-7B, Technology Innovation Institute · 94.63 · 04/24/2023RedPajama-INCITE-7B-Base · 89.87 · 05/04/2023MPT-7B, MosaicML · 94.11 · 05/05/2023PaLM 2-L · 114.88 · 05/17/2023PaLM 2-M · 108.04 · 05/17/2023PaLM 2-S · 105.88 · 05/17/2023Falcon-40B, Technology Innovation Institute · 104.13 · 05/25/2023Baichuan1-7B, Baichuan · 89.92 · 06/01/2023open_llama_7b · 91.12 · 06/07/2023GPT-3.5 Turbo (Jun 2023), OpenAI · 113.19 · 06/13/2023GPT-4 (Jun 2023), OpenAI · 123.1 · 06/13/2023MPT-30B, MosaicML · 100.28 · 06/22/2023chatglm2-6b · 98.53 · 06/24/2023XGen-7B, Salesforce · 92.9 · 06/27/2023internlm-7b · 102.53 · 07/05/2023Claude 2, Anthropic · 120.09 · 07/11/2023Llama 2-13B, Meta AI · 105.88 · 07/18/2023Llama 2-34B, Meta AI · 104.9 · 07/18/2023Llama 2-70B, Meta AI · 113.63 · 07/18/2023Llama 2-7B, Meta AI · 98.66 · 07/18/2023Stable Beluga 2, Stability AI · 117 · 07/20/2023Claude Instant, Anthropic · 120.2 · 08/09/2023Baichuan2-13B, Baichuan · 102.86 · 09/06/2023Falcon-180B, Technology Innovation Institute · 111.94 · 09/06/2023Phi-1.5, Microsoft · 90.99 · 09/11/2023internlm-20b · 111.89 · 09/18/2023Baichuan 2-7B, Baichuan · 95.88 · 09/20/2023Qwen-14B, Alibaba · 112.85 · 09/24/2023Mistral 7B v0.1, Mistral AI · 112.02 · 09/27/2023Qwen-7B, Alibaba · 106.55 · 09/28/2023DeepSeek Coder 1.3B, DeepSeek,Peking University · 62.54 · 11/02/2023DeepSeek Coder 33B, DeepSeek,Peking University · 95.87 · 11/02/2023DeepSeek Coder 6.7B, DeepSeek,Peking University · 89.05 · 11/02/2023Yi-34B, 01.AI · 117.3 · 11/02/2023GPT-3.5 Turbo (Nov 2023), OpenAI · 118.49 · 11/06/2023Claude 2.1, Anthropic · 119.22 · 11/21/2023Yi 6B, 01.AI · 104.46 · 11/22/2023DeepSeek LLM 67B, DeepSeek · 110.56 · 11/29/2023Qwen-1_8B · 92.43 · 11/30/2023Mixtral 8x7B, Mistral AI · 118.4 · 12/11/2023Phi-2, Microsoft · 107.69 · 12/12/2023Gemini 1.0 Pro, Google DeepMind · 116.97 · 12/13/2023GPT-3.5 Turbo (Jan 2024), OpenAI · 115.61 · 01/25/2024GPT-4 Turbo (Nov 2023), OpenAI · 126.45 · 01/25/2024StarCoder 2 15B, Hugging Face,ServiceNow,NVIDIA,BigCode · 104.7 · 02/20/2024StarCoder 2 7B, Hugging Face,ServiceNow,NVIDIA,BigCode · 93.06 · 02/20/2024Gemma 2B, Google DeepMind · 93.72 · 02/21/2024Gemma 7B, Google DeepMind · 111.79 · 02/21/2024StarCoder 2 3B, Hugging Face,ServiceNow,NVIDIA,BigCode · 88.18 · 02/22/2024Mistral Large, Mistral AI · 122 · 02/26/2024Nemotron-4 15B, NVIDIA · 107.45 · 02/26/2024Claude 3 Opus, Anthropic · 126.91 · 02/29/2024Claude 3 Sonnet, Anthropic · 120.68 · 02/29/2024Yi-9B · 107.35 · 03/01/2024Claude 3 Haiku, Anthropic · 118.3 · 03/07/2024GPT-4 Turbo (Apr 2024), OpenAI · 127.25 · 04/09/2024CodeQwen1.5-7B · 94.38 · 04/15/2024Mixtral 8x22B, Mistral AI · 122 · 04/17/2024Llama 3-70B, Meta AI · 122.91 · 04/18/2024Llama 3-8B, Meta AI · 116.35 · 04/18/2024phi-3-medium 14B, Microsoft · 121.19 · 04/23/2024phi-3-mini 3.8B, Microsoft · 117.25 · 04/23/2024phi-3-small 7.4B, Microsoft · 121.8 · 04/23/2024DeepSeek-V2 (MoE-236B, May 2024), DeepSeek · 124.77 · 05/07/2024Falcon 2 11B, Technology Innovation Institute · 109.32 · 05/09/2024GPT-4o (May 2024), OpenAI · 128.97 · 05/13/2024Gemini 1.5 Pro (May 2024), Google DeepMind · 126.9 · 05/14/2024Gemini 1.5 Flash (May 2024), Google DeepMind · 122.57 · 05/23/2024Mistral 7B v0.3, Mistral AI · 108.78 · 05/27/2024Qwen2-72B, Alibaba · 125.27 · 06/07/2024DeepSeek-Coder-V2-Lite-Base · 108.64 · 06/13/2024Claude 3.5 Sonnet, Anthropic · 130 · 06/20/2024Gemma 2 27B, Google DeepMind · 122.05 · 06/24/2024Gemma 2 9B, Google DeepMind · 119.79 · 06/24/2024GPT-4o mini, OpenAI · 126.56 · 07/18/2024Mistral NeMo, Mistral AI · 118.64 · 07/18/2024Llama 3.1-405B, Meta AI · 128.75 · 07/23/2024Llama 3.1-70B, Meta AI · 125.91 · 07/23/2024Llama 3.1-8B, Meta AI · 116.5 · 07/23/2024Mistral Large 2 (Jul 2024), Mistral AI · 127.54 · 07/24/2024GPT-4o (Aug 2024), OpenAI · 128.77 · 08/06/2024Command R+, Cohere,Cohere Labs (formerly Cohere for AI) · 119.25 · 08/30/2024o1-mini, OpenAI · 135.82 · 09/12/2024o1-preview, OpenAI · 134.78 · 09/12/2024Qwen2.5-32B, Alibaba · 128.52 · 09/17/2024Qwen2.5-Coder (1.5B), Alibaba · 102.66 · 09/18/2024Qwen2.5-Coder (7B), Alibaba · 112.97 · 09/18/2024Qwen2.5-Coder-0.5B · 87.56 · 09/18/2024Qwen2.5-Coder-14B · 116.28 · 09/18/2024Qwen2.5-Coder-32B, Alibaba · 119.42 · 09/18/2024Qwen2.5-Coder-3B · 107.48 · 09/18/2024Qwen2.5-72B, Alibaba · 129 · 09/19/2024Qwen2.5-7B, Alibaba · 118.44 · 09/19/2024Gemini 1.5 Flash (Sep 2024), Google DeepMind · 129.36 · 09/24/2024Gemini 1.5 Pro (Sept 2024), Google DeepMind · 131.73 · 09/24/2024Llama 3.2 1B, Meta AI · 102.43 · 09/24/2024Llama 3.2 90B, Meta AI · 125.5 · 09/24/2024Ministral 3B, Mistral AI · 118.06 · 10/16/2024Claude 3.5 Haiku, Anthropic · 127.15 · 10/22/2024Claude 3.5 Sonnet (October 2024), Anthropic · 133.55 · 10/22/2024Mistral Large 2 (Nov 2024), Mistral AI · 128.52 · 11/18/2024GPT-4o (Nov 2024), OpenAI · 128.81 · 11/20/2024INTELLECT-1, Prime Intellect,Hugging Face,Arcee AI · 100.49 · 11/29/2024Amazon Nova Pro, Amazon · 123.77 · 12/03/2024Llama 3.3 70B, Meta AI · 127.32 · 12/06/2024Gemini 2.0 Flash (Dec 2024), Google DeepMind,Google · 134.71 · 12/11/2024Grok-2 (Dec 2024), xAI · 130.48 · 12/12/2024Phi-4, Microsoft Research · 130.42 · 12/12/2024o1, OpenAI · 141.86 · 12/17/2024DeepSeek-V3, DeepSeek · 132.35 · 12/26/2024DeepSeek-R1, DeepSeek · 138.97 · 01/20/2025DeepSeek-R1-Distill-Qwen-14B, DeepSeek · 135.43 · 01/20/2025DeepSeek-R1-Distill-Qwen-32B, DeepSeek · 137.42 · 01/20/2025Gemini 2.0 Flash Thinking (Jan 2025), Google DeepMind,Google · 135.37 · 01/21/2025Qwen2.5-Max, Alibaba · 132.53 · 01/25/2025Mistral Small 3, Mistral AI · 127.07 · 01/30/2025o3-mini, OpenAI · 140.35 · 01/31/2025Gemini 2.0 Flash (Feb 2025), Google DeepMind,Google · 134.69 · 02/05/2025Gemini 2.0 Pro, Google DeepMind · 135.06 · 02/05/2025Claude 3.7 Sonnet, Anthropic · 141.16 · 02/24/2025GPT-4.5, OpenAI · 136.75 · 02/27/2025QwQ-32B, Alibaba · 137.6 · 03/05/2025Gemma 3 12B, Google DeepMind · 123.46 · 03/12/2025Gemma 3 27B, Google DeepMind · 130.03 · 03/12/2025Gemma 3 4B, Google DeepMind · 115.97 · 03/12/2025Mistral Small 3.1, Mistral AI · 127.48 · 03/17/2025DeepSeek-V3 (Mar 2025), DeepSeek · 135.95 · 03/24/2025Gemini 2.5 Pro (Mar 2025), Google DeepMind · 144.16 · 03/31/2025Llama 4 Scout, Meta AI · 129.64 · 04/05/2025Llama 4 Maverick, Meta AI · 132.2 · 04/06/2025Grok 3, xAI · 138.3 · 04/09/2025GPT-4.1, OpenAI · 136.8 · 04/14/2025GPT-4.1 mini, OpenAI · 135.02 · 04/14/2025GPT-4.1 nano, OpenAI · 129.63 · 04/14/2025o3, OpenAI · 146.91 · 04/16/2025o4-mini, OpenAI · 145.65 · 04/16/2025Gemini 2.5 Flash (Apr 2025), Google DeepMind · 139.97 · 04/17/2025Qwen3-235B-A22B, Alibaba · 139.35 · 04/28/2025Qwen3-8B, Alibaba · 136.18 · 04/28/2025Qwen3-14B, Alibaba · 138.24 · 04/29/2025Qwen3-30B-A3B, Alibaba · 136.19 · 04/29/2025Qwen3-32B, Alibaba · 138.51 · 04/29/2025Gemini 2.5 Pro (May 2025), Google DeepMind · 142.47 · 05/06/2025Mistral Medium 3, Mistral AI · 134.07 · 05/07/2025Gemini 2.5 Flash (May 2025), Google DeepMind · 141.54 · 05/20/2025Claude Opus 4, Anthropic · 142.68 · 05/22/2025Claude Sonnet 4, Anthropic · 141.69 · 05/22/2025DeepSeek-R1 (May 2025), DeepSeek · 141.29 · 05/28/2025Gemini 2.5 Pro (Jun 2025), Google DeepMind · 145.26 · 06/05/2025Magistral Small 1.0, Mistral AI · 133.19 · 06/10/2025o3-pro, OpenAI · 147.45 · 06/10/2025Gemini 2.5 Flash (Jun 2025), Google DeepMind · 140.53 · 06/17/2025Gemini 2.5 Flash-Lite (Jun 2025), Google DeepMind · 133.93 · 06/17/2025Mistral Small 3.2, Mistral AI · 131.74 · 06/20/2025Grok-3 mini, xAI · 140.35 · 06/24/2025Grok 4, xAI · 146.45 · 07/09/2025Kimi K2 (Jul 2025), Moonshot · 140.11 · 07/12/2025Qwen3-235B-A22B-Instruct (Jul 2025), Alibaba · 138.92 · 07/25/2025Qwen3-235B-A22B-Thinking (Jul 2025), Alibaba · 143.88 · 07/25/2025Qwen3-30B-A3B-Instruct (Jul 2025), Alibaba · 137.42 · 07/29/2025Qwen3-30B-A3B-Thinking (Jul 2025), Alibaba · 139.64 · 07/30/2025Claude Opus 4.1, Anthropic · 144.11 · 08/05/2025gpt-oss-120b, OpenAI · 140.1 · 08/05/2025gpt-oss-20b, OpenAI · 137.8 · 08/05/2025GPT-5, OpenAI · 150 · 08/07/2025GPT-5 mini, OpenAI · 145.52 · 08/07/2025GPT-5 nano, OpenAI · 139.39 · 08/07/2025DeepSeek-V3.1, DeepSeek · 139.92 · 08/21/2025Magistral Small 1.2, Mistral AI · 131.41 · 09/18/2025Grok 4 Fast, xAI · 144.21 · 09/19/2025Qwen3-Max, Alibaba · 142.43 · 09/24/2025Gemini 2.5 Flash (Sep 2025), Google DeepMind · 142.99 · 09/25/2025Claude Sonnet 4.5, Anthropic · 146.84 · 09/29/2025DeepSeek-V3.2-Exp, DeepSeek · 145.02 · 09/29/2025GLM-4.6, Z.ai (Zhipu AI),Tsinghua University · 140.78 · 09/30/2025GPT-5 Pro, OpenAI · 150.29 · 10/07/2025Claude Haiku 4.5, Anthropic · 142.37 · 10/15/2025Kimi K2 Thinking, Moonshot · 145.76 · 11/06/2025GPT-5.1, OpenAI · 149.64 · 11/13/2025Gemini 3 Pro, Google DeepMind · 153 · 11/18/2025Claude Opus 4.5, Anthropic · 150.11 · 11/24/2025DeepSeek-V3.2, DeepSeek · 146.19 · 12/01/2025GPT-5.2, OpenAI · 153.51 · 12/11/2025GPT-5.2 Pro, OpenAI · 155.38 · 12/11/2025Gemini 3 Flash, Google DeepMind · 151.84 · 12/17/2025GLM-4.7, Z.ai (Zhipu AI) · 143.46 · 12/22/2025Kimi K2.5, Moonshot · 148.01 · 01/27/2026Claude Opus 4.6, Anthropic · 155.34 · 02/05/2026GPT-5.3 Codex, OpenAI · 156.58 · 02/05/2026GLM-5, Z.ai (Zhipu AI) · 145.85 · 02/11/2026MiniMax-M2.5, MiniMax · 146.51 · 02/12/2026Qwen3.5 397B-A17B, Alibaba · 146.96 · 02/13/2026Qwen 3.5 Plus (hosted 397B-A17B), Alibaba · 146.73 · 02/16/2026Claude Sonnet 4.6, Anthropic · 152.25 · 02/17/2026Grok 4.20, xAI · 152.03 · 02/17/2026Gemini 3.1 Pro, Google DeepMind · 155 · 02/19/2026Qwen3.5-35B-A3B, Alibaba · 142.53 · 02/24/2026Qwen3.5-9B, Alibaba · 139.44 · 02/24/2026Qwen 3.5 Flash (hosted 35B-A3B), Alibaba · 144 · 02/25/2026Gemini 3.1 Flash-Lite, Google · 144.49 · 03/03/2026GPT-5.4, OpenAI · 156.86 · 03/05/2026GPT-5.4 Pro, OpenAI · 158.95 · 03/05/2026GPT-5.4 Mini, OpenAI · 148.96 · 03/17/2026GPT-5.4 Nano, OpenAI · 145.85 · 03/17/2026MiniMax-M2.7, MiniMax · 145.8 · 03/18/2026Qwen 3.6 Plus, Alibaba · 147.67 · 03/31/2026Gemma 4 26B A4B, Google DeepMind · 141.86 · 04/02/2026Gemma 4 31B IT, Google DeepMind · 142.66 · 04/02/2026GLM-5.1, Z.ai (Zhipu AI) · 149.71 · 04/07/2026Muse Spark, Meta AI · 152.12 · 04/08/2026Qwen 3.6 35B-A3B, Alibaba · 143.86 · 04/14/2026Claude Opus 4.7, Anthropic · 156.34 · 04/16/2026Grok 4.3 Beta, xAI · 149.17 · 04/17/2026Kimi K2.6, Moonshot · 151 · 04/20/2026Qwen 3.6 Max (Preview), Alibaba · 149.27 · 04/20/2026Qwen3.6 27B, Alibaba · 146.49 · 04/22/2026GPT-5.5, OpenAI · 159.12 · 04/23/2026GPT-5.5 Pro, OpenAI · 162.25 · 04/23/2026DeepSeek-V4-Flash, DeepSeek · 146.11 · 04/24/2026DeepSeek-V4-Pro, DeepSeek · 149.09 · 04/24/2026Qwen 3.6 Flash, Alibaba · 143.27 · 04/27/2026Mistral Medium 3.5, Mistral AI · 141.28 · 04/28/2026GPT-5.5 Instant, OpenAI · 142.5 · 05/05/2026Gemini 3.5 Flash, Google DeepMind · 154.69 · 05/19/2026Qwen3.7-Max, Alibaba · 153.73 · 05/19/2026Claude Opus 4.8, Anthropic · 158.3 · 05/28/2026MiniMax-M3, MiniMax · 146.47 · 06/01/2026Qwen3.7-Plus, Alibaba · 147.41 · 06/02/2026Nemotron 3 Ultra, NVIDIA · 146.27 · 06/04/2026Claude Fable 5, Anthropic · 163.27 · 06/09/2026Kimi K2.7 Code, Moonshot · 150.07 · 06/12/2026GLM-5.2, Z.ai (Zhipu AI) · 151.86 · 06/16/2026Claude Sonnet 5, Anthropic · 156.22 · 06/30/2026Grok 4.5, xAI · 153.92 · 07/08/2026GPT-5.6 Luna, OpenAI · 156.31 · 07/09/2026GPT-5.6 Sol, OpenAI · 161.81 · 07/09/2026GPT-5.6 Terra, OpenAI · 159.14 · 07/09/2026Muse Spark 1.1, Meta AI · 154.62 · 07/09/2026Inkling, Thinking Machines · 148.64 · 07/15/2026Inkling-Small, Thinking Machines · 150.15 · 07/15/2026Kimi K3, Moonshot · 157.63 · 07/16/2026Gemini 3.5 Flash-Lite, Google DeepMind · 145.13 · 07/21/2026Gemini 3.6 Flash, Google DeepMind · 154.33 · 07/21/2026Claude Opus 5, Anthropic · 162.3 · 07/24/2026Qwen3.7 Flash, Alibaba · 144.63 · 07/27/2026DeepSeek V4 Flash 0731, DeepSeek · 154.49 · 07/31/2026Qwen 3.8 Max, Alibaba · 156.62 · 08/02/2026Muse Spark 1.2, Meta AI · 155.48 · 08/05/2026Grok 4.6, xAI · 156.35 · 08/12/2026DeepSeek V4 Pro 0813, DeepSeek · 155.47 · 08/13/2026Gemini 3.7 Flash, Google DeepMind · 157.44 · 08/13/2026GLM-5.3, Z.ai (Zhipu AI) · 155.25 · 08/14/2026GLM-5.3-Flash, Z.ai (Zhipu AI) · 151.45 · 08/20/2026Claude Fable 5.1, Anthropic · 164.47 · 09/01/2026Qwen3.8 Max (0902), Alibaba · 155.22 · 09/01/2026Gemini 3.8 Flash, Google DeepMind · 156.54 · 09/02/2026GPT-6 Astra, OpenAI · 166.31 · 09/03/2026New high: LLaMA-65B, Meta AI · 109.94 · 02/24/2023New high: GPT-4 (Mar 2023), OpenAI · 125.89 · 03/14/2023New high: GPT-4 Turbo (Nov 2023), OpenAI · 126.45 · 01/25/2024New high: Claude 3 Opus, Anthropic · 126.91 · 02/29/2024New high: GPT-4 Turbo (Apr 2024), OpenAI · 127.25 · 04/09/2024New high: GPT-4o (May 2024), OpenAI · 128.97 · 05/13/2024New high: Claude 3.5 Sonnet, Anthropic · 130 · 06/20/2024New high: o1-mini, OpenAI · 135.82 · 09/12/2024New high: o1, OpenAI · 141.86 · 12/17/2024New high: Gemini 2.5 Pro (Mar 2025), Google DeepMind · 144.16 · 03/31/2025New high: o3, OpenAI · 146.91 · 04/16/2025New high: o3-pro, OpenAI · 147.45 · 06/10/2025New high: GPT-5, OpenAI · 150 · 08/07/2025New high: GPT-5 Pro, OpenAI · 150.29 · 10/07/2025New high: Gemini 3 Pro, Google DeepMind · 153 · 11/18/2025New high: GPT-5.2 Pro, OpenAI · 155.38 · 12/11/2025New high: GPT-5.3 Codex, OpenAI · 156.58 · 02/05/2026New high: GPT-5.4 Pro, OpenAI · 158.95 · 03/05/2026New high: GPT-5.5 Pro, OpenAI · 162.25 · 04/23/2026New high: Claude Fable 5, Anthropic · 163.27 · 06/09/2026New high: Claude Fable 5.1, Anthropic · 164.47 · 09/01/2026New high: GPT-6 Astra, OpenAI · 166.31 · 09/03/2026LLaMA-65BClaude 3.5 SonnetFable 5GPT-6 Astra1Level 1, Direct prompting: ChatGPT, Nov 30, 20222Level 2, Added context: Perplexity, Dec 7, 20223Level 3, Workflows: OpenAI steps in Zapier, Dec 9, 20224Level 4, Tool use: ChatGPT plugins, Mar 23, 20235Level 5, Agent loops: OpenAI Operator, Jan 23, 20256Level 6, Teams of Agents: How we built our multi-agent research system, Jun 13, 20257Level 7, Always-on agents: Grok Bot, Aug 11, 2026Selected milestones by level
  • A model, closed weights
  • A model, open weights
  • Best score to date
  • A selected milestone for a level

Anthropic · ECI 163.27 · Highest score Jun 9, 2026–Aug 31, 2026

Epoch AI’s Capabilities Index combines scores from many different AI benchmarks into a single “general capability” scale, allowing comparisons between models even over timespans long enough for single benchmarks to reach saturation. Epoch adds that “Absolute ECI values are meaningless by themselves, but meaningful comparisons can be made between models.” So read the shape, not the numbers. Record periods are reconstructed from this snapshot’s scores and release dates, not historical leaderboard snapshots. Highest ECI does not mean best at every task.

The shape has a bend in it. From GPT-4 on Mar 14, 2023 to the day before the first reasoning model, 18 months, the best score rose 4.1 points, from 125.89 to 130. In the 24 months since, it has risen 36.3, to 166.31 (GPT-6 Astra, Sep 3, 2026). The selected milestones for levels 1 to 4 fall before the bend; the selected milestones for levels 5 to 7 fall after it. This reflects the editorial selection, not the first emergence of those patterns. Both figures are this site’s subtraction over Epoch’s scores. The index starts in early 2023, so it says nothing about what came before.

The 22 models that set a new high
ReleasedModelMakerScore
02/24/2023LLaMA-65BMeta AI109.94
03/14/2023GPT-4 (Mar 2023)OpenAI125.89
01/25/2024GPT-4 Turbo (Nov 2023)OpenAI126.45
02/29/2024Claude 3 OpusAnthropic126.91
04/09/2024GPT-4 Turbo (Apr 2024)OpenAI127.25
05/13/2024GPT-4o (May 2024)OpenAI128.97
06/20/2024Claude 3.5 SonnetAnthropic130
09/12/2024o1-miniOpenAI135.82
12/17/2024o1OpenAI141.86
03/31/2025Gemini 2.5 Pro (Mar 2025)Google DeepMind144.16
04/16/2025o3OpenAI146.91
06/10/2025o3-proOpenAI147.45
08/07/2025GPT-5OpenAI150
10/07/2025GPT-5 ProOpenAI150.29
11/18/2025Gemini 3 ProGoogle DeepMind153
12/11/2025GPT-5.2 ProOpenAI155.38
02/05/2026GPT-5.3 CodexOpenAI156.58
03/05/2026GPT-5.4 ProOpenAI158.95
04/23/2026GPT-5.5 ProOpenAI162.25
06/09/2026Claude Fable 5Anthropic163.27
09/01/2026Claude Fable 5.1Anthropic164.47
09/03/2026GPT-6 AstraOpenAI166.31

Data: Epoch Capabilities Index, Epoch AI, reused under CC BY 4.0, retrieved 09/18/2026; 266 models. Scores, release dates and makers are Epoch’s, unchanged apart from rounding. The level markers and the reasoning line are this site’s. Best score on the day of the snapshot: 166.31.

A map, not a ladder you must climb

Our eight levels organize patterns by the decisions a model can make. Higher does not mean better, and a system can combine several patterns. Start with your task, then explore the concepts it needs.

Leave with something you can use

Describe your task, prepare a brief, and give it to your own model with representative files. The model can consult this guide and recommend an approach; the site itself does not run it.