Best AI Models in 2026
Rankings, comparisons, and deep dives, updated monthly as new models ship.
August 2026 opens with Claude Opus 5 on top and two crowns changing hands. Opus 5 leads Artificial Analysis's Intelligence Index at 61 and its Agentic Index at 55.3 at $5 / $25 per 1M tokens, half the price of Claude Fable 5, and it has now taken the coding crown after finishing first on both of Arena's vote-based coding boards. The other move is at the cheap end, where DeepSeek V4-Flash 0731 arrived on July 31 as the new price-performance pick at $0.14 / $0.28. Below are the category winners for August 2026. Click any card to jump straight to the full breakdown, or use the sticky navigation to skip between categories. Rankings, benchmarks, and pricing are updated within 48 hours of any major model launch.
Aug 5 Muse Spark 1.2 and Muse Code, Intelligence Index 51 to 54 at an unchanged $1.25 / $4.25, plus a terminal coding agent New
Aug 3 Alibaba Qwen 3.8 (Qwen3.8-Max), 2.4-trillion-parameter multimodal flagship now generally available at $2 / $6 New
Jul 31 DeepSeek DeepSeek V4-Flash 0731, same price, same size, Intelligence Index 40 to 50 New
Jul 31 MiniMax MiniMax H3, 2K video with native stereo audio from $0.13 a second New
Late Jul Thinking Machines Inkling Small, a quarter the size of Inkling at nearly the same score New
Jul 30 OpenAI GPT-5.6 price cut, Luna 80% cheaper, Terra 20% cheaper, Sol unchanged
Jul 24 Anthropic Claude Opus 5, new #1 on Artificial Analysis, tops both the Intelligence Index (61) and the Agentic Index (55.3) New
Jul 24 Sakana AI Fugu-Ultra v1.1, orchestration-engine refresh with vendor-reported gains of up to 7.9 points over v1.0 at the same price New
Jul 21 Google Gemini 3.6 Flash, cheaper output, faster, same Intelligence Index
Jul 16 Moonshot AI Kimi K3, 2.8T params, the largest open-weight model ever released (weights shipped July 27) New
Jul 9 OpenAI GPT-5.6 Sol, Terra, and Luna, next-gen family live across ChatGPT, Codex, and the API
Pending Google Gemini 3.5 Pro, still unreleased, months behind schedule per Bloomberg Delayed
Best AI for Writing
The best AI for writing is Claude Fable 5, the only model in the top three of all three independent writing boards, with Claude Sonnet 5 as the free-tier value pick and GPT-5.5 as the alternative for fact-anchored business writing. Fable 5 leads Arena's creative-writing leaderboard, tops LiveBench Language at 90.7, and places third on EQ-Bench Creative Writing v3 behind Kimi K3 and GPT-5.6 Sol. No other model is top-three on more than one of them. This is a change from last month, when Claude Sonnet 5 held this slot on GDPval-AA; the preference boards do not support that placing, so Sonnet 5 is now the value pick rather than the quality leader. Fable 5 costs $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits. If your writing is a work deliverable rather than prose, Claude Opus 5 tops Artificial Analysis's GDPval-AA v2 professional-deliverables board outright at 1858, well clear of Fable 5's 1746.
| Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|---|---|---|---|
| Claude Fable 5 | Best writing overall | #1 Arena text overall (1508.6), #1 LiveBench Language (90.7), #3 EQ-Bench | Priciest option here | $10 / $50 |
| Kimi K3 | Creative fiction and voice | #1 EQ-Bench Creative Writing (2377), 234 Elo clear of second | Only #10 on Arena creative writing | $3 / $15 |
| Claude Sonnet 5 | Free everyday writing | Free and default on claude.ai, 1M context | #53 Arena creative writing; 75.0 LiveBench Language | $2 / $10 intro (then $3 / $15) |
| Claude Opus 5 | Professional deliverables | #1 GDPval-AA v2 at 1858, ahead of Fable 5 (1746) | Behind Fable 5 on Arena text and Humanity's Last Exam | $5 / $25 |
| GPT-5.5 | Fact-anchored business writing | Documented factual-reliability gains over GPT-5.4 | Reasoning tiers now marked deprecated | $5 / $30 |
| Gemini 3.6 Flash | Bulk drafts at scale | 17% fewer output tokens than 3.5 Flash | Weaker on hardest reasoning | $1.50 / $7.50 |
The writing crown stays with Claude Fable 5, and it is the strongest-supported call on the page. Fable 5 is now #1 on Arena's text leaderboard at 1508.6 on the August 1 cutoff, #1 on Humanity's Last Exam at 53.3% and #1 on AA-Omniscience at 40, which is three different houses agreeing. Claude Opus 5 keeps the professional-deliverables lane on GDPval-AA v2, where the re-fitted board now reads 1858 against Fable 5's 1746.
Best AI for Chat & Daily Assistant
The best AI for everyday chat is GPT-5.6, and the honest reason is reach rather than board position. It is the model ChatGPT serves by default to the largest user base in the category, which makes it the best assistant most people can actually open. On raw human preference it is not the leader: GPT-5.6 Sol sits #14 on Arena's text leaderboard at 1482.8, where Claude Fable 5 leads at 1508.6. Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 and, since the July 30, 2026 price cut, costs 60% less than it. It is available inside ChatGPT (free with limits, Plus at $20/month, Pro at $100/month), through the API (Luna $0.20 / $1.20, Terra $2 / $12, Sol $5 / $30 per 1M tokens), and bundled inside Fello AI alongside Claude, Gemini, Grok, and DeepSeek. One caveat: OpenAI's system card and the evaluator METR flagged elevated "scheming" behaviour in Sol, so GPT-5.5 Instant stays the safer pick for hallucination-sensitive work.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| GPT-5.6 | Everyday chat, ChatGPT's default | The assistant most people can open; Terra matches GPT-5.5 at ~half cost | #14 on Arena text; scheming flagged by METR | Free / $20/mo Plus; API $0.20 / $1.20 to $5 / $30 |
| Claude Fable 5 | Highest-rated conversation | #1 on Arena text overall (1508.6) and 6 of 7 subcategories | No free tier; usage-credit access on Pro | $10 / $50 API |
| Claude Opus 5 | Thoughtful, nuanced answers | #1 Artificial Analysis Intelligence Index (61) and Agentic Index (55.3) | #6 on Arena text, behind Claude Fable 5 | $20/mo Pro, $5 / $25 API |
| GPT-5.5 Instant | Hallucination-sensitive daily work | 52.5% fewer hallucinated claims vs 5.3 Instant | Reasoning tiers now marked deprecated | $20/mo Plus; API $5 / $30 |
| Gemini 3.6 Flash | Fast, free, multimodal | Free in the Gemini app, 1M context, #12 on Arena text | Weaker on hardest reasoning | Free / $1.50 / $7.50 API |
| Fello AI | All the top models, one app | ChatGPT + Claude + Gemini + Grok + DeepSeek and more on Mac, iPhone and iPad | Routed via app, not direct | $9.99/mo |
GPT-5.6 keeps the chat pick on reach, and the gap to the preference leader widened rather than closed. On the August 1 cutoff Sol sits #14 on Arena text at 1482.8, down from #11, while Claude Fable 5 leads at 1508.6. Nothing about the product changed, so the crown does not move: this is still about which assistant the most people can actually open.
Best AI for Images
The best AI for image generation is ChatGPT Images 2.0, and it is the least controversial crown on this page. GPT Image 2 leads Arena's text-to-image board at 1385 Elo and its image-editing board at 1463, and Artificial Analysis puts it first on its own image arena at 1339.4. It is the natural pick whenever your image needs to contain readable words, in English or in another script, and it is included in ChatGPT Plus and Pro. The runner-ups have changed. Reve 2.1 (July 9) is the real #2 on text-to-image at 1302, and Reve 2.0 now sits behind it on both boards. Meta's Muse Image is #3 on Arena's text-to-image board and #2 on image editing, the strongest showing any Meta image model has managed. Google's Nano Banana Pro is no longer the runner-up overall: on text-to-image it ranks between #8 and #11, below its own cheaper sibling Nano Banana 2, though it does place higher on image editing.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| ChatGPT Images 2.0 | Images with readable text | #1 text-to-image (1385) and #1 image editing (1463) | Less photoreal than the Gemini image line | Included in ChatGPT Plus |
| Reve 2.1 | Layout, typography, native 4K | #2 text-to-image at 1302, layout-preserving editing | Smaller ecosystem | Free / from $7.99/mo |
| Muse Image | Image editing, Meta ecosystem | #3 text-to-image, #2 image editing (1407) | New, thin tooling around it | Meta AI app |
| Nano Banana 2 (Gemini 3.1 Flash Image) | Photoreal portraits and products | Outranks Nano Banana Pro on both boards | Weaker on text in image | Gemini app / AI Studio |
| Seedream 5.0 Pro | Multilingual text + region editing | 10+ languages incl. Arabic RTL, lasso and layer editing | No independent benchmarks; copyright cloud | BytePlus / Magnific |
| Midjourney v8 | Stylized art, illustration | Aesthetic baseline most artists prefer | Weaker on text in image | $10-$120/mo |
| Grok Imagine | NSFW / Spicy Mode | Most permissive guardrails | Smaller model behind it | $30/mo SuperGrok |
The image crown is unchanged and GPT Image 2 still leads all three boards we track, on Arena text-to-image (1385), Arena image editing (1463) and Artificial Analysis's image arena (1339.4). The one addition is Microsoft's MAI-Image-2.5, which sits third on Artificial Analysis's image board at 1269.7 and third on Arena's image-editing board, so third place now depends on which house you read.
Best AI for Video
The best AI for video generation is Gemini Omni Flash, which leads Arena's text-to-video board at 1527 Elo, a full 45 points clear of second place, and also tops Artificial Analysis's video arena. It is #1 on both houses, which no other video model manages. Pricing runs $1.50 in and $17.50 per 1M video output tokens, which works out at roughly $0.10 per second of finished video, and it supports conversational editing. The one real limit is length: Omni Flash generates 10-second clips. If you need longer takes, Veo 3.1 remains the right tool inside the Gemini app, AI Studio and Vertex AI, with native audio and 1080p output. This replaces Veo 3.1 at the top of the category; Veo 3.1 is a good model but not the leading one, with its best variant at #6 on Arena's text-to-video board. The contender to watch is MiniMax H3 (July 31), which generates 2K clips of 4 to 15 seconds with native stereo audio and is billed per second from $0.13 at 2K, already #2 on Artificial Analysis's video board at 1241.5. We are not moving the crown on that: the score is one day old, comes from the single board that has not reproduced between parses, and Arena's text-to-video vote cutoff predates H3.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| Gemini Omni Flash | Best AI video overall | #1 on both video boards (1527 Arena), conversational editing | Caps at 10-second generations | ~$0.10/sec; Gemini app / AI Studio |
| MiniMax H3 | 2K clips with native audio | 4-15s at 2K, native stereo audio; #2 on AA video (1241.5) | One day old, no Arena votes yet; weights not published | $0.13/sec at 2K |
| Dreamina Seedance 2.0 | Closest challenger | #2 text-to-video (1482) and #1 on image-to-video | ByteDance ecosystem, limited Western access | Dreamina / BytePlus |
| Muse Video | Meta ecosystem video | #3 text-to-video at 1459 | Newest of the group, thin tooling | Meta AI app |
| Veo 3.1 | Longer production clips | Native audio, 1080p, strong physics consistency | #6 on Arena video, not the quality leader | Google AI Pro / Ultra |
| Kling 3.0 / 3.0 Turbo | Fast iteration at lower cost | Native 4K, 60fps, 15-second clips; Turbo shipped June 17 | Outside the top 16 on Arena text-to-video | From $10/mo |
| Luma Ray 3 | Photoreal scenes | Strong realism for landscapes | Smaller community | Free / from $9.99/mo |
Gemini Omni Flash keeps the video crown, and it is still #1 on both boards, at 1527.5 on Arena and 1244.8 on Artificial Analysis. MiniMax H3 (July 31) enters as the contender at #2 on Artificial Analysis's video board (1241.5), 3.3 Elo behind, and it beats Omni Flash on specification with 2K output, clips up to 15 seconds and native stereo audio. We are holding the crown because that margin is one day old, single-board, and carries no votes on Arena, whose text-to-video cutoff predates the launch.
Best AI for Coding
The best AI for coding is Claude Opus 5, and it wins on the two boards where developers vote on the finished result rather than a script measuring a harness. It is #1 on Arena's WebDev board at 1,702.9 and #1 on image-to-WebDev at 1,668.6, on vote cutoffs of August 1 and July 31. Price is the second half of the argument: Opus 5 runs $5 / $25 per 1M tokens against Claude Fable 5's $10 / $50, and Artificial Analysis measures it at $2.34 per index task against Fable 5's $3.15. Anthropic's own docs tell developers to start with Opus 5 for complex agentic coding and reserve Fable 5 for workloads that need the highest available capability. Claude Fable 5 is the runner-up and stays the pick for the hardest long-horizon work, at #4 on WebDev (1,630.7) and #2 on image-to-WebDev (1,625.7). The contender is Kimi K3, which takes #2 on WebDev at 1,675.5 and #3 on Arena's Agent board, the strongest open-weight coder on the boards, though self-hosting it means 1.56 TB of weights. On Artificial Analysis's Coding Index it is not a Claude sweep: GPT-5.6 Sol (xhigh) leads at 78.3 with Opus 5 (max) at 78.0, close enough to call a tie. The cheapest serious contender is now DeepSeek V4-Flash 0731 at $0.14 / $0.28.
| Model | Best For | Strength | Weakness | Price (per 1M tokens) |
|---|---|---|---|---|
| Claude Opus 5 | Best coding overall | #1 Arena WebDev (1,702.9) and #1 image-to-WebDev (1,668.6); $2.34 per index task | Artificial Analysis Coding Index has GPT-5.6 Sol a shade ahead | $5 / $25 |
| Claude Fable 5 | Hardest long-horizon agentic work | #2 Arena image-to-WebDev (1,625.7), #4 WebDev; 1M context | Priciest; Artificial Analysis Coding Index puts it 7th | $10 / $50 |
| Kimi K3 | Web app building and agents | #2 Arena WebDev (1,675.5), #3 Arena Agent, highest open Intelligence Index (57) | 1.56 TB to self-host | $3 / $15 |
| GPT-5.6 Sol | OpenAI flagship, agentic coding | #1 Artificial Analysis Coding Index at 78.3 (xhigh) | Absent from the official Terminal-Bench board; eval-gaming flagged by METR | $5 / $30 |
| Muse Spark 1.2 | Cheap agentic coding | Intelligence Index 54 at $1.25 / $4.25; 80% on Terminal-Bench 2.1 (Artificial Analysis) | Second to Opus 5 on all three of Meta's own coding charts (82.9% vs 86.7%); Scale SEAL has rated only 1.1 | $1.25 / $4.25 |
| Grok 4.5 | Cheap value coder | #4 on the official Terminal-Bench 2.1 board at 79.3% via Cursor CLI | Higher hallucination rate; EU API console still closed | $2 / $6 |
| Gemini 3.6 Flash | Agent coding at scale | Intelligence Index 50, 17% fewer output tokens than 3.5 Flash | Weaker on hardest reasoning | $1.50 / $7.50 |
| GLM-5.2 | Best open-weight coder you can host | Intelligence Index 51, #6 Arena WebDev (1,587.1), MIT licence | Kimi K3 outscores it; self-host or provider only | Open weights (MIT) |
The coding crown moved to Claude Opus 5. Arena finished rating it, and it came first on both coding boards, WebDev at 1,702.9 and image-to-WebDev at 1,668.6, with Claude Fable 5 fourth and second. Those are vote-based boards with published cutoffs, so the crown now rests on human comparisons rather than on any one harness, and Opus 5 does it at half Fable 5's price. Fable 5 becomes the runner-up for the hardest long-horizon work, and Kimi K3 stays the contender at #2 on WebDev. Meta's Muse Spark 1.2 arrived on August 5 and does not change that, because on Meta's own charts it comes second to Opus 5 on every coding benchmark it published.
Best AI for Creativity
The best AI for unfiltered, on-trend creative work is Grok 4.5, and we want to be exact about why. This pick is about the product, not the prose quality. Grok 4.5 carries the fewest content restrictions of any frontier model and the only native real-time X integration, which makes it the one model that will engage with edgy, topical or deliberately provocative briefs that the others decline. It is the default in the Grok app for SuperGrok and X Premium+ subscribers at $30/month. It is not the best writer, and the boards are blunt about it. Grok 4.5 sits #33 on EQ-Bench Creative Writing and #41 on Arena's creative-writing leaderboard, losing on both to the older Grok 4.20-beta1. If you are picking on output quality alone, Claude Fable 5 wins outright at #1 on Arena creative writing, and Kimi K3 wins EQ-Bench. Choose Grok 4.5 for what it will let you make, not for how well it writes.
| Model | Best For | Strength | Weakness | Price |
|---|---|---|---|---|
| Grok 4.5 | Unfiltered, opinionated, on-trend | Fewest content restrictions, native real-time X grounding | #33 EQ-Bench, #41 Arena creative writing | $30/mo SuperGrok |
| Claude Fable 5 | Highest-quality creative prose | #1 Arena creative writing, #1 LiveBench Language | Cautious guardrails on edgy briefs | $10 / $50 API |
| Kimi K3 | Fiction and distinctive voice | #1 EQ-Bench Creative Writing at 2377 | Only #10 on Arena creative writing | $3 / $15 |
| Claude Opus 5 | Long-form structured creativity | Holds long threads and self-edits; #1 Intelligence Index | Most cautious of the group | $20/mo Pro; $5 / $25 API |
| Gemini 3.1 Pro | Multimodal creative | Strong text, image and video chain | Quotas inside the Gemini app | Free / $2.00-$4.00 API in |
| Grok Imagine (Spicy Mode) | NSFW / adult creative | Most permissive image generation | Niche use case | $30/mo SuperGrok |
Grok 4.5 keeps this pick, and nothing about the product moved. It is here for its permissiveness and its live X access, not for board position, and Claude Fable 5 remains the model to use when you want the better writing. One naming note for August: xAI now trades as SpaceXAI on the leaderboards, and the model is unchanged.
Best AI for Accuracy & Research
The best AI for accuracy and research is Gemini 3.1 Pro. Its strongest result is on ARC Prize's ARC-AGI-1, where it scores 98% and ties the human panel, and it does that at $0.52 per task. That combination is the argument: several models are close on capability, none matches it on cost for reliable factual work. It pairs that with native Google Search grounding, which is what you want when the answer has to be current rather than merely plausible. It also scores 94.1% on GPQA Diamond, where GPT-5.6 Sol now matches it rather than trailing it, and 44.4% on Humanity's Last Exam, and tops Scale SEAL's HLE board at 46.44. We have dropped the previous ARC-AGI-2 framing: its 77.1% is still correct, but the board has moved and that score now places it around 14th. Two honest caveats. On grounded search specifically, Arena's search leaderboard is led by Anthropic, not Google, with Gemini 3.1 Pro grounding at #7. And on novel reasoning, GPT-5.6 Sol leads ARC-AGI-2 and Claude Opus 5 leads ARC-AGI-3 at 30%, roughly 3.75x the next-best model. We did not move the crown to Opus 5 because Artificial Analysis measures its hallucination rate at 50% and places it below Fable 5 on AA-Omniscience.
| Model | Best For | Key Benchmark | Weakness | Price |
|---|---|---|---|---|
| Gemini 3.1 Pro | Cheap, reliable factual work | 98% ARC-AGI-1 (ties human panel) at $0.52/task, 94.1% GPQA (tied by Sol) | ARC-AGI-2 77.1% now ranks ~14th; #7 on Arena search | $2.00-$4.00 / $12.00-$18.00 (tiered) |
| Claude Fable 5 | Grounded search | #1 and #3 on Arena's search leaderboard, ahead of Google | No single cheap tier | $10 / $50 (Fable 5) |
| GPT-5.6 Sol | Novel reasoning | #1 ARC-AGI-2 at 93%, against a 100% human panel | Scheming flagged by METR | $5 / $30 |
| Claude Opus 5 | Hardest unseen problems | #1 ARC-AGI-3 at 30%, ~3.75x the next model (ARC Prize) | Artificial Analysis measures a 50% hallucination rate | $5 / $25 |
| Qwen 3.7 Max | Frontier accuracy at value pricing | 92.4 GPQA Diamond, 200 free requests/day | API-only, no chat front-end | $1.25 / $3.75 promo; $2.50 / $7.50 list |
| Claude Opus 4.6 | Honesty under pressure | #1 on Scale SEAL's MASK board at 96.28; Anthropic holds the top 5 | Superseded as a flagship | Legacy Anthropic model |
Gemini 3.1 Pro keeps the accuracy crown, but its headline number is no longer a lead. Its GPQA Diamond score re-fitted to 94.1% and GPT-5.6 Sol now matches it exactly, so the crown rests on the ARC-AGI-1 result at $0.52 a task plus Search grounding rather than on GPQA. This is the next crown we re-test: Claude Fable 5 tops all three boards that measure how often a model is simply wrong, at 40 on AA-Omniscience, 61% on Omniscience Accuracy and 53.3% on Humanity's Last Exam.
Best AI for Problem Solving
The best AI for hard problem solving is GPT-5.6 Sol, and it is the best-supported crown on this page. It takes #1 on LiveBench Mathematics at 96.2, #1 on LiveBench Reasoning at 91.7, and #1 on ARC-AGI-2 at 93%, the closest any model has come to the 100% human panel. Three separate houses put it first on the reasoning tasks that matter, which is more agreement than any other category on this page produces. OpenAI has still not published Sol's FrontierMath score, so the verified OpenAI mark remains GPT-5.5 Pro's 39.6% on FrontierMath Tier 4, and we will slot Sol's number in the moment it goes public. Qwen 3.7 Max is the value alternative for competition-style problems at 97.1 on the February 2026 HMMT index and 44.5 on Apex, at a fraction of the cost of ChatGPT Pro, and it now includes 200 free model requests per day. Claude Opus 5 is the alternative for long agentic reasoning chains, leading Artificial Analysis's Agentic Index at 55.3 and ARC Prize's ARC-AGI-3 at 30%, roughly 3.75x the next-best model.
| Model | Best For | Key Benchmark | Weakness | Price |
|---|---|---|---|---|
| GPT-5.6 Sol | Hardest math, science and reasoning | #1 LiveBench Mathematics (96.2), #1 LiveBench Reasoning (91.7), #1 ARC-AGI-2 (93%) | FrontierMath still unpublished; scheming flagged by METR | $100/mo ChatGPT Pro; API $5 / $30 |
| Claude Opus 5 | Long agentic reasoning chains | #1 Agentic Index (55.3), #1 ARC-AGI-3 (30%) | Second on LiveBench Reasoning; 50% hallucination rate | $5 / $25 |
| GPT-5.5 Pro | Verified FrontierMath leader | 39.6% FrontierMath Tier 4 | Superseded by Sol as flagship | $100/mo ChatGPT Pro |
| Qwen 3.7 Max | Competition math on a budget | 97.1 HMMT 2026 Feb, 44.5 Apex, 200 free requests/day | API-only | $1.25 / $3.75 promo; $2.50 / $7.50 list |
| Claude Fable 5 | Math inside a coding workflow | #1 on Arena's math subcategory (1543), 96.0 LiveBench Mathematics | Priciest option here | $10 / $50 |
| GLM-5.2 | Open-weight problem solving | Highest open Intelligence Index at 51, MIT, 1M context | Self-host or provider only | Open weights (MIT) |
GPT-5.6 Sol keeps this crown, and it picked up a second argument. Sol now also leads Artificial Analysis's Terminal-Bench v2.1 at 89.5% and its Coding Index at 78.3, and it matches Gemini 3.1 Pro on GPQA Diamond at 94.1%. Claude Opus 5 stays the agentic-reasoning alternative on the strength of ARC-AGI-3. On August 1 OpenAI named its next model family Astra and published ten new mathematical results from an internal version, though Astra is unreleased and has no date, price or availability.
Best AI Agent
The best AI agent right now is Gemini Spark for 24/7 cloud-resident work and Claude Cowork for desktop-resident work, with ChatGPT Codex as the alternative for coding agents and OpenAI Operator-class browser agents as the alternative for web tasks. AI agents are the fastest-moving category of 2026: each top vendor now ships an agent product, and the practical choice is between agents that live in the cloud (run while your laptop is closed) and agents that live on your desktop (drive your apps directly). Gemini Spark launched at Google I/O on May 19, 2026 and is the first 24/7 cloud agent. Claude Cowork launched in general availability on April 9, 2026 and runs as a desktop agent that drives your local apps. ChatGPT Codex Mobile (May 14) is the pick for coding-agent work. Read the full Gemini Spark vs Claude Cowork comparison.
| Agent | Best For | Where It Runs | Strength | Price |
|---|---|---|---|---|
| Gemini Spark | 24/7 cloud tasks, Workspace workflows | Google Cloud VM (always-on) | First true 24/7 agent, deep Workspace integration | $99.99/mo Google AI Ultra |
| Claude Cowork | Desktop, app-driving, design + code | Your Mac/Windows desktop | Drives local apps, sees your screen | $20/mo Claude Pro |
| ChatGPT Codex Mobile | Coding agent on phone | OpenAI cloud + iOS/Android | Approve diffs and redirect work from phone | Included in ChatGPT plans |
| Grok Agentic (Grok 4.5) | Real-time research, X scraping | xAI cloud | Native X integration | $30/mo SuperGrok |
| OpenAI Operator-class | Browser tasks, web forms | OpenAI cloud + your browser | Web automation | ChatGPT Pro |
No new consumer agents shipped, and Meta's Muse Code (August 5) is a terminal developer tool rather than a consumer one, so the Gemini Spark (cloud) versus Claude Cowork (desktop) choice still drives most agent decisions for individual users. The model layer underneath them moved again: Claude Opus 5 (July 24) took #1 on Artificial Analysis's Agentic Index at 55.3, ahead of GPT-5.6 Sol at 54.0 and Claude Fable 5 at 52.8, and it costs $5 / $25. Opus 5 also tops Arena's Agent board, with Kimi K3 in third. For teams building their own agents, Meta's Muse Spark 1.2 (August 5) is a cheap agent-native option at $1.25 / $4.25 whose GDPval-AA v2 agentic Elo jumped from 1371 to 1631, though Scale SEAL has rated only the older 1.1, which leads its MCP Atlas tool-use board at 88.1. GLM-5.2 (MIT) is the strongest open-weight agent model you can host on modest hardware, at #5 on Arena's Agent board.
Best AI for Students
The best AI for students is GPT-5.5 Free inside ChatGPT for general coursework and Gemini 3.6 Flash Free inside the Gemini app for STEM and multimodal study, with Qwen 3.7 Max as the API alternative for harder problem sets (200 free requests a day) and Claude Opus 5 as the alternative for essay editing. Most students don't need to pay: the free ChatGPT tier now defaults to GPT-5.6 (with GPT-5.5 still available), Gemini 3.6 Flash is in the free Gemini app and AI Studio, Claude Sonnet 5 is the new free Claude default, and DeepSeek V4 is free on DeepSeek's chat site. For step-by-step working on the hardest math, GPT-5.6 Sol is OpenAI's new flagship (its FrontierMath score is not yet published, with GPT-5.5 Pro's verified 39.6% on FrontierMath Tier 4 the current mark), though both are paid-only; Qwen 3.7 Max is the value alternative at 97.1 HMMT 2026 February with API pricing at $1.25 / $3.75 on its current 50% promo ($2.50 / $7.50 list).
| Task | Best Model | Why | Free? | Alternative |
|---|---|---|---|---|
| Essays & coursework | GPT-5.5 | Free in ChatGPT, improved factual reliability vs 5.4 | Yes | Claude Sonnet 5 (free Claude) |
| STEM problem-solving | GPT-5.6 Sol / Qwen 3.7 Max | New STEM flagship (5.5 Pro: 39.6% FrontierMath) / 97.1 HMMT 2026 Feb | Pro paid / Qwen API paid | Gemini 3.6 Flash (free) |
| Research & accuracy | Gemini 3.1 Pro | 98% ARC-AGI-1 at $0.52/task, native Google Search grounding | Yes (Gemini app) | Claude Opus 5 |
| Writing editing | Claude Sonnet 5 | Free and default on claude.ai; Claude Fable 5 is the quality leader | Yes (Claude free) | Claude Fable 5 |
| Multimodal study (PDFs, slides, images) | Gemini 3.6 Flash | 1M context, free in Gemini app | Yes | NotebookLM (Google) |
Best AI for Work & Professionals
The best AI for professional work is GPT-5.6 (ChatGPT's default since July 9) for daily knowledge work, Claude Opus 5 for coding and high-stakes writing, and Gemini Spark for 24/7 agentic workflows. Most professionals get the most out of running two paid subscriptions (ChatGPT Plus at $20/month plus Claude Pro at $20/month, total $40/month), or consolidating with Fello AI at $9.99/month for all five top models in one Mac/iOS app. For agentic work that runs while you sleep, Gemini Spark on Google AI Ultra at $99.99/month is the only true 24/7 cloud agent.
| Use Case | Best Model | Key Stat | Price | Alternative |
|---|---|---|---|---|
| Daily knowledge work | GPT-5.6 | ChatGPT's default since July 9; the assistant most people can open | $20/mo ChatGPT Plus | Claude Opus 5 |
| Coding (proprietary) | Claude Opus 5 | #1 on Arena WebDev (1,702.9) and image-to-WebDev (1,668.6); Anthropic's recommended default | $20/mo Claude Pro | Claude Fable 5 |
| Coding (cost-effective) | Qwen 3.7 Max | 80.4 SWE-Verified, 1M context | $1.25 / $3.75 promo; $2.50 / $7.50 list | DeepSeek V4-Flash 0731 |
| Research & briefings | Gemini 3.1 Pro | 98% ARC-AGI-1 at $0.52/task, Google grounding | Google AI Pro / Ultra | Claude Opus 5 |
| Hard math, physics, finance modelling | GPT-5.6 Sol | OpenAI's new STEM flagship (5.5 Pro verified at 39.6% FrontierMath) | $100/mo ChatGPT Pro | Qwen 3.7 Max |
| Always-on agent workflows | Gemini Spark | First 24/7 cloud agent | $99.99/mo Google AI Ultra | Claude Cowork |
| Live news, X-context creative | Grok 4.5 | Opus-class + native X grounding | $30/mo SuperGrok | Gemini 3.1 Pro |
| All-in-one consolidation | Fello AI | ChatGPT + Claude + Gemini + Grok + DeepSeek | $9.99/mo | Pay each vendor separately |
AI Model Pricing in August 2026
From $0 free tiers to $199.99/month Google AI Ultra. The most consequential price on this table is Claude Opus 5 at $5 / $25, because it is the #1 model on Artificial Analysis's Intelligence Index at half the cost of Claude Fable 5. Meta's Muse Spark 1.2 lists at $1.25 / $4.25, with a Contributor tier at $0.10 / $0.20 for anyone willing to let Meta train on their prompts, and Grok 4.5 lists at $2 / $6. The price-performance pick is now DeepSeek V4-Flash 0731 at $0.14 / $0.28, which matches Gemini 3.6 Flash's Intelligence Index of 50 and costs $0.03 per index task against Flash's $0.56. Among closed models GPT-5.6 Luna is the cheapest at $0.20 / $1.20 after OpenAI's July 30 cut. For a deeper breakdown see our full AI Pricing Comparison Guide.
| Model | Input (per 1M) | Output (per 1M) | Context | Free Access? |
|---|---|---|---|---|
| GPT-5.5 | $5.00 | $30.00 | 1M (400K in Codex) | ChatGPT Free; API paid |
| GPT-5.5 Pro | $30.00 | $180.00 | 1M | ChatGPT Pro from $100/mo ($200 higher-usage tier) |
| GPT-5.6 Sol | $5.00 | $30.00 | Not published | Live in ChatGPT, Codex & API (July 9) |
| GPT-5.6 Terra | $2.00 | $12.00 | Not published | Live in ChatGPT, Codex & API (July 9) |
| GPT-5.6 Luna | $0.20 | $1.20 | Not published | Live in ChatGPT, Codex & API (July 9) |
| Claude Opus 5 | $5.00 | $25.00 | 1M | Claude Pro/Max default; API paid |
| Claude Opus 4.8 | $5.00 | $25.00 | 1M | Legacy model at Anthropic; Pro/Max/API |
| Claude Fable 5 | $10.00 | $50.00 | 1M | Permanent in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits |
| Claude Sonnet 5 | $2.00 intro / $3.00 list | $10.00 intro / $15.00 list | 1M | Claude Free & Pro default; API paid |
| Claude Sonnet 4.6 | $3.00 | $15.00 | 1M | API paid (superseded by Sonnet 5) |
| Gemini 3.1 Pro | $2.00 (≤200K) / $4.00 (>200K) | $12.00 (≤200K) / $18.00 (>200K) | 1M | Limited Gemini app; API paid |
| Gemini 3.6 Flash | $1.50 | $7.50 | 1M | Gemini app/AI Studio; free API tier + paid API |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | AI Studio; free API tier + paid API |
| Qwen 3.7 Max | $1.25 promo / $2.50 list | $3.75 promo / $7.50 list | 1M | 200 free requests/day; API paid beyond that |
| MiniMax M3 | $0.30 (50% off $0.60) | $1.20 (≤512K) | 1M | Open weights; hosting costs apply |
| LongCat-2.0 | Provider-dependent | Provider-dependent | 1M | Open weights (MIT); hosting costs apply |
| NVIDIA Nemotron 3 Ultra | Provider-dependent | Provider-dependent | 1M | Open weights (OpenMDW); hosting costs apply |
| Qwen 3.5 (open-weight) | Self-host / Together | Self-host / Together | 1M | Open weights; hosting costs apply |
| Nex-N2-Pro | Self-host / providers | Self-host / providers | 1M | Open weights (Apache 2.0); hosting costs apply |
| Rio 3.5 Open 397B | Self-host / providers | Self-host / providers | 1M | Open weights (MIT); hosting costs apply |
| Grok 4.3 | $1.25 | $2.50 | 1M | Free consumer plan; API paid |
| Grok 4.5 | $2.00 | $6.00 | 500K | Grok Build / Cursor / xAI console; EU partial, API console still closed |
| Muse Spark 1.2 | $1.25 ($0.10 Contributor tier) | $4.25 ($0.20 Contributor tier) | 1M | Paid API; Muse Code, Meta Model API and OpenRouter; Contributor tier trades training rights for a ~92% discount |
| Kimi K3 | $3.00 ($0.30 cache-hit) | $15.00 | 1M | Free basic tier in the Kimi app; open weights on Hugging Face (Kimi K3 License) |
| Gemini Omni Flash (video) | $1.50 | $17.50 (video output) | 10-second clips | Gemini app / Flow; AI Studio + API |
| DeepSeek V4-Pro | $0.435 ($0.0036 cache-hit) | $0.87 | 1M | DeepSeek Chat free; API paid |
| DeepSeek V4-Flash 0731 | $0.14 | $0.28 | 1M | DeepSeek Chat free; API paid |
| Kimi K2.7 Code | Provider-dependent | Provider-dependent | 256K | Open weights; hosting costs apply |
| GLM-5.2 | Provider-dependent | Provider-dependent | 1M | Open weights; hosting costs apply |
| ERNIE 5.1 | China-region pricing | China-region pricing | 256K | Baidu free tier |
| Gemini Spark (agent) | Not API-priced | Not API-priced | 1M (Gemini base) | Google AI Ultra $99.99 or $199.99/mo |
| Fello AI (aggregator) | Routed via app | Routed via app | Model-dependent | $9.99/mo, free tier available |
The GPT-5.5 and GPT-5.5 Pro rates above are short-context prices; OpenAI no longer publishes the specific long-context figures. The GPT-5.6 tiers are billed at 2x input and 1.5x output once a prompt passes 272K input tokens, which puts long-context Terra at $4 / $18 and Luna at $0.40 / $1.80. If you want access to multiple AI models without managing separate subscriptions, Fello AI provides GPT, Claude, Gemini, Grok, Perplexity, and more in a single app for Mac, iPhone, and iPad from $9.99/month.
Best Open-Weight Models in August 2026
The best open-weight model in August 2026 is Kimi K3, and it took the lead the moment Moonshot published the weights on July 27, 2026. It holds the highest Intelligence Index of any open model at 57 on Artificial Analysis v4.1, six points clear of GLM-5.2, and on Arena it is still the highest-placed open entry, at #2 on WebDev (1,675.5) behind only Claude Opus 5 and #3 on the Agent board. One catch decides which of the two you should actually use. The K3 download is 96 safetensors shards and about 1.56 TB, which needs a multi-node GPU cluster rather than a workstation, and it ships under a custom Kimi K3 License rather than MIT. So GLM-5.2 (Z.ai, MIT) stays our practical recommendation for teams running their own weights, at Intelligence Index 51, #6 on Arena WebDev (1,587.1) and #5 on the Agent board. Third place changed hands on the last day of July: DeepSeek V4-Flash 0731 was re-post-trained and jumped from Intelligence Index 40 to 50, at $0.14 / $0.28 under MIT.
| Model | Best For | Key Benchmark | Context / License | Where To Run |
|---|---|---|---|---|
| Kimi K3 | Highest-scoring open model, agentic and web work | II 57 (v4.1), highest of any open model; #2 Arena WebDev, #3 Arena Agent; 2.8T/104B active | 1M / Kimi K3 License | Hugging Face (96 shards, 1.56 TB), Moonshot API, providers |
| GLM-5.2 | Long-horizon agentic coding, 1M context | II 51 (v4.1), highest you can host on modest hardware; 744B/40B active | 1M / MIT | Z.ai, Hugging Face, OpenRouter |
| DeepSeek V4-Flash 0731 | Best value of any model on the page | II 50 (v4.1), up from 40 on July 31; $0.03 per index task, 284B/13B active | 1M / MIT | DeepSeek API ($0.14/$0.28), local |
| LongCat-2.0 | Frontier open coder trained on Chinese chips | II 33 (v4.1); 59.5% SWE-Bench Pro (vendor), 1.6T/~48B active | 1M / MIT | Hugging Face, GitHub, OpenRouter |
| MiniMax M3 | Cheap frontier-class multimodal | II 44 (v4.1), 59% SWE-Bench Pro, multimodal | 1M / license TBD | Hugging Face, API $0.30/1M (50% off) |
| Nex-N2-Pro | Strongest open coding score | II 41 (v4.1); 80.8 SWE-Bench Verified, 397B/17B active | Qwen-based / Apache 2.0 | Hugging Face, providers, self-host |
| Kimi K2.7 Code | Strongest commercially-licensed open coder | +21.8% on Kimi Code Bench v2 vs K2.6 (vendor); 1T/32B active | 256K / Modified MIT | Hugging Face, DeepInfra, providers |
| DeepSeek V4-Pro | Agentic real-world work | II 44 (v4.1), 1.6T/49B active | 1M / MIT | DeepSeek API ($0.435/$0.87), local |
| Hy3 | Newest permissive-licence entrant | II 41 (v4.1); #22 on Arena WebDev (1,516.4) | Apache 2.0 | Hugging Face, providers, self-host |
| Inkling | Thinking Machines' first open model | II 41 (v4.1), agentic 32.3, released July 15 | Open weights | Hugging Face, providers, self-host |
| Inkling Small | Same family at a quarter the size | II 40 (v4.1), agentic 30.8; 276B/12B active, text, image and audio in | Apache 2.0 | Hugging Face (BF16 and NVFP4), providers, self-host |
| NVIDIA Nemotron 3 Ultra | NVIDIA-tuned, fully permissive license | II 38 (v4.1), 65-70.4 SWE-Bench Verified, 550B/55B active | 1M / OpenMDW | OpenRouter, Hugging Face, AWS (8x B200 self-host) |
| Qwen 3.5 (397B / 17B active) | Multimodal, fast decode | 88.4 GPQA, 91.3 AIME 2026, 83.6 LiveCodeBench v6 | 1M / open | Together, OpenRouter, local |
| Qwen3.6-35B-A3B | Efficient open agentic coder (3B active) | 86.0 GPQA Diamond, 92.7 AIME 2026, 35B/3B active | 262K (to 1M YaRN) / Apache 2.0 | Hugging Face, OpenRouter, local |
| Qwen3.6-27B | Laptop-runnable dense coder | 87.8 GPQA Diamond, dense 27B, multimodal | 256K / Apache 2.0 | Local Mac/PC, Hugging Face, OpenRouter |
| Rio 3.5 Open 397B | Qwen 3.5 fine-tune, multilingual reasoning | 70.8 Terminal-Bench 2.1 (first-party), beats Qwen 3.7 Plus on 4/5 | 397B/17B active / MIT | Hugging Face, providers, self-host |
| Llama 4 Maverick | Meta-line flagship | 17B active / 400B total params | Llama 4 license | Meta cloud, Hugging Face, local |
| NVIDIA Nemotron 3 Nano Omni | Edge / low-power | Multimodal, very small footprint | Compact / open | Local, NVIDIA tool |
Licensing matters as much as raw score here: Kimi K2.7 (Modified MIT), DeepSeek V4 (MIT), GLM-5.2 (MIT), LongCat-2.0 (MIT), Hy3 (Apache 2.0), Nex-N2-Pro (Apache 2.0) and Nemotron 3 Ultra (OpenMDW) all clearly allow commercial use, while MiniMax M3 ships under its own community license and Kimi K3 sits on its own custom licence with a revenue threshold for anyone reselling it as a service. No open model is first on any Arena board we track any more: Claude Opus 5 took the Agent board in July, the last one an open model led.
Benchmarks, Prices, and Hands-On Use
Every ranking on this page combines three inputs: public benchmarks from seven independent houses (Artificial Analysis, Arena formerly LMArena, Scale SEAL, LiveBench, EQ-Bench, ARC Prize and the official Terminal-Bench 2.1 board, covering the Intelligence and Agentic indexes, GPQA Diamond, ARC-AGI-1 through 3, Humanity's Last Exam, GDPval-AA, FrontierMath, HMMT, MCP Atlas, SWE Atlas and the Remote Labor Index), published API and subscription pricing from each vendor's official pricing page, and hands-on use by the FelloAI editorial team running real prompts across the same task on every model. We re-fetch official pricing and benchmark sources before every monthly update.
Benchmarks are weighted to the use case: SWE-bench and Terminal-Bench drive coding, GPQA Diamond and ARC-AGI-2 drive accuracy, GDPval-AA (Artificial Analysis's professional-deliverables benchmark) informs professional-task quality while writing style is judged primarily by hands-on testing, FrontierMath and HMMT drive problem-solving. We disclose when a benchmark is vendor-reported but not independently verified, and we strip any claim we cannot reproduce against a live source. We do not move a category crown on the strength of a single leaderboard, and where a new model is too recent for the human-preference boards to have rated it, we say so rather than crowning it on benchmark scores alone. When a model goes through a major upgrade between updates, we re-rank the category and add a "What changed this month" line at the bottom of the deep-dive.





