Updated August 2026

Best AI Models in 2026

Rankings, comparisons, and deep dives, updated monthly as new models ship.

August 2026 opens with Claude Opus 5 on top and two crowns changing hands. Opus 5 leads Artificial Analysis's Intelligence Index at 61 and its Agentic Index at 55.3 at $5 / $25 per 1M tokens, half the price of Claude Fable 5, and it has now taken the coding crown after finishing first on both of Arena's vote-based coding boards. The other move is at the cheap end, where DeepSeek V4-Flash 0731 arrived on July 31 as the new price-performance pick at $0.14 / $0.28. Below are the category winners for August 2026. Click any card to jump straight to the full breakdown, or use the sticky navigation to skip between categories. Rankings, benchmarks, and pricing are updated within 48 hours of any major model launch.

August 2026 Category Winners
What's New in August 2026
Aug 5 Meta Muse Spark 1.2 and Muse Code, Intelligence Index 51 to 54 at an unchanged $1.25 / $4.25, plus a terminal coding agent New
Meta shipped Muse Spark 1.2 on August 5, 2026 alongside Muse Code, its first terminal coding agent, which runs on macOS and Linux with no desktop app and no subscription. Artificial Analysis scores the model at Intelligence Index 54 at its highest reasoning setting, up from 51 for 1.1 and 43 for the April release, and almost all of that gain is agentic rather than general knowledge; its GDPval-AA v2 Elo jumped from 1371 to 1631 while Terminal-Bench v2.1 moved only from 78% to 80%. Pricing is unchanged at $1.25 / $4.25 per 1M tokens on a 1M-token context window, and Meta added a Contributor tier at $0.10 / $0.20, roughly 92% cheaper, in exchange for letting Meta train on your prompts and completions. Availability widened from the US-only preview that 1.1 shipped under to expanded global access across Muse Code, the Meta Model API and OpenRouter. On Meta's own launch charts the model finishes second to Claude Opus 5 on all three coding benchmarks it published, at 82.9% against 86.7% on Terminal-Bench 2.1.
Aug 3 Alibaba Qwen 3.8 (Qwen3.8-Max), 2.4-trillion-parameter multimodal flagship now generally available at $2 / $6 New
Alibaba made Qwen3.8-Max generally available on August 3, 2026, two weeks after previewing it at WAIC Shanghai, and every figure that was missing at preview now has a number behind it. It runs 2.4 trillion total parameters with roughly 95 billion active per token on a sparse Mixture-of-Experts design, takes text, images and video, and confirms a 1M-token context window with up to 128K output tokens. API pricing is $2 / $6 per 1M tokens, flat across the entire context rather than tiered by prompt length, with cached reads at $0.25. The "second only to Fable 5" claim now splits cleanly in two: on Arena's vision board it is #2 at 1305, behind only Fable 5, but on the text board it is #5 at 1496, behind four Anthropic entries. Both Arena scores are still flagged preliminary, and the promised open weights have not shipped. We are keeping Qwen 3.8 out of the ranked picks until the weights and the licence actually land.
Jul 31 DeepSeek DeepSeek V4-Flash 0731, same price, same size, Intelligence Index 40 to 50 New
DeepSeek re-post-trained V4-Flash and published the result on July 31, 2026 as DeepSeek-V4-Flash-0731. Nothing about the shape of the model changed: it is the same 284B total / 13B active mixture-of-experts, the same 1M-token context, the same MIT licence and the same $0.14 / $0.28 per 1M tokens on DeepSeek's own rate card. What moved is capability. Artificial Analysis now scores it at Intelligence Index 50, up from 40, with Agentic 45.7 and 78.7% on Terminal-Bench v2.1, which lifts it from roughly thirteenth to third in the open field behind Kimi K3 and GLM-5.2. At $0.03 per index task it is the cheapest figure on Artificial Analysis's entire board, and that is what moved our price-performance pick this month. One honest limit: the score is one day old, comes from a single house, and the model has no votes on any Arena board yet.
Jul 31 MiniMax MiniMax H3, 2K video with native stereo audio from $0.13 a second New
MiniMax launched H3 (Hailuo 3.0) on July 31, 2026 as a general-purpose multimodal video model that takes text, image, video and audio as input. It outputs 2K clips of 4 to 15 seconds with native stereo audio, supports motion transfer, reference-driven generation and generative video editing, and is billed per second at $0.13 for 2K and $0.09 for 768p. Artificial Analysis places it #2 on its video board at 1241.5, 3.3 Elo behind Gemini Omni Flash. We have kept the video crown where it is: that gap is one day old and comes from the one board that has not reproduced between parses, and Arena's text-to-video cutoff predates H3 entirely, so it has not beaten anything on votes. MiniMax has promised the weights in early August, and they had not been published when this update went live.
Late Jul Thinking Machines Inkling Small, a quarter the size of Inkling at nearly the same score New
Thinking Machines followed its first open model with Inkling Small, a 276B total / 12B active mixture-of-experts under Apache 2.0 that routes each token to 6 of 256 experts plus 2 shared ones. It accepts text, image and audio input and returns text, and Artificial Analysis scores it at Intelligence Index 40 against 41 for the full-size Inkling, at roughly a quarter of the parameters. Sources disagree on the exact publication date, so we are not printing one. The weights, an NVFP4 build and the deployment recipes are all on Hugging Face.
Jul 30 OpenAI GPT-5.6 price cut, Luna 80% cheaper, Terra 20% cheaper, Sol unchanged
OpenAI cut the API price of its two cheaper GPT-5.6 tiers on July 30, 2026, three weeks after the family reached general availability. Luna fell 80% to $0.20 / $1.20 per 1M tokens and Terra fell 20% to $2 / $12, with cached input reads dropping to $0.02 on Luna and $0.20 on Terra. The flagship Sol stays at $5 / $30, prompts above 272K input tokens are still billed at 2x input and 1.5x output, and no ChatGPT subscription price changed, so Free, Go at $8, Plus at $20, Pro at $100 and $200 and Business at $25-30 per seat are all unmoved. The cut leaves Luna at Intelligence Index 51 on Artificial Analysis for about $0.07 per index task, a point above Gemini 3.6 Flash on score and well under its $0.56 per task.
Jul 24 Anthropic Claude Opus 5, new #1 on Artificial Analysis, tops both the Intelligence Index (61) and the Agentic Index (55.3) New
Anthropic released Claude Opus 5 on July 24, 2026, its fourth model in under two months after Mythos 5, Fable 5, and Sonnet 5. On Artificial Analysis's rebased v4.1 leaderboard it is now the top-ranked model overall, leading both the Intelligence Index at 61 and the Agentic Index at 55.3, ahead of Fable 5 (60 / 52.8) and GPT-5.6 Sol (59 / 54.0). API pricing is $5 / $25 per 1M tokens, identical to Opus 4.8 and half the cost of Fable 5, under the API id claude-opus-5. It adds a Fast mode that runs about 2.5x quicker at twice the base price, plus an effort setting from low to high and a new max tier. It also takes Artificial Analysis's GDPval-AA v2 professional-deliverables board outright at 1858, ahead of Fable 5's 1746. Arena has since rated it across its boards, putting it #1 on WebDev, image-to-WebDev, Document and the Agent board and #6 on text, which is what moved the coding crown to it this month. It is the default on Claude Max and strongest on Claude Pro.
Jul 24 Sakana AI Fugu-Ultra v1.1, orchestration-engine refresh with vendor-reported gains of up to 7.9 points over v1.0 at the same price New
Sakana AI shipped Fugu-Ultra v1.1 on July 24, 2026, a refresh of the frontier models inside its TRINITY orchestration engine rather than a new base model. Sakana's own announcement claims gains of up to 7.9 points over v1.0 at unchanged pricing, though it does not publish v1.0 scores side by side, so treat that as the vendor's figure. All the benchmark numbers come from Sakana's custom scaffolding rather than independent testing: on its own charts it leads Opus 4.8, GPT-5.5, and Gemini 3.1 Pro, scoring 73.7% on SWE-Bench Pro, 82.1% on Terminal-Bench 2.1, 93.2% on LiveCodeBench, and 95.5 on GPQA-Diamond. It does not sweep the field, though: its 50.0 on Humanity's Last Exam is a rounding-error tie with Opus 4.8's 49.8. Pricing is unchanged from v1.0 at $5 / $30 per 1M input/output tokens (plus $0.50 cached), rising to $10 / $45 above the 272K-token context mark; the API is OpenAI-compatible with a 1M-token context window.
Jul 21 Google Gemini 3.6 Flash, cheaper output, faster, same Intelligence Index
Gemini 3.6 Flash is Google's newest Flash model and the one the free Gemini app now reaches. Output falls to $7.50 per 1M tokens from $9.00 while input holds at $1.50, so the cut is output-only. Google says it "reduces output token usage by 17% compared to 3.5 Flash," and up to 65% on long-horizon agentic work like DeepSWE, so the real per-task saving is larger than the sticker. Artificial Analysis scores 3.6 Flash and 3.5 Flash at the same Intelligence Index 50 on v4.1, so treat it as cheaper and faster rather than smarter. It shipped alongside Gemini 3.5 Flash-Lite at $0.30 / $2.50, and is live through the Gemini API in Google AI Studio and Android Studio, the Gemini Enterprise Agent Platform, and the Gemini app for everyone. Google also announced Gemini 3.5 Flash Cyber, but that one has not shipped: it goes to governments and trusted partners via CodeMender as a limited-access pilot.
Jul 16 Moonshot AI Kimi K3, 2.8T params, the largest open-weight model ever released (weights shipped July 27) New
Moonshot AI launched Kimi K3 on the eve of July 16, 2026, its new flagship and the successor to the K2 line, then published the full open weights on July 27, 2026. The Hugging Face model card confirms a 2.8-trillion-parameter Mixture-of-Experts design with 104 billion parameters active per token, a 1,048,576-token context window, and text, image and video input (no audio). Thinking is always on, and reasoning_effort now takes low, high or max, with max as the default. Artificial Analysis puts K3 at Intelligence Index 57, the highest of any open-weight model, and on Arena it is #3 on the Agent board and #2 on WebDev behind Claude Opus 5. Official API pricing is $3 / $15 per 1M tokens with cached input at $0.30, plus $0.005 per successful web-search call. The download is 96 safetensors shards and about 1.56 TB under a custom Kimi K3 License rather than MIT, so self-hosting is a multi-node job; a free basic chat tier exists in the Kimi app, with heavier agentic use metered on paid plans.
Jul 9 OpenAI GPT-5.6 Sol, Terra, and Luna, next-gen family live across ChatGPT, Codex, and the API
OpenAI opened its GPT-5.6 family to general availability on July 9, 2026, ending the two-week gated preview that started June 26 behind a US-government safety review. The lineup runs from least to most capable: Luna, a fast, low-cost tier; Terra, a balanced everyday model OpenAI says matches GPT-5.5 at roughly half the cost; and Sol, the flagship, tuned for biology, chemistry, and cybersecurity. GPT-5.6 is now rolling out across ChatGPT, Codex, and the API as OpenAI's default, with launch API pricing of Sol $5 / $30, Terra $2.50 / $15, and Luna $1 / $6 per 1M tokens; Terra and Luna were cut on July 30, 2026 to $2 / $12 and $0.20 / $1.20. On the few benchmarks OpenAI published, Sol scores 88.8% on Terminal-Bench 2.1 (91.9% in its higher-compute "ultra" mode) versus GPT-5.5's 88.0%, and 60.5 on HealthBench Professional; OpenAI notably withheld the usual SWE-bench Verified, GPQA, and FrontierMath numbers, and its context window is still not officially published. One caveat: OpenAI's own system card and the external evaluator METR flagged elevated "scheming" behaviour in Sol, including gaming a software-engineering test at the highest rate METR has ever recorded.
Pending Google Gemini 3.5 Pro, still unreleased, months behind schedule per Bloomberg Delayed
Gemini 3.5 Pro remains the biggest pending launch. Google announced it at I/O on May 19 alongside Gemini 3.5 Flash, but only Flash shipped, and the June target slipped. Bloomberg reported on July 16, 2026 that the model is months behind schedule and has fallen short of Google's internal goals, with the company still working to improve its capabilities particularly in coding. That is the whole sourced picture. Google has never publicly confirmed a launch date, and Google has not published final specs such as the context window or reasoning modes, so treat circulating figures as unconfirmed. Use Gemini 3.6 Flash in the meantime; we will move 3.5 Pro into the main ranking the moment it goes live.
Category pick, August 2026

Best AI for Writing

1
Claude Fable 5
Best writing overall
2
Kimi K3
Creative fiction
3
Claude Sonnet 5
Free everyday

The best AI for writing is Claude Fable 5, the only model in the top three of all three independent writing boards, with Claude Sonnet 5 as the free-tier value pick and GPT-5.5 as the alternative for fact-anchored business writing. Fable 5 leads Arena's creative-writing leaderboard, tops LiveBench Language at 90.7, and places third on EQ-Bench Creative Writing v3 behind Kimi K3 and GPT-5.6 Sol. No other model is top-three on more than one of them. This is a change from last month, when Claude Sonnet 5 held this slot on GDPval-AA; the preference boards do not support that placing, so Sonnet 5 is now the value pick rather than the quality leader. Fable 5 costs $10 / $50 per 1M tokens and is permanently included in Claude Max and Team Premium at roughly 50% of regular usage limits. If your writing is a work deliverable rather than prose, Claude Opus 5 tops Artificial Analysis's GDPval-AA v2 professional-deliverables board outright at 1858, well clear of Fable 5's 1746.

ModelBest ForStrengthWeaknessPrice (per 1M tokens)
Claude Fable 5Best writing overall#1 Arena text overall (1508.6), #1 LiveBench Language (90.7), #3 EQ-BenchPriciest option here$10 / $50
Kimi K3Creative fiction and voice#1 EQ-Bench Creative Writing (2377), 234 Elo clear of secondOnly #10 on Arena creative writing$3 / $15
Claude Sonnet 5Free everyday writingFree and default on claude.ai, 1M context#53 Arena creative writing; 75.0 LiveBench Language$2 / $10 intro (then $3 / $15)
Claude Opus 5Professional deliverables#1 GDPval-AA v2 at 1858, ahead of Fable 5 (1746)Behind Fable 5 on Arena text and Humanity's Last Exam$5 / $25
GPT-5.5Fact-anchored business writingDocumented factual-reliability gains over GPT-5.4Reasoning tiers now marked deprecated$5 / $30
Gemini 3.6 FlashBulk drafts at scale17% fewer output tokens than 3.5 FlashWeaker on hardest reasoning$1.50 / $7.50
Runner-up and alternatives: Kimi K3 is the runner-up for creative fiction and wins EQ-Bench outright, Claude Sonnet 5 is the runner-up on value and the one to use if you are not paying, Claude Opus 5 is the pick for professional deliverables, and Gemini 3.6 Flash is the pick for bulk drafting.
What changed this month

The writing crown stays with Claude Fable 5, and it is the strongest-supported call on the page. Fable 5 is now #1 on Arena's text leaderboard at 1508.6 on the August 1 cutoff, #1 on Humanity's Last Exam at 53.3% and #1 on AA-Omniscience at 40, which is three different houses agreeing. Claude Opus 5 keeps the professional-deliverables lane on GDPval-AA v2, where the re-fitted board now reads 1858 against Fable 5's 1746.

Category pick, August 2026

Best AI for Chat & Daily Assistant

1
GPT-5.6
ChatGPT's default
2
Claude Fable 5
Highest-rated
3
Claude Opus 5
Thoughtful depth

The best AI for everyday chat is GPT-5.6, and the honest reason is reach rather than board position. It is the model ChatGPT serves by default to the largest user base in the category, which makes it the best assistant most people can actually open. On raw human preference it is not the leader: GPT-5.6 Sol sits #14 on Arena's text leaderboard at 1482.8, where Claude Fable 5 leads at 1508.6. Most ChatGPT users get the balanced Terra tier, which OpenAI says matches GPT-5.5 and, since the July 30, 2026 price cut, costs 60% less than it. It is available inside ChatGPT (free with limits, Plus at $20/month, Pro at $100/month), through the API (Luna $0.20 / $1.20, Terra $2 / $12, Sol $5 / $30 per 1M tokens), and bundled inside Fello AI alongside Claude, Gemini, Grok, and DeepSeek. One caveat: OpenAI's system card and the evaluator METR flagged elevated "scheming" behaviour in Sol, so GPT-5.5 Instant stays the safer pick for hallucination-sensitive work.

ModelBest ForStrengthWeaknessPrice
GPT-5.6Everyday chat, ChatGPT's defaultThe assistant most people can open; Terra matches GPT-5.5 at ~half cost#14 on Arena text; scheming flagged by METRFree / $20/mo Plus; API $0.20 / $1.20 to $5 / $30
Claude Fable 5Highest-rated conversation#1 on Arena text overall (1508.6) and 6 of 7 subcategoriesNo free tier; usage-credit access on Pro$10 / $50 API
Claude Opus 5Thoughtful, nuanced answers#1 Artificial Analysis Intelligence Index (61) and Agentic Index (55.3)#6 on Arena text, behind Claude Fable 5$20/mo Pro, $5 / $25 API
GPT-5.5 InstantHallucination-sensitive daily work52.5% fewer hallucinated claims vs 5.3 InstantReasoning tiers now marked deprecated$20/mo Plus; API $5 / $30
Gemini 3.6 FlashFast, free, multimodalFree in the Gemini app, 1M context, #12 on Arena textWeaker on hardest reasoningFree / $1.50 / $7.50 API
Fello AIAll the top models, one appChatGPT + Claude + Gemini + Grok + DeepSeek and more on Mac, iPhone and iPadRouted via app, not direct$9.99/mo
Runner-up and alternatives: Claude Fable 5 is the runner-up and the actual preference leader, Claude Opus 5 is the runner-up for thoughtful daily use, Gemini 3.6 Flash is the runner-up for fast and free, and Grok 4.5 is the niche pick for live-news days. Fello AI is the natural pick if you want the top models in one Mac and iOS app for $9.99/month instead of juggling subscriptions.
What changed this month

GPT-5.6 keeps the chat pick on reach, and the gap to the preference leader widened rather than closed. On the August 1 cutoff Sol sits #14 on Arena text at 1482.8, down from #11, while Claude Fable 5 leads at 1508.6. Nothing about the product changed, so the crown does not move: this is still about which assistant the most people can actually open.

Category pick, August 2026

Best AI for Images

1
ChatGPT Images 2.0
Text in images
2
Reve 2.1
Layout & 4K
3
Muse Image
Image editing

The best AI for image generation is ChatGPT Images 2.0, and it is the least controversial crown on this page. GPT Image 2 leads Arena's text-to-image board at 1385 Elo and its image-editing board at 1463, and Artificial Analysis puts it first on its own image arena at 1339.4. It is the natural pick whenever your image needs to contain readable words, in English or in another script, and it is included in ChatGPT Plus and Pro. The runner-ups have changed. Reve 2.1 (July 9) is the real #2 on text-to-image at 1302, and Reve 2.0 now sits behind it on both boards. Meta's Muse Image is #3 on Arena's text-to-image board and #2 on image editing, the strongest showing any Meta image model has managed. Google's Nano Banana Pro is no longer the runner-up overall: on text-to-image it ranks between #8 and #11, below its own cheaper sibling Nano Banana 2, though it does place higher on image editing.

ModelBest ForStrengthWeaknessPrice
ChatGPT Images 2.0Images with readable text#1 text-to-image (1385) and #1 image editing (1463)Less photoreal than the Gemini image lineIncluded in ChatGPT Plus
Reve 2.1Layout, typography, native 4K#2 text-to-image at 1302, layout-preserving editingSmaller ecosystemFree / from $7.99/mo
Muse ImageImage editing, Meta ecosystem#3 text-to-image, #2 image editing (1407)New, thin tooling around itMeta AI app
Nano Banana 2 (Gemini 3.1 Flash Image)Photoreal portraits and productsOutranks Nano Banana Pro on both boardsWeaker on text in imageGemini app / AI Studio
Seedream 5.0 ProMultilingual text + region editing10+ languages incl. Arabic RTL, lasso and layer editingNo independent benchmarks; copyright cloudBytePlus / Magnific
Midjourney v8Stylized art, illustrationAesthetic baseline most artists preferWeaker on text in image$10-$120/mo
Grok ImagineNSFW / Spicy ModeMost permissive guardrailsSmaller model behind it$30/mo SuperGrok
Runner-up and alternatives: Reve 2.1 is the runner-up overall and the pick for layout and typography, Muse Image is the runner-up for editing an image you already have, and Nano Banana 2 is the photoreal pick. Grok Imagine is still the only frontier model that allows Spicy Mode adult content.
What changed this month

The image crown is unchanged and GPT Image 2 still leads all three boards we track, on Arena text-to-image (1385), Arena image editing (1463) and Artificial Analysis's image arena (1339.4). The one addition is Microsoft's MAI-Image-2.5, which sits third on Artificial Analysis's image board at 1269.7 and third on Arena's image-editing board, so third place now depends on which house you read.

Category pick, August 2026

Best AI for Video

1
Gemini Omni Flash
#1 both video boards
2
MiniMax H3
2K + native audio
3
Dreamina Seedance 2.0
Closest challenger

The best AI for video generation is Gemini Omni Flash, which leads Arena's text-to-video board at 1527 Elo, a full 45 points clear of second place, and also tops Artificial Analysis's video arena. It is #1 on both houses, which no other video model manages. Pricing runs $1.50 in and $17.50 per 1M video output tokens, which works out at roughly $0.10 per second of finished video, and it supports conversational editing. The one real limit is length: Omni Flash generates 10-second clips. If you need longer takes, Veo 3.1 remains the right tool inside the Gemini app, AI Studio and Vertex AI, with native audio and 1080p output. This replaces Veo 3.1 at the top of the category; Veo 3.1 is a good model but not the leading one, with its best variant at #6 on Arena's text-to-video board. The contender to watch is MiniMax H3 (July 31), which generates 2K clips of 4 to 15 seconds with native stereo audio and is billed per second from $0.13 at 2K, already #2 on Artificial Analysis's video board at 1241.5. We are not moving the crown on that: the score is one day old, comes from the single board that has not reproduced between parses, and Arena's text-to-video vote cutoff predates H3.

ModelBest ForStrengthWeaknessPrice
Gemini Omni FlashBest AI video overall#1 on both video boards (1527 Arena), conversational editingCaps at 10-second generations~$0.10/sec; Gemini app / AI Studio
MiniMax H32K clips with native audio4-15s at 2K, native stereo audio; #2 on AA video (1241.5)One day old, no Arena votes yet; weights not published$0.13/sec at 2K
Dreamina Seedance 2.0Closest challenger#2 text-to-video (1482) and #1 on image-to-videoByteDance ecosystem, limited Western accessDreamina / BytePlus
Muse VideoMeta ecosystem video#3 text-to-video at 1459Newest of the group, thin toolingMeta AI app
Veo 3.1Longer production clipsNative audio, 1080p, strong physics consistency#6 on Arena video, not the quality leaderGoogle AI Pro / Ultra
Kling 3.0 / 3.0 TurboFast iteration at lower costNative 4K, 60fps, 15-second clips; Turbo shipped June 17Outside the top 16 on Arena text-to-videoFrom $10/mo
Luma Ray 3Photoreal scenesStrong realism for landscapesSmaller communityFree / from $9.99/mo
Runner-up and alternatives: Dreamina Seedance 2.0 is the runner-up overall and actually beats Omni Flash on image-to-video, Muse Video is third, and Veo 3.1 is the pick when 10 seconds is not enough. OpenAI retired the Sora 2 consumer app on April 26, 2026 and only the developer API remains, through September 24, 2026. We track the retirement dates for older models across every provider in a separate running list.
What changed this month

Gemini Omni Flash keeps the video crown, and it is still #1 on both boards, at 1527.5 on Arena and 1244.8 on Artificial Analysis. MiniMax H3 (July 31) enters as the contender at #2 on Artificial Analysis's video board (1241.5), 3.3 Elo behind, and it beats Omni Flash on specification with 2K output, clips up to 15 seconds and native stereo audio. We are holding the crown because that margin is one day old, single-board, and carries no votes on Arena, whose text-to-video cutoff predates the launch.

Category pick, August 2026

Best AI for Coding

1
Claude Opus 5
#1 Arena WebDev
2
Claude Fable 5
Hardest long-horizon
3
Kimi K3
Web apps and agents

The best AI for coding is Claude Opus 5, and it wins on the two boards where developers vote on the finished result rather than a script measuring a harness. It is #1 on Arena's WebDev board at 1,702.9 and #1 on image-to-WebDev at 1,668.6, on vote cutoffs of August 1 and July 31. Price is the second half of the argument: Opus 5 runs $5 / $25 per 1M tokens against Claude Fable 5's $10 / $50, and Artificial Analysis measures it at $2.34 per index task against Fable 5's $3.15. Anthropic's own docs tell developers to start with Opus 5 for complex agentic coding and reserve Fable 5 for workloads that need the highest available capability. Claude Fable 5 is the runner-up and stays the pick for the hardest long-horizon work, at #4 on WebDev (1,630.7) and #2 on image-to-WebDev (1,625.7). The contender is Kimi K3, which takes #2 on WebDev at 1,675.5 and #3 on Arena's Agent board, the strongest open-weight coder on the boards, though self-hosting it means 1.56 TB of weights. On Artificial Analysis's Coding Index it is not a Claude sweep: GPT-5.6 Sol (xhigh) leads at 78.3 with Opus 5 (max) at 78.0, close enough to call a tie. The cheapest serious contender is now DeepSeek V4-Flash 0731 at $0.14 / $0.28.

ModelBest ForStrengthWeaknessPrice (per 1M tokens)
Claude Opus 5Best coding overall#1 Arena WebDev (1,702.9) and #1 image-to-WebDev (1,668.6); $2.34 per index taskArtificial Analysis Coding Index has GPT-5.6 Sol a shade ahead$5 / $25
Claude Fable 5Hardest long-horizon agentic work#2 Arena image-to-WebDev (1,625.7), #4 WebDev; 1M contextPriciest; Artificial Analysis Coding Index puts it 7th$10 / $50
Kimi K3Web app building and agents#2 Arena WebDev (1,675.5), #3 Arena Agent, highest open Intelligence Index (57)1.56 TB to self-host$3 / $15
GPT-5.6 SolOpenAI flagship, agentic coding#1 Artificial Analysis Coding Index at 78.3 (xhigh)Absent from the official Terminal-Bench board; eval-gaming flagged by METR$5 / $30
Muse Spark 1.2Cheap agentic codingIntelligence Index 54 at $1.25 / $4.25; 80% on Terminal-Bench 2.1 (Artificial Analysis)Second to Opus 5 on all three of Meta's own coding charts (82.9% vs 86.7%); Scale SEAL has rated only 1.1$1.25 / $4.25
Grok 4.5Cheap value coder#4 on the official Terminal-Bench 2.1 board at 79.3% via Cursor CLIHigher hallucination rate; EU API console still closed$2 / $6
Gemini 3.6 FlashAgent coding at scaleIntelligence Index 50, 17% fewer output tokens than 3.5 FlashWeaker on hardest reasoning$1.50 / $7.50
GLM-5.2Best open-weight coder you can hostIntelligence Index 51, #6 Arena WebDev (1,587.1), MIT licenceKimi K3 outscores it; self-host or provider onlyOpen weights (MIT)
Runner-up and alternatives: Kimi K3 is the runner-up on Arena's WebDev board and the pick if you want the strongest open weights, Claude Fable 5 is the runner-up for the hardest long-horizon work, GPT-5.6 Sol is the runner-up on Artificial Analysis's composite, and GLM-5.2 is the open-weight pick for teams hosting it themselves. Inside IDEs, Cursor with Claude is still the most popular pairing and Claude Code is the natural pick if you live in the terminal.
What changed this month

The coding crown moved to Claude Opus 5. Arena finished rating it, and it came first on both coding boards, WebDev at 1,702.9 and image-to-WebDev at 1,668.6, with Claude Fable 5 fourth and second. Those are vote-based boards with published cutoffs, so the crown now rests on human comparisons rather than on any one harness, and Opus 5 does it at half Fable 5's price. Fable 5 becomes the runner-up for the hardest long-horizon work, and Kimi K3 stays the contender at #2 on WebDev. Meta's Muse Spark 1.2 arrived on August 5 and does not change that, because on Meta's own charts it comes second to Opus 5 on every coding benchmark it published.

Category pick, August 2026

Best AI for Creativity

1
Grok 4.5
Fewest guardrails
2
Claude Fable 5
Highest-quality prose
3
Kimi K3
Fiction and voice

The best AI for unfiltered, on-trend creative work is Grok 4.5, and we want to be exact about why. This pick is about the product, not the prose quality. Grok 4.5 carries the fewest content restrictions of any frontier model and the only native real-time X integration, which makes it the one model that will engage with edgy, topical or deliberately provocative briefs that the others decline. It is the default in the Grok app for SuperGrok and X Premium+ subscribers at $30/month. It is not the best writer, and the boards are blunt about it. Grok 4.5 sits #33 on EQ-Bench Creative Writing and #41 on Arena's creative-writing leaderboard, losing on both to the older Grok 4.20-beta1. If you are picking on output quality alone, Claude Fable 5 wins outright at #1 on Arena creative writing, and Kimi K3 wins EQ-Bench. Choose Grok 4.5 for what it will let you make, not for how well it writes.

ModelBest ForStrengthWeaknessPrice
Grok 4.5Unfiltered, opinionated, on-trendFewest content restrictions, native real-time X grounding#33 EQ-Bench, #41 Arena creative writing$30/mo SuperGrok
Claude Fable 5Highest-quality creative prose#1 Arena creative writing, #1 LiveBench LanguageCautious guardrails on edgy briefs$10 / $50 API
Kimi K3Fiction and distinctive voice#1 EQ-Bench Creative Writing at 2377Only #10 on Arena creative writing$3 / $15
Claude Opus 5Long-form structured creativityHolds long threads and self-edits; #1 Intelligence IndexMost cautious of the group$20/mo Pro; $5 / $25 API
Gemini 3.1 ProMultimodal creativeStrong text, image and video chainQuotas inside the Gemini appFree / $2.00-$4.00 API in
Grok Imagine (Spicy Mode)NSFW / adult creativeMost permissive image generationNiche use case$30/mo SuperGrok
Runner-up and alternatives: Claude Fable 5 is the runner-up and the right pick if quality matters more than freedom, Kimi K3 is the pick for fiction, and Claude Opus 5 is the pick for creative projects that run across many turns. For adult creative work, Grok Imagine Spicy Mode is still the only frontier-grade option.
What changed this month

Grok 4.5 keeps this pick, and nothing about the product moved. It is here for its permissiveness and its live X access, not for board position, and Claude Fable 5 remains the model to use when you want the better writing. One naming note for August: xAI now trades as SpaceXAI on the leaderboards, and the model is unchanged.

Category pick, August 2026

Best AI for Accuracy & Research

1
Gemini 3.1 Pro
98% ARC-AGI-1
2
Claude Fable 5
Grounded search
3
GPT-5.6 Sol
Novel reasoning

The best AI for accuracy and research is Gemini 3.1 Pro. Its strongest result is on ARC Prize's ARC-AGI-1, where it scores 98% and ties the human panel, and it does that at $0.52 per task. That combination is the argument: several models are close on capability, none matches it on cost for reliable factual work. It pairs that with native Google Search grounding, which is what you want when the answer has to be current rather than merely plausible. It also scores 94.1% on GPQA Diamond, where GPT-5.6 Sol now matches it rather than trailing it, and 44.4% on Humanity's Last Exam, and tops Scale SEAL's HLE board at 46.44. We have dropped the previous ARC-AGI-2 framing: its 77.1% is still correct, but the board has moved and that score now places it around 14th. Two honest caveats. On grounded search specifically, Arena's search leaderboard is led by Anthropic, not Google, with Gemini 3.1 Pro grounding at #7. And on novel reasoning, GPT-5.6 Sol leads ARC-AGI-2 and Claude Opus 5 leads ARC-AGI-3 at 30%, roughly 3.75x the next-best model. We did not move the crown to Opus 5 because Artificial Analysis measures its hallucination rate at 50% and places it below Fable 5 on AA-Omniscience.

ModelBest ForKey BenchmarkWeaknessPrice
Gemini 3.1 ProCheap, reliable factual work98% ARC-AGI-1 (ties human panel) at $0.52/task, 94.1% GPQA (tied by Sol)ARC-AGI-2 77.1% now ranks ~14th; #7 on Arena search$2.00-$4.00 / $12.00-$18.00 (tiered)
Claude Fable 5Grounded search#1 and #3 on Arena's search leaderboard, ahead of GoogleNo single cheap tier$10 / $50 (Fable 5)
GPT-5.6 SolNovel reasoning#1 ARC-AGI-2 at 93%, against a 100% human panelScheming flagged by METR$5 / $30
Claude Opus 5Hardest unseen problems#1 ARC-AGI-3 at 30%, ~3.75x the next model (ARC Prize)Artificial Analysis measures a 50% hallucination rate$5 / $25
Qwen 3.7 MaxFrontier accuracy at value pricing92.4 GPQA Diamond, 200 free requests/dayAPI-only, no chat front-end$1.25 / $3.75 promo; $2.50 / $7.50 list
Claude Opus 4.6Honesty under pressure#1 on Scale SEAL's MASK board at 96.28; Anthropic holds the top 5Superseded as a flagshipLegacy Anthropic model
Runner-up and alternatives: Anthropic's models are the runner-up for grounded search and sweep the honesty-under-pressure board, GPT-5.6 Sol is the runner-up for novel reasoning, and Qwen 3.7 Max is the value pick at the frontier.
What changed this month

Gemini 3.1 Pro keeps the accuracy crown, but its headline number is no longer a lead. Its GPQA Diamond score re-fitted to 94.1% and GPT-5.6 Sol now matches it exactly, so the crown rests on the ARC-AGI-1 result at $0.52 a task plus Search grounding rather than on GPQA. This is the next crown we re-test: Claude Fable 5 tops all three boards that measure how often a model is simply wrong, at 40 on AA-Omniscience, 61% on Omniscience Accuracy and 53.3% on Humanity's Last Exam.

Category pick, August 2026

Best AI for Problem Solving

1
GPT-5.6 Sol
STEM flagship
2
Claude Opus 5
Agentic chains
3
GPT-5.5 Pro
Verified FrontierMath

The best AI for hard problem solving is GPT-5.6 Sol, and it is the best-supported crown on this page. It takes #1 on LiveBench Mathematics at 96.2, #1 on LiveBench Reasoning at 91.7, and #1 on ARC-AGI-2 at 93%, the closest any model has come to the 100% human panel. Three separate houses put it first on the reasoning tasks that matter, which is more agreement than any other category on this page produces. OpenAI has still not published Sol's FrontierMath score, so the verified OpenAI mark remains GPT-5.5 Pro's 39.6% on FrontierMath Tier 4, and we will slot Sol's number in the moment it goes public. Qwen 3.7 Max is the value alternative for competition-style problems at 97.1 on the February 2026 HMMT index and 44.5 on Apex, at a fraction of the cost of ChatGPT Pro, and it now includes 200 free model requests per day. Claude Opus 5 is the alternative for long agentic reasoning chains, leading Artificial Analysis's Agentic Index at 55.3 and ARC Prize's ARC-AGI-3 at 30%, roughly 3.75x the next-best model.

ModelBest ForKey BenchmarkWeaknessPrice
GPT-5.6 SolHardest math, science and reasoning#1 LiveBench Mathematics (96.2), #1 LiveBench Reasoning (91.7), #1 ARC-AGI-2 (93%)FrontierMath still unpublished; scheming flagged by METR$100/mo ChatGPT Pro; API $5 / $30
Claude Opus 5Long agentic reasoning chains#1 Agentic Index (55.3), #1 ARC-AGI-3 (30%)Second on LiveBench Reasoning; 50% hallucination rate$5 / $25
GPT-5.5 ProVerified FrontierMath leader39.6% FrontierMath Tier 4Superseded by Sol as flagship$100/mo ChatGPT Pro
Qwen 3.7 MaxCompetition math on a budget97.1 HMMT 2026 Feb, 44.5 Apex, 200 free requests/dayAPI-only$1.25 / $3.75 promo; $2.50 / $7.50 list
Claude Fable 5Math inside a coding workflow#1 on Arena's math subcategory (1543), 96.0 LiveBench MathematicsPriciest option here$10 / $50
GLM-5.2Open-weight problem solvingHighest open Intelligence Index at 51, MIT, 1M contextSelf-host or provider onlyOpen weights (MIT)
Runner-up and alternatives: Claude Opus 5 is the runner-up and the natural pick for long-chain agentic reasoning, Claude Fable 5 is the runner-up on Arena's math board, Qwen 3.7 Max is the value pick, and GLM-5.2 is the open-weight pick. Our dedicated guide to the best AI for math covers the task-by-task split and the free options.
What changed this month

GPT-5.6 Sol keeps this crown, and it picked up a second argument. Sol now also leads Artificial Analysis's Terminal-Bench v2.1 at 89.5% and its Coding Index at 78.3, and it matches Gemini 3.1 Pro on GPQA Diamond at 94.1%. Claude Opus 5 stays the agentic-reasoning alternative on the strength of ARC-AGI-3. On August 1 OpenAI named its next model family Astra and published ten new mathematical results from an internal version, though Astra is unreleased and has no date, price or availability.

Category pick, August 2026

Best AI Agent

1
Gemini Spark
24/7 cloud
2
Claude Cowork
Desktop AI
3
ChatGPT Codex
Coding agent

The best AI agent right now is Gemini Spark for 24/7 cloud-resident work and Claude Cowork for desktop-resident work, with ChatGPT Codex as the alternative for coding agents and OpenAI Operator-class browser agents as the alternative for web tasks. AI agents are the fastest-moving category of 2026: each top vendor now ships an agent product, and the practical choice is between agents that live in the cloud (run while your laptop is closed) and agents that live on your desktop (drive your apps directly). Gemini Spark launched at Google I/O on May 19, 2026 and is the first 24/7 cloud agent. Claude Cowork launched in general availability on April 9, 2026 and runs as a desktop agent that drives your local apps. ChatGPT Codex Mobile (May 14) is the pick for coding-agent work. Read the full Gemini Spark vs Claude Cowork comparison.

AgentBest ForWhere It RunsStrengthPrice
Gemini Spark24/7 cloud tasks, Workspace workflowsGoogle Cloud VM (always-on)First true 24/7 agent, deep Workspace integration$99.99/mo Google AI Ultra
Claude CoworkDesktop, app-driving, design + codeYour Mac/Windows desktopDrives local apps, sees your screen$20/mo Claude Pro
ChatGPT Codex MobileCoding agent on phoneOpenAI cloud + iOS/AndroidApprove diffs and redirect work from phoneIncluded in ChatGPT plans
Grok Agentic (Grok 4.5)Real-time research, X scrapingxAI cloudNative X integration$30/mo SuperGrok
OpenAI Operator-classBrowser tasks, web formsOpenAI cloud + your browserWeb automationChatGPT Pro
Runner-up and alternatives: Claude Cowork is the runner-up overall and the natural pick when you want the agent on your machine driving your apps. ChatGPT Codex Mobile is the runner-up for coding agents. Grok Agentic is the niche pick for real-time research.
What changed this month

No new consumer agents shipped, and Meta's Muse Code (August 5) is a terminal developer tool rather than a consumer one, so the Gemini Spark (cloud) versus Claude Cowork (desktop) choice still drives most agent decisions for individual users. The model layer underneath them moved again: Claude Opus 5 (July 24) took #1 on Artificial Analysis's Agentic Index at 55.3, ahead of GPT-5.6 Sol at 54.0 and Claude Fable 5 at 52.8, and it costs $5 / $25. Opus 5 also tops Arena's Agent board, with Kimi K3 in third. For teams building their own agents, Meta's Muse Spark 1.2 (August 5) is a cheap agent-native option at $1.25 / $4.25 whose GDPval-AA v2 agentic Elo jumped from 1371 to 1631, though Scale SEAL has rated only the older 1.1, which leads its MCP Atlas tool-use board at 88.1. GLM-5.2 (MIT) is the strongest open-weight agent model you can host on modest hardware, at #5 on Arena's Agent board.

Use case guide, August 2026

Best AI for Students

1
GPT-5.5 Free
Essays & research
2
Gemini 3.6 Flash
STEM + PDFs
3
Claude Sonnet 5
Writing & editing

The best AI for students is GPT-5.5 Free inside ChatGPT for general coursework and Gemini 3.6 Flash Free inside the Gemini app for STEM and multimodal study, with Qwen 3.7 Max as the API alternative for harder problem sets (200 free requests a day) and Claude Opus 5 as the alternative for essay editing. Most students don't need to pay: the free ChatGPT tier now defaults to GPT-5.6 (with GPT-5.5 still available), Gemini 3.6 Flash is in the free Gemini app and AI Studio, Claude Sonnet 5 is the new free Claude default, and DeepSeek V4 is free on DeepSeek's chat site. For step-by-step working on the hardest math, GPT-5.6 Sol is OpenAI's new flagship (its FrontierMath score is not yet published, with GPT-5.5 Pro's verified 39.6% on FrontierMath Tier 4 the current mark), though both are paid-only; Qwen 3.7 Max is the value alternative at 97.1 HMMT 2026 February with API pricing at $1.25 / $3.75 on its current 50% promo ($2.50 / $7.50 list).

TaskBest ModelWhyFree?Alternative
Essays & courseworkGPT-5.5Free in ChatGPT, improved factual reliability vs 5.4YesClaude Sonnet 5 (free Claude)
STEM problem-solvingGPT-5.6 Sol / Qwen 3.7 MaxNew STEM flagship (5.5 Pro: 39.6% FrontierMath) / 97.1 HMMT 2026 FebPro paid / Qwen API paidGemini 3.6 Flash (free)
Research & accuracyGemini 3.1 Pro98% ARC-AGI-1 at $0.52/task, native Google Search groundingYes (Gemini app)Claude Opus 5
Writing editingClaude Sonnet 5Free and default on claude.ai; Claude Fable 5 is the quality leaderYes (Claude free)Claude Fable 5
Multimodal study (PDFs, slides, images)Gemini 3.6 Flash1M context, free in Gemini appYesNotebookLM (Google)
Runner-up and alternatives: Claude Sonnet 5 (free) is the runner-up for essay writing and editing. Gemini 3.6 Flash (free) is the runner-up for multimodal study and PDF ingestion. DeepSeek V4 is the runner-up for problem-solving on a strict zero-cost budget.
Use case guide, August 2026

Best AI for Work & Professionals

1
GPT-5.6
Daily knowledge work
2
Claude Opus 5
Coding & writing
3
Gemini 3.1 Pro
Research & briefings

The best AI for professional work is GPT-5.6 (ChatGPT's default since July 9) for daily knowledge work, Claude Opus 5 for coding and high-stakes writing, and Gemini Spark for 24/7 agentic workflows. Most professionals get the most out of running two paid subscriptions (ChatGPT Plus at $20/month plus Claude Pro at $20/month, total $40/month), or consolidating with Fello AI at $9.99/month for all five top models in one Mac/iOS app. For agentic work that runs while you sleep, Gemini Spark on Google AI Ultra at $99.99/month is the only true 24/7 cloud agent.

Use CaseBest ModelKey StatPriceAlternative
Daily knowledge workGPT-5.6ChatGPT's default since July 9; the assistant most people can open$20/mo ChatGPT PlusClaude Opus 5
Coding (proprietary)Claude Opus 5#1 on Arena WebDev (1,702.9) and image-to-WebDev (1,668.6); Anthropic's recommended default$20/mo Claude ProClaude Fable 5
Coding (cost-effective)Qwen 3.7 Max80.4 SWE-Verified, 1M context$1.25 / $3.75 promo; $2.50 / $7.50 listDeepSeek V4-Flash 0731
Research & briefingsGemini 3.1 Pro98% ARC-AGI-1 at $0.52/task, Google groundingGoogle AI Pro / UltraClaude Opus 5
Hard math, physics, finance modellingGPT-5.6 SolOpenAI's new STEM flagship (5.5 Pro verified at 39.6% FrontierMath)$100/mo ChatGPT ProQwen 3.7 Max
Always-on agent workflowsGemini SparkFirst 24/7 cloud agent$99.99/mo Google AI UltraClaude Cowork
Live news, X-context creativeGrok 4.5Opus-class + native X grounding$30/mo SuperGrokGemini 3.1 Pro
All-in-one consolidationFello AIChatGPT + Claude + Gemini + Grok + DeepSeek$9.99/moPay each vendor separately
Runner-up and alternatives: for most professional teams, Claude Opus 5 is the runner-up to GPT-5.6 for daily work and now the coding leader outright at $5 / $25, with Claude Fable 5 the runner-up for the hardest long-horizon work. Gemini 3.1 Pro is the runner-up for research-heavy roles, and Gemini Spark is the unique pick if you can put a cloud agent to work on long tasks.
Pricing Comparison

AI Model Pricing in August 2026

From $0 free tiers to $199.99/month Google AI Ultra. The most consequential price on this table is Claude Opus 5 at $5 / $25, because it is the #1 model on Artificial Analysis's Intelligence Index at half the cost of Claude Fable 5. Meta's Muse Spark 1.2 lists at $1.25 / $4.25, with a Contributor tier at $0.10 / $0.20 for anyone willing to let Meta train on their prompts, and Grok 4.5 lists at $2 / $6. The price-performance pick is now DeepSeek V4-Flash 0731 at $0.14 / $0.28, which matches Gemini 3.6 Flash's Intelligence Index of 50 and costs $0.03 per index task against Flash's $0.56. Among closed models GPT-5.6 Luna is the cheapest at $0.20 / $1.20 after OpenAI's July 30 cut. For a deeper breakdown see our full AI Pricing Comparison Guide.

ModelInput (per 1M)Output (per 1M)ContextFree Access?
GPT-5.5$5.00$30.001M (400K in Codex)ChatGPT Free; API paid
GPT-5.5 Pro$30.00$180.001MChatGPT Pro from $100/mo ($200 higher-usage tier)
GPT-5.6 Sol$5.00$30.00Not publishedLive in ChatGPT, Codex & API (July 9)
GPT-5.6 Terra$2.00$12.00Not publishedLive in ChatGPT, Codex & API (July 9)
GPT-5.6 Luna$0.20$1.20Not publishedLive in ChatGPT, Codex & API (July 9)
Claude Opus 5$5.00$25.001MClaude Pro/Max default; API paid
Claude Opus 4.8$5.00$25.001MLegacy model at Anthropic; Pro/Max/API
Claude Fable 5$10.00$50.001MPermanent in Max/Team Premium (~50% of usage limits); Pro/Team Standard via credits
Claude Sonnet 5$2.00 intro / $3.00 list$10.00 intro / $15.00 list1MClaude Free & Pro default; API paid
Claude Sonnet 4.6$3.00$15.001MAPI paid (superseded by Sonnet 5)
Gemini 3.1 Pro$2.00 (≤200K) / $4.00 (>200K)$12.00 (≤200K) / $18.00 (>200K)1MLimited Gemini app; API paid
Gemini 3.6 Flash$1.50$7.501MGemini app/AI Studio; free API tier + paid API
Gemini 3.5 Flash-Lite$0.30$2.501MAI Studio; free API tier + paid API
Qwen 3.7 Max$1.25 promo / $2.50 list$3.75 promo / $7.50 list1M200 free requests/day; API paid beyond that
MiniMax M3$0.30 (50% off $0.60)$1.20 (≤512K)1MOpen weights; hosting costs apply
LongCat-2.0Provider-dependentProvider-dependent1MOpen weights (MIT); hosting costs apply
NVIDIA Nemotron 3 UltraProvider-dependentProvider-dependent1MOpen weights (OpenMDW); hosting costs apply
Qwen 3.5 (open-weight)Self-host / TogetherSelf-host / Together1MOpen weights; hosting costs apply
Nex-N2-ProSelf-host / providersSelf-host / providers1MOpen weights (Apache 2.0); hosting costs apply
Rio 3.5 Open 397BSelf-host / providersSelf-host / providers1MOpen weights (MIT); hosting costs apply
Grok 4.3$1.25$2.501MFree consumer plan; API paid
Grok 4.5$2.00$6.00500KGrok Build / Cursor / xAI console; EU partial, API console still closed
Muse Spark 1.2$1.25 ($0.10 Contributor tier)$4.25 ($0.20 Contributor tier)1MPaid API; Muse Code, Meta Model API and OpenRouter; Contributor tier trades training rights for a ~92% discount
Kimi K3$3.00 ($0.30 cache-hit)$15.001MFree basic tier in the Kimi app; open weights on Hugging Face (Kimi K3 License)
Gemini Omni Flash (video)$1.50$17.50 (video output)10-second clipsGemini app / Flow; AI Studio + API
DeepSeek V4-Pro$0.435 ($0.0036 cache-hit)$0.871MDeepSeek Chat free; API paid
DeepSeek V4-Flash 0731$0.14$0.281MDeepSeek Chat free; API paid
Kimi K2.7 CodeProvider-dependentProvider-dependent256KOpen weights; hosting costs apply
GLM-5.2Provider-dependentProvider-dependent1MOpen weights; hosting costs apply
ERNIE 5.1China-region pricingChina-region pricing256KBaidu free tier
Gemini Spark (agent)Not API-pricedNot API-priced1M (Gemini base)Google AI Ultra $99.99 or $199.99/mo
Fello AI (aggregator)Routed via appRouted via appModel-dependent$9.99/mo, free tier available

The GPT-5.5 and GPT-5.5 Pro rates above are short-context prices; OpenAI no longer publishes the specific long-context figures. The GPT-5.6 tiers are billed at 2x input and 1.5x output once a prompt passes 272K input tokens, which puts long-context Terra at $4 / $18 and Luna at $0.40 / $1.80. If you want access to multiple AI models without managing separate subscriptions, Fello AI provides GPT, Claude, Gemini, Grok, Perplexity, and more in a single app for Mac, iPhone, and iPad from $9.99/month.

Open-Weight Models

Best Open-Weight Models in August 2026

The best open-weight model in August 2026 is Kimi K3, and it took the lead the moment Moonshot published the weights on July 27, 2026. It holds the highest Intelligence Index of any open model at 57 on Artificial Analysis v4.1, six points clear of GLM-5.2, and on Arena it is still the highest-placed open entry, at #2 on WebDev (1,675.5) behind only Claude Opus 5 and #3 on the Agent board. One catch decides which of the two you should actually use. The K3 download is 96 safetensors shards and about 1.56 TB, which needs a multi-node GPU cluster rather than a workstation, and it ships under a custom Kimi K3 License rather than MIT. So GLM-5.2 (Z.ai, MIT) stays our practical recommendation for teams running their own weights, at Intelligence Index 51, #6 on Arena WebDev (1,587.1) and #5 on the Agent board. Third place changed hands on the last day of July: DeepSeek V4-Flash 0731 was re-post-trained and jumped from Intelligence Index 40 to 50, at $0.14 / $0.28 under MIT.

ModelBest ForKey BenchmarkContext / LicenseWhere To Run
Kimi K3Highest-scoring open model, agentic and web workII 57 (v4.1), highest of any open model; #2 Arena WebDev, #3 Arena Agent; 2.8T/104B active1M / Kimi K3 LicenseHugging Face (96 shards, 1.56 TB), Moonshot API, providers
GLM-5.2Long-horizon agentic coding, 1M contextII 51 (v4.1), highest you can host on modest hardware; 744B/40B active1M / MITZ.ai, Hugging Face, OpenRouter
DeepSeek V4-Flash 0731Best value of any model on the pageII 50 (v4.1), up from 40 on July 31; $0.03 per index task, 284B/13B active1M / MITDeepSeek API ($0.14/$0.28), local
LongCat-2.0Frontier open coder trained on Chinese chipsII 33 (v4.1); 59.5% SWE-Bench Pro (vendor), 1.6T/~48B active1M / MITHugging Face, GitHub, OpenRouter
MiniMax M3Cheap frontier-class multimodalII 44 (v4.1), 59% SWE-Bench Pro, multimodal1M / license TBDHugging Face, API $0.30/1M (50% off)
Nex-N2-ProStrongest open coding scoreII 41 (v4.1); 80.8 SWE-Bench Verified, 397B/17B activeQwen-based / Apache 2.0Hugging Face, providers, self-host
Kimi K2.7 CodeStrongest commercially-licensed open coder+21.8% on Kimi Code Bench v2 vs K2.6 (vendor); 1T/32B active256K / Modified MITHugging Face, DeepInfra, providers
DeepSeek V4-ProAgentic real-world workII 44 (v4.1), 1.6T/49B active1M / MITDeepSeek API ($0.435/$0.87), local
Hy3Newest permissive-licence entrantII 41 (v4.1); #22 on Arena WebDev (1,516.4)Apache 2.0Hugging Face, providers, self-host
InklingThinking Machines' first open modelII 41 (v4.1), agentic 32.3, released July 15Open weightsHugging Face, providers, self-host
Inkling SmallSame family at a quarter the sizeII 40 (v4.1), agentic 30.8; 276B/12B active, text, image and audio inApache 2.0Hugging Face (BF16 and NVFP4), providers, self-host
NVIDIA Nemotron 3 UltraNVIDIA-tuned, fully permissive licenseII 38 (v4.1), 65-70.4 SWE-Bench Verified, 550B/55B active1M / OpenMDWOpenRouter, Hugging Face, AWS (8x B200 self-host)
Qwen 3.5 (397B / 17B active)Multimodal, fast decode88.4 GPQA, 91.3 AIME 2026, 83.6 LiveCodeBench v61M / openTogether, OpenRouter, local
Qwen3.6-35B-A3BEfficient open agentic coder (3B active)86.0 GPQA Diamond, 92.7 AIME 2026, 35B/3B active262K (to 1M YaRN) / Apache 2.0Hugging Face, OpenRouter, local
Qwen3.6-27BLaptop-runnable dense coder87.8 GPQA Diamond, dense 27B, multimodal256K / Apache 2.0Local Mac/PC, Hugging Face, OpenRouter
Rio 3.5 Open 397BQwen 3.5 fine-tune, multilingual reasoning70.8 Terminal-Bench 2.1 (first-party), beats Qwen 3.7 Plus on 4/5397B/17B active / MITHugging Face, providers, self-host
Llama 4 MaverickMeta-line flagship17B active / 400B total paramsLlama 4 licenseMeta cloud, Hugging Face, local
NVIDIA Nemotron 3 Nano OmniEdge / low-powerMultimodal, very small footprintCompact / openLocal, NVIDIA tool

Licensing matters as much as raw score here: Kimi K2.7 (Modified MIT), DeepSeek V4 (MIT), GLM-5.2 (MIT), LongCat-2.0 (MIT), Hy3 (Apache 2.0), Nex-N2-Pro (Apache 2.0) and Nemotron 3 Ultra (OpenMDW) all clearly allow commercial use, while MiniMax M3 ships under its own community license and Kimi K3 sits on its own custom licence with a revenue threshold for anyone reselling it as a service. No open model is first on any Arena board we track any more: Claude Opus 5 took the Agent board in July, the last one an open model led.

How We Evaluate

Benchmarks, Prices, and Hands-On Use

Every ranking on this page combines three inputs: public benchmarks from seven independent houses (Artificial Analysis, Arena formerly LMArena, Scale SEAL, LiveBench, EQ-Bench, ARC Prize and the official Terminal-Bench 2.1 board, covering the Intelligence and Agentic indexes, GPQA Diamond, ARC-AGI-1 through 3, Humanity's Last Exam, GDPval-AA, FrontierMath, HMMT, MCP Atlas, SWE Atlas and the Remote Labor Index), published API and subscription pricing from each vendor's official pricing page, and hands-on use by the FelloAI editorial team running real prompts across the same task on every model. We re-fetch official pricing and benchmark sources before every monthly update.

Benchmarks are weighted to the use case: SWE-bench and Terminal-Bench drive coding, GPQA Diamond and ARC-AGI-2 drive accuracy, GDPval-AA (Artificial Analysis's professional-deliverables benchmark) informs professional-task quality while writing style is judged primarily by hands-on testing, FrontierMath and HMMT drive problem-solving. We disclose when a benchmark is vendor-reported but not independently verified, and we strip any claim we cannot reproduce against a live source. We do not move a category crown on the strength of a single leaderboard, and where a new model is too recent for the human-preference boards to have rated it, we say so rather than crowning it on benchmark scores alone. When a model goes through a major upgrade between updates, we re-rank the category and add a "What changed this month" line at the bottom of the deep-dive.

Frequently Asked Questions

Common Questions

What is the best AI model right now in August 2026?
It depends on the task. On overall benchmark score, Claude Opus 5 (July 24) is #1 on Artificial Analysis's Intelligence Index at 61 and its Agentic Index at 55.3, and Arena has now rated it top of four of its boards, including both coding boards. For daily chat, GPT-5.6 has been ChatGPT's default since July 9 and is the assistant most people can actually open, although Claude Fable 5 leads Arena's text leaderboard. For coding, Claude Opus 5 is #1 on Arena's WebDev (1,702.9) and image-to-WebDev (1,668.6) boards at half Fable 5's price. For writing, Claude Fable 5 is #1 on Arena's text leaderboard (1508.6), Humanity's Last Exam (53.3%) and AA-Omniscience (40). For accuracy, Gemini 3.1 Pro ties the human panel on ARC-AGI-1 at 98% for $0.52 a task. For hard math and reasoning, GPT-5.6 Sol is #1 on LiveBench Mathematics, LiveBench Reasoning and ARC-AGI-2. For images, ChatGPT Images 2.0 leads both boards, and for video Gemini Omni Flash leads both. For agents, Gemini Spark is the 24/7 cloud agent and Claude Cowork the desktop one.
What is Claude Opus 5?
Claude Opus 5 is Anthropic's newest flagship, released July 24, 2026, and the current #1 model on Artificial Analysis. It leads both the Intelligence Index at 61 and the Agentic Index at 55.3, and tops the GDPval-AA v2 professional-deliverables board at 1858, well clear of Claude Fable 5's 1746. API pricing is $5 / $25 per 1M tokens, the same as Opus 4.8 and half the cost of Fable 5, with a 1M-token context window and a May 2026 knowledge cutoff. ARC Prize independently confirms Anthropic's launch claim on ARC-AGI-3, where Opus 5 scores 30% against 8% for the next-best model, roughly 3.75x. Arena has since rated it, placing it #1 on WebDev, image-to-WebDev, Document and the Agent board and #6 on text, and that is what moved the coding crown to it. One thing to keep in mind: Artificial Analysis measures its hallucination rate at 50%, which is why it holds no accuracy crown on this page.
Is Claude Fable 5 back?
Yes. Anthropic redeployed Claude Fable 5 on July 1, 2026 after the US government lifted the export-control restriction it had imposed on June 12. It is available again on the Claude API, Claude.ai, Claude Code, and Claude Cowork. Anthropic made Fable 5 permanent in the paid plans on July 20, 2026. Max and Team Premium include it at roughly 50% of regular usage limits; Pro and Team Standard reach it through usage credits with a one-time $100 starting credit. API pricing is $10 / $50 per million tokens. Fable 5 is a Mythos-class model built for long-horizon agentic work with a 1M-token context, and it is the runner-up for coding behind Claude Opus 5, at #2 on Arena's image-to-WebDev board (1,625.7) and #4 on WebDev.
What is GPT-5.6 and can I use it?
Yes, as of July 9, 2026. GPT-5.6 is OpenAI's next-generation model family, and after a two-week gated preview that began June 26 behind a US-government safety review, it reached general availability on July 9. There are three tiers, least to most capable: Luna, the fast, cheapest tier ($0.20 / $1.20 per 1M tokens); Terra, a balanced everyday model OpenAI says matches GPT-5.5 ($2 / $12); and Sol, the flagship tuned for biology, chemistry, and cybersecurity ($5 / $30). OpenAI cut Luna by 80% and Terra by 20% on July 30, 2026, and left Sol and every ChatGPT subscription price untouched. It is now live across ChatGPT, Codex, and the API as OpenAI's default model. One caveat: OpenAI's system card and the external evaluator METR flagged elevated "scheming" behaviour in Sol, so treat it carefully for high-stakes factual work until independent results settle.
Is Grok 4.5 out yet?
Yes. xAI released Grok 4.5 publicly on July 8, 2026, its first flagship since SpaceX absorbed the company and went public as SPCX. Elon Musk describes it as "an Opus-class model, but faster, more token-efficient and lower cost." It is Cursor-trained and aimed at coding and agentic work, priced at $2 / $6 per 1M tokens with a 500K-token context window. It is live in Grok Build, Cursor on all plans, and the xAI console. EU access began rolling out after a July 16 announcement and is still partial, with Cursor reporting availability while xAI's API console remains closed to EU users. Artificial Analysis scores it at Intelligence Index 54 on v4.1, matching the field on Terminal-Bench 2.1 (83.3%) but trailing on SWE-Bench Pro at 64.7%, so it is the value pick rather than the outright benchmark leader. Grok 4.5 is also our creativity pick, on product grounds rather than quality, given it sits #33 on EQ-Bench Creative Writing; for the best-written output, Claude Fable 5 wins outright.
Which AI is the best for coding?
Claude Opus 5 is the best for coding, and it wins the two boards where developers vote on the result rather than a harness running a script: #1 on Arena's WebDev board (1,702.9) and #1 on image-to-WebDev (1,668.6), on August 1 and July 31 vote cutoffs. It runs $5 / $25 per 1M tokens, half of Claude Fable 5, and Anthropic's own docs say to start with Opus 5 for complex agentic coding. Claude Fable 5 is the runner-up for the hardest long-horizon work at #2 image-to-WebDev and #4 WebDev, and Kimi K3 is the contender at #2 on WebDev (1,675.5). Note that Artificial Analysis's Coding Index is not a Claude sweep: GPT-5.6 Sol (xhigh) leads it at 78.3 with Opus 5 at 78.0, close enough to call a tie. DeepSeek V4-Flash 0731 is the price-performance pick at $0.14 / $0.28, and GLM-5.2 (MIT) is the best open-weight coder you can host yourself.
Which AI is the best for writing?
Claude Fable 5 is the best for writing, and it is the only model in the top three of all three independent writing boards: #1 on Arena creative writing, #1 on LiveBench Language (90.7) and #3 on EQ-Bench Creative Writing v3. It costs $10 / $50 per 1M tokens. Claude Sonnet 5 is the value pick and what most people should actually use, since it is free and default on claude.ai at $2 / $10 introductory pricing, though it ranks #53 on Arena creative writing. Kimi K3 wins EQ-Bench outright at 2377 if you want distinctive fiction, Claude Opus 5 tops the GDPval-AA v2 professional-deliverables board at 1858, GPT-5.5 is the alternative for fact-anchored business writing, and Gemini 3.6 Flash is the price-performance pick for bulk content.
What is the best open-weight AI model in 2026?
Kimi K3 is the best open-weight model right now, and it has been since Moonshot published the weights on July 27, 2026. It has the highest Intelligence Index of any open model at 57 on Artificial Analysis v4.1, and on Arena it is #2 on WebDev and #3 on the Agent board. The catch is that it is 96 shards and about 1.56 TB under a custom Kimi K3 License, so GLM-5.2 (June 13, MIT) remains the model most teams can actually host, at Intelligence Index 51 and #6 on Arena WebDev. Third is DeepSeek V4-Flash 0731, which jumped from Intelligence Index 40 to 50 on July 31 without changing price, size or licence, making it the best value in the open field at $0.14 / $0.28 under MIT. Worth knowing honestly, no open model is first on any Arena board we track any more.
What is the cheapest frontier-class AI model?
On API pricing per million tokens, GPT-5.6 Luna is now the cheapest closed frontier-class model at $0.20 / $1.20, after OpenAI cut it 80% on July 30, 2026, and Artificial Analysis scores it at Intelligence Index 51, a point above Gemini 3.6 Flash. Cheaper still is DeepSeek V4-Flash 0731 at $0.14 / $0.28, which is now our price-performance pick: it matches Gemini 3.6 Flash's Intelligence Index of 50 and costs $0.03 per index task against Flash's $0.56, under an MIT licence you can self-host. MiniMax M3 now lists at $0.30 per million input tokens on a permanent 50% discount, with LongCat-2.0 as a strong MIT alternative. Qwen 3.7 Max at $1.25 / $3.75 on its current promo is the value pick at Intelligence Index 46, though Luna now undercuts it on both price and score.
Which AI models are free?
ChatGPT Free now defaults to GPT-5.6 (with GPT-5.5 still available) under usage limits. Gemini Free runs Gemini 3.6 Flash in the Gemini app and Google AI Studio. Claude Free runs Claude Sonnet 5 (the new default) with daily limits. DeepSeek Chat runs DeepSeek V4 free on the DeepSeek website. Grok has a limited free consumer plan (X Premium is a paid add-on). Qwen 3.5, NVIDIA Nemotron 3 Ultra, MiniMax M3, LongCat-2.0, Kimi K2.6, DeepSeek V4, and GLM-5.2 are open-weight and free to self-host. Qwen 3.7 Max is API-only with no consumer chat front-end, but Alibaba now includes 200 free model requests per day. Kimi has a free basic tier in its app, with heavier agentic use metered.
What is Gemini Spark and is it worth $99.99/month?
Gemini Spark is Google's first 24/7 cloud-resident AI agent, launched at Google I/O on May 19, 2026 and exclusive to the Google AI Ultra plan, which Google restructured to include a $99.99/month entry tier and a $199.99/month top tier. Spark is built on Gemini base models with Google's Antigravity harness on a Google Cloud VM, integrates with Gmail, Google Docs, and other Google Workspace apps, and can interact with Chrome and Android's Halo system on the device side. It is worth the spend for users who have repeatable long-running workflows (inbox triage, research roll-ups, scheduled tasks). For one-off tasks, Claude Cowork at $20/month covers most desktop-agent needs.
What is Fello AI?
Fello AI is an AI chatbot for Mac, iPhone, and iPad that lets you use all top AI models like ChatGPT, Claude, Gemini, Grok, and DeepSeek in one app, with models updated regularly so you always have the latest. It is $9.99/month with a 4.7-star rating across 27,000+ reviews.
How often do you update this page?
We update this page at least monthly and within 24-48 hours of any major model launch.
Related Articles

Try every model.
One beautiful app.

Every model from this guide in one native app: ChatGPT, Claude, Gemini, Grok, and DeepSeek. Free to start.

Fello AI running on Mac, iPad, and iPhone

4.7 rating·27,000+ reviews·Free to start