Blog

/

Guides

/

Claude Sonnet 5.5 Alternatives Compared

Claude Sonnet 5.5 Alternatives Compared

Claude Sonnet 5.5 is Anthropic's new model for everyday coding and document work. It runs on Anthropic's servers. This guide covers Claude Sonnet 5.5 alternatives, starting with benchmarks against Opus, Fable and GPT-6 Astra, then the local models you can run on your own hardware.

Claude Sonnet 5.5 Alternatives Compared
Alex Shapiro
Alex Shapiro
Calendar icon

September 29, 2026

Table of Contents

Sonnet 5.5 comes close to Opus 5.5 on some of Anthropic's published benchmarks, at half the standard API token price. Opus still leads on FrontierCode, while Fable 5.1 and GPT-6 Astra offer other cloud options to compare. The results depend on the task and the evaluation settings, so the tables below explain where those comparisons hold.

If you want to run a model on your own computer, the choice changes. Anthropic does not publish Sonnet's weights. Open-weight alternatives range from large models such as GLM-5.3 and Kimi K3, which need server hardware, to Qwen3.8 27B and Ornith 1.5 9B for a desktop. We'll compare their published results before choosing a model and download for your hardware.

In this guide, you'll learn:

What is Claude Sonnet 5.5?

Anthropic released Claude Sonnet 5.5 on September 28, 2026, as the successor to Sonnet 5. It sits below Opus 5.5 in the lineup, with lower token prices.

Claude Sonnet 5.5 main specs, from Anthropic's model documentation:

Swipe to compare →

SpecificationClaude Sonnet 5.5
API model IDclaude-sonnet-5-5
Context window1 million tokens
Maximum output128K tokens
ModalitiesText and image input; text output
API input price$2 per million tokens
API output price$10 per million tokens
Local weightsNo public checkpoint

You can use Sonnet through Claude or a supported cloud provider. There is no Sonnet GGUF to download. If you want the model to work offline, you need one of the open-weight alternatives later in this guide.

How much does Sonnet 5.5 cost compared with Opus and Fable?

These are the standard API token rates:

Swipe to compare →

ModelInput, per 1M tokensOutput, per 1M tokens
Claude Sonnet 5.5$2$10
Claude Opus 5.5$4$20
Claude Fable 5.1$10$50

A request with 100,000 uncached input tokens and 10,000 output tokens costs $0.30 on Sonnet at those rates, before any additional tool charges; running the same token counts through Opus costs $0.60.

The amount you pay for a task also depends on how long the model reasons and how many attempts it makes. In Artificial Analysis's own evaluation, Sonnet 5.5 at max effort cost $7.60 per Intelligence Index task, about 50% more than Sonnet 5, in an evaluation using a prerelease deployment with a subsequently fixed structured-output bug.

Start with the effort setting you expect to use and compare the total bill for completed tasks. Opus is the next model to test when Sonnet needs too many retries; Anthropic's model-selection guidance reserves Fable for demanding work where higher-effort Opus still falls short.

Claude Sonnet 5.5 benchmarks

Anthropic's launch table compares Sonnet 5.5 with Sonnet 5, Opus 5.5 and GPT-6 Sol:

Swipe to compare →

Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Terminal-Bench 4.0
Terminal agents
70.610.366.4-
FrontierCode v1.1 Main
Code changes
46.242.454.449.3
CursorBench 4.0
Coding agents
55.534.157.8-
GDPval-AA v2.1
Professional work
1844144918461487
HLE with tools
Tool-assisted reasoning
64.554.967.7-

Sonnet nearly matches Opus on GDPval-AA and leads Terminal-Bench, while Opus keeps the higher FrontierCode score.

FrontierCode reports Sonnet at max effort; xhigh scores 52.1. Opus's Terminal-Bench result uses xhigh. GDPval-AA is an Elo score; the other rows are percentages. The launch footnotes disclose a fixed prerelease bug affecting structured outputs and a recent GPT-6 Sol vision fix. A hyphen means no result in that source.

Sonnet 5.5 vs Fable 5.1 and GPT-6 Astra

Fable and Astra appear in Anthropic's Opus 5.5 announcement. The following table places those published results beside Sonnet's launch scores:

Swipe to compare →

Benchmark Sonnet 5.5 Fable 5.1 GPT-6 Astra
Terminal-Bench 4.0
Terminal agents
70.655.857.9
FrontierCode v1.1 Main
Code changes
46.250.353.3
HLE with tools
Tool-assisted reasoning
64.565.657.2

These results come from two Anthropic release tables with different evaluation settings. Astra's Terminal-Bench score uses high effort; Sonnet's FrontierCode result above uses max. Anthropic attributes Astra's results to OpenAI.

Terminal-Bench tests completing terminal tasks; FrontierCode tests producing changes a reviewer would merge. Our GPT-6 Astra alternatives guide covers Astra and its local options in more detail.

Open-weight alternatives to Claude Sonnet 5.5

Downloading an open-weight model lets you run inference on your own hardware. For codebase edits, you also need a coding tool that can use the model; downloading weights alone does not reproduce Claude Code's workflow.

GLM-5.3 is a large open-weight coding model with published vLLM and SGLang deployment instructions. Kimi K3 has 2.8 trillion total parameters and 104 billion active per token. Both need server hardware. Qwen3.8 27B and Ornith 1.5 9B are the smaller models to look at for a desktop.

MoE models use only some of their parameters for each token. The download still contains the full model. Our Kimi K3 guide covers the hardware and setup for that class of model.

Sonnet 5.5 vs open-weight models

Here's how the vendors score their models. Each lab uses its own evaluation setup:

Swipe to compare →

Benchmark Sonnet 5.5 GLM-5.3 Kimi K3 Qwen3.8 Flash Next Qwen3.8 27B Ornith 1.5 9B
HLE with tools
Tool-assisted reasoning
64.562.5---30.5
DeepSWE v1.1
Software engineering
-66.967.558.742.2-
GPQA Diamond
Expert science
--93.591.789.286.4

Sources: Anthropic, Z.ai, Moonshot, Qwen Flash Next, Qwen 27B and Ornith.

Tools, judges and reasoning budgets differ across vendors, so a small gap between columns does not establish a winner. Kimi separately reports 56.0 on HLE-Full with tools; we leave its HLE cell empty because the source labels do not establish the same test edition. For DeepSWE, Kimi uses Kimi Code, GLM uses mini-swe-agent, and Qwen takes the higher score from Claude Code and mini-SWE-agent. These scores describe the original models rather than the quantized GGUF files recommended below.

Sonnet's launch table does not report DeepSWE or GPQA Diamond. Qwen's HLE results are not labeled as tool-assisted, so they do not belong in the first row.

Qwen reports 58.7 on DeepSWE for Flash Next and 42.2 for the 27B under its published evaluation procedure. Flash Next runs on a 64 GB Mac with our split build; the 27B fits a 24 GB GPU.

Local Claude Sonnet 5.5 alternatives for your hardware

The system requirement to check is memory. These are specific downloads, with room still needed for the runtime and conversation cache:

Swipe to compare →

Your hardwareModel and buildDownload size
8 GB GPU, start with text and short contextOrnith 1.5 9B, AD-Q5_K-Q4_K5.93 GB
24 GB GPU or 32 GB Apple Silicon MacQwen3.8 27B, AD-Q4_K_M17.1 GB
64 GB Apple Silicon Mac with SSD spaceFlash Next, AD-3.84bpw-IQ4_XS-M6484.9 GB total; 45.8 GB GPU-resident weights

The Ornith and Qwen 27B vision projectors add about 0.92 GB and 0.93 GB respectively if you use images. Flash Next stores its large n-gram table on SSD with the split Atomic build, which is why its download can exceed the Mac's memory.

Qwen3.8 27B for a 24 GB GPU or 32 GB Mac

For this hardware, Qwen3.8 27B is the model to start with. It accepts text and images, and the 17.1 GB Atomic AD-Q4_K_M file leaves room on a 24 GB GPU for the runtime and a conversation starting at 8,192 tokens of context. Longer conversations use more memory through the KV cache, so a model fitting at 8K does not mean its full context window will fit on the same card.

Our Qwen3.8 27B guide covers the other builds and the llama.cpp commands. Use that guide if you want a larger quant or need to split the model between GPU and system memory.

Qwen3.8 Flash Next for a 64 GB Mac

Flash Next has a 125B core with 6B active parameters, plus separate n-gram embeddings and a prediction head; the Atomic build keeps the n-gram table on SSD while the main weights remain in memory.

Our GGUF repository reports generation at 36 tokens per second for the AD-3.84bpw-IQ4_XS-M64 build on a 64 GB MacBook Pro M5 Max.

Follow the Flash Next setup guide for the split files and memory settings. The 45.8 GB figure covers the GPU-resident weights. The runtime and context cache still need memory.

Ornith 1.5 9B for an 8 GB GPU

Ornith 1.5 9B is a dense model with a much smaller download: our 5.93 GB AD-Q5_K-Q4_K build is the repository's recommended option for 8 GB of VRAM. Start with text and a short context before adding the vision projector.

Its vendor reports 70.6 on SWE-bench Verified and 47.5 on SWE-bench Pro, which test different tasks from Sonnet's Terminal-Bench result. For the files and settings, see our Ornith 1.5 guide.

Bonsai 2 27B if you can use Prism's runtime

Bonsai 2 27B compresses Qwen3.8 27B into ternary weights. Its PTQ1_0 language-model file is 5.95 GB; the PQ2_0 packing is 7.21 GB.

These files require Prism ML's llama.cpp build because stock llama.cpp cannot read their ternary packing formats; follow the Bonsai 2 guide for the fork and its setup. The local Atomic Chat steps below are for Qwen's standard GGUF build.

Its vendor benchmarks measure compression against the original Qwen model; they do not establish that Bonsai matches Sonnet 5.5.

How to use Claude Sonnet 5.5 via API in Atomic Chat

You can connect Anthropic's API to Atomic Chat and use Sonnet alongside your local models. Sonnet runs through Anthropic's hosted service. API usage is billed separately from a Claude subscription.

Step 1: Get an Anthropic API key

Open the Claude Console and create an API key. Set up API billing if you have not used the service before.

Step 2: Connect Anthropic

Open Cloud in Atomic Chat and select Anthropic. Paste your key into API key, then click Connect. Keep the default Base URL.

Step 3: Add Sonnet 5.5

Click Reload models and look for claude-sonnet-5-5. If it is missing, click the + button beside Reload models. Enter claude-sonnet-5-5 in Model ID, then click Add Model.

Step 4: Start a chat

Open a chat and select Sonnet 5.5 from the model picker. Your Anthropic account must have API access to the model.

You can switch to a downloaded local model in the same app when you want your computer to generate the answers.

How to run a local alternative with Atomic Chat

Atomic Chat includes a Hugging Face model browser and a built-in chat. You can download a local model without building llama.cpp yourself.

Here's how to run Qwen3.8 27B:

Step 1: Install Atomic Chat

Download the desktop app from atomic.chat and install the build for your platform.

Step 2: Find the Qwen GGUF

Open Models and search for AtomicChat/Qwen3.8-27B-GGUF. Choose the AtomicChat repository, then expand Download Options.

Atomic Chat model search showing the Qwen3.8-27B-GGUF repository published by AtomicChat
Choose the Qwen3.8-27B-GGUF repository published by AtomicChat.

Step 3: Download the build for your memory

For a 24 GB GPU or 32 GB Mac, choose AD-Q4_K_M, whose full filename is Qwen3.8-27B-AD-Q4_K_M.gguf; the picker can show the shorter Q4_K_M tag, and displayed download sizes can differ from the weight-file sizes in the table.

Atomic Chat Download Options showing Qwen3.8 27B quantization variants including Q4_K_M
Download Options in our Qwen guide. The screenshot shows IQ3_S selected; choose Q4_K_M for the configuration above.

Step 4: Set the context and start a chat

Open Settings > Model Providers > Llama.cpp, find the downloaded model, and click the gear icon on its row. Set Context Size to 8192. If memory is tight, turn off Auto Increase Context Size so a longer chat does not allocate more memory than your machine has.

Load the downloaded model and open a chat. Your computer generates the answers. To connect a compatible coding tool, Atomic Chat also exposes a local OpenAI-compatible server at http://localhost:1337/v1; the local LLM setup guide covers that route.

Frequently asked questions

Can Claude Sonnet 5.5 run locally?

No. There is no public Sonnet 5.5 checkpoint to download. Calling Claude from a desktop app still uses a hosted model; Qwen, Ornith and Bonsai provide downloadable alternatives.

Is Sonnet 5.5 better than Opus 5.5?

Sonnet leads Terminal-Bench 4.0 in Anthropic's launch table, while Opus leads the other four rows shown above. Compare completed tasks at the effort setting you plan to use.

What is the best local Sonnet 5.5 alternative for a 24 GB GPU?

Start with Qwen3.8 27B and the 17.1 GB Atomic AD-Q4_K_M build at 8K context. It is a practical local option for that hardware, with text and image support.

Can I run a local alternative on an 8 GB GPU?

Yes. Ornith 1.5 9B has a 5.93 GB Atomic build recommended for that memory tier. Begin with text and a short context; images require an additional projector and more working memory.

Can I load Bonsai 2 like a normal GGUF?

Its ternary GGUF files require a runtime that implements their packing formats. Prism documents its own llama.cpp build for this purpose. Follow that setup instead of assuming every GGUF loader supports the files.

Do local models use my Claude subscription?

No. Inference with a downloaded model runs on your computer and does not consume your Claude allowance. Your hardware still uses electricity, and any external tools you connect can have their own costs.

Which model should you use?

Keep Sonnet or Opus for work where their results justify cloud access. For local use on a 24 GB GPU or 32 GB Mac, download Qwen3.8 27B and start at 8K context. Choose Ornith 1.5 9B when memory is tighter. Flash Next is the larger option for a 64 GB Mac, with its split build and SSD-backed table.

Best Local AI Video Generators Compared on Quality, Speed and VRAM

Best Local AI Video Generators Compared on Quality, Speed and VRAM

Compare local AI video generators on quality, speed and VRAM, with RTX 4090 benchmarks, example clips and an Atomic Chat setup guide.

10/9/26

16 min

Best Local LLMs for 12GB VRAM in 2026

Best Local LLMs for 12GB VRAM in 2026

Eight local LLMs for 12GB VRAM: exact GGUF files, RTX 3080 Ti speed and memory figures, plus Snake and physics tests on an RTX 4070.

9/30/26

15 min

Jev 1.13: Can You Run It Locally? Laya Setup Guide

Jev 1.13: Can You Run It Locally? Laya Setup Guide

Can you run Jev 1.13 locally? Learn how it works, see our Jev vs Laya Tetris demo, and set up Laya as an independent local alternative.

9/25/26

12 min

Best Local AI Image Generators in 2026: Apps and Models

Best Local AI Image Generators in 2026: Apps and Models

Compare local AI image generators, with seven models benchmarked on an RTX 5090. See image examples, generation speed, memory use, and license limits.

9/25/26

14 min