Blog

/

Guides

/

Best Local LLMs for RTX 5090: Benchmarks

Best Local LLMs for RTX 5090: Benchmarks

If you're wondering which AI model is best for a desktop GeForce RTX 5090, you're in the right place. We tested six local models on the 32GB card, measuring generation speed, GPU memory use, and how long each takes to answer questions about the same 128k-token document.

Best Local LLMs for RTX 5090: Benchmarks
Alex Shapiro
Alex Shapiro
Calendar icon

September 18, 2026

Table of Contents

This is a guide to the desktop RTX 5090. The RTX 5090 Laptop GPU has a different memory budget and is not part of this ranking.

Quick Answer

Qwen3.8-27B is our general-purpose pick for the RTX 5090. It generated 79.4 tokens/s and processed a 4k prompt at 5,100 tokens/s. If output speed matters most, Nemotron 3.5 Lightning reached 334.3 tokens/s in the tested AtomicChat build. Gemma 4 26B-A4B offers a strong balance of speed and long-document work: it generated 248.6 tokens/s and took 24.4 seconds to answer the first question after our 128k-token document. All six models answered the document questions correctly, so the useful differences were wait time and VRAM use. Every number on this page came off a real RTX 5090 running the model, not from a memory-bandwidth estimate.

Best Local LLMs for RTX 5090

ModelTested GGUF fileFile (GiB)VRAM at 32k (MiB)4k prompt (tok/s)Generation (tok/s)Best for
Qwen3.8-27BQwen3.8-27B-iMatrix-NVFP4-MTP.gguf15.9518,0485,10079.4General-purpose use
Gemma 4 31B itgemma-4-31B-it-Q4_K_M.gguf17.4022,3723,69868.5Dense Gemma alternative
Gemma 4 26B-A4B itgemma-4-26B-A4B-it-Q4_K_M.gguf15.6417,70211,829248.6Fast long-document work
Gemma 4 12B itgemma-4-12b-it-Q4_K_M.gguf6.878,7588,659145.4Lowest VRAM use
Muse Glimmer 30BMuse-Glimmer-30B-AD-Q4_K_M.gguf17.7518,3924,38773.9Reasoning-first assistant
Nemotron 3.5 Lightning 30B-A3BNVIDIA-Nemotron-3.5-Lightning-30B-A3B-AD-IQ4_NL.gguf18.3018,61212,047334.3Fast output and first token

1. Qwen3.8-27B

Qwen3.8-27B is our all-round pick for a desktop RTX 5090. The dense model combines conventional attention with Gated DeltaNet layers and supports a native 262k-token context. Its strong performance in coding benchmarks makes it a great starting point for development work.

Model factValue
DeveloperQwen / Alibaba
ReleaseAugust 14, 2026
Parameters and architecture27.32B, dense, 64 hybrid Gated DeltaNet / attention layers
Native context262,144 tokens
Base-model inputsText, images, and video
ThinkingSwitchable, reasoning depth can be adjusted
Recommended tested buildcdiamond/Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF then Qwen3.8-27B-iMatrix-NVFP4-MTP.gguf, 15.95 GiB

Released in August 2026, it is the 27B dense member of the Qwen3.8 family. The original model accepts text, images, and video and has switchable thinking. Our GGUF tests used text with thinking off. For this card we recommend the cdiamond community file: 70.8% of its weights are NVFP4, with the rest in Q5_K, Q6_K, and Q8_0.

BenchmarkResult
GPQA Diamond
Expert science
89.2
SWE-bench Pro
Harder engineering
61.7
Terminal-Bench 2.1
Terminal agents
73.0
LiveCodeBench v6
Competitive coding
90.3

The results above are for the original Qwen model.

MeasurementAtomicChat AD-Q4_K_MAtomicChat AD-Q5_K_Mcdiamond NVFP4
Exact filenameQwen3.8-27B-AD-Q4_K_M.ggufQwen3.8-27B-AD-Q5_K_M.ggufQwen3.8-27B-iMatrix-NVFP4-MTP.gguf
File size15.94 GiB18.84 GiB15.95 GiB
GPU memory at 32k context18,356 MiBnot measured18,048 MiB
4k prompt processing3,830 tok/s3,540 tok/s5,100 tok/s
128-token generation74.2 tok/s69.2 tok/s79.4 tok/s

Verdict: Choose the exact cdiamond NVFP4 GGUF for Qwen3.8-27B on Blackwell. It matched the AD-Q4 file size and processed the prompt faster. The 128k document took 62.6 seconds to produce its first token and used 24,773 MiB with the tested AtomicChat AD-Q4 build; do not assume those long-context figures transfer unchanged to the cdiamond file.

2. Gemma 4 31B it

Gemma 4 31B it is Google's largest dense Gemma 4 model. It is a useful alternative if you prefer Gemma's output style, but it is also the most memory-hungry of our six at long context.

Model factValue
DeveloperGoogle DeepMind
ReleaseApril 2026
Parameters and architecture30.7B, dense, 60 layers
Native context262,144 tokens
Original-model inputsText and images
Tested buildAtomicChat/gemma-4-31B-it-GGUF then gemma-4-31B-it-Q4_K_M.gguf, 17.40 GiB

Gemma 4 31B uses local sliding-window and global attention across its layers. Unlike the 26B-A4B model below, it uses all its language-model weights to produce each token. The table below shows how it performs in benchmarks. These are official Google numbers, and Google doesn't specify a per-row thinking setting:

BenchmarkResult
GPQA Diamond
Expert science
84.3%
LiveCodeBench v6
Competitive coding
80.0%
MRCR v2, eight needles at 128k
66.4%

Our RTX 5090 test:

MeasurementResult
GPU memory at 32k context22,372 MiB
4k prompt processing3,698 tok/s
128-token generation68.5 tok/s
First token after our 128k document85.0 s
GPU memory with the 128k document30,361 MiB
Document facts / summary topics at 128k4/4 / 4/4

Verdict: The dense Gemma alternative works at 128k on the RTX 5090, but it leaves only 2,246 MiB of the card's 32,607 MiB free in our test. Choose a lighter model if long-document response time or room for other GPU tasks matters more. Among our six it also holds the best published long-context score, 66.4% on MRCR v2 at 128k.

3. Gemma 4 26B-A4B it

If you need more generation speed at the small cost of reasoning performance, and want to stick to the Gemma family, choose Gemma 4 26B-A4B it. Its mixture-of-experts design stores about 25.2B parameters but activates roughly 4B for each token, helping it generate much faster than the dense 31B model.

Model factValue
DeveloperGoogle DeepMind
ReleaseApril 2026
Parameters and architecture25.2B stored, about 4B active per token; 30-layer MoE
Native context262,144 tokens
Original-model inputsText and images
Tested buildAtomicChat/gemma-4-26B-A4B-it-GGUF then gemma-4-26B-A4B-it-Q4_K_M.gguf, 15.64 GiB

The smaller active parameter count reduces computation per generated token, which is what makes generation fast here. Its lighter memory use at long context comes from a different place: a smaller file plus the way its attention layers store the KV cache. In our setup, this model used 19,831 MiB at 128k against 30,361 MiB for Gemma 4 31B. And here's how this model performs in the benchmarks, according to Google's official measurements for the original full precision file:

BenchmarkResult
GPQA Diamond
Expert science
82.3%
LiveCodeBench v6
Competitive coding
77.1%
MRCR v2, eight needles at 128k
44.1%

The table below displays results during our RTX 5090 test:

MeasurementResult
GPU memory at 32k context17,702 MiB
4k prompt processing11,829 tok/s
128-token generation248.6 tok/s
First token after our 128k document24.4 s
GPU memory with the 128k document19,831 MiB
Document facts / summary topics at 128k4/4 / 4/4

Verdict: Choose Gemma 4 26B-A4B when you want fast generation and a 128k document without consuming most of the RTX 5090's VRAM. On Google's own MRCR v2 test at 128k it scores 44.1% against 66.4% for the dense 31B, so pick the 31B when the answer matters more than the wait.

4. Gemma 4 12B it

Gemma 4 12B it is the smallest model in our lineup of the best models to run on an RTX 5090. As the smallest file, it makes the most sense when you need an AI model to run alongside other apps that also need GPU memory. The GGUF file we tested occupied 8,758 MiB at a 32k allocation and 10,447 MiB with our 128k document.

Model factValue
DeveloperGoogle DeepMind
ReleaseApril 2026
Parameters and architectureAbout 11.9B, dense
Native context262,144 tokens
Original-model inputsText and images
Tested buildAtomicChat/gemma-4-12B-it-GGUF then gemma-4-12b-it-Q4_K_M.gguf, 6.87 GiB

This is a much lighter model than the two larger Gemmas we've tested, yet it completed all of the planted-fact questions and preserved all four required topics in our 128k summary.

BenchmarkResult
GPQA Diamond
Expert science
78.8%
LiveCodeBench v6
Competitive coding
72.0%
MRCR v2, eight needles at 128k
43.4%

And here are the results from our test run:

MeasurementResult
GPU memory at 32k context8,758 MiB
4k prompt processing8,659 tok/s
128-token generation145.4 tok/s
First token after our 128k document30.1 s
GPU memory with the 128k document10,447 MiB
Document facts / summary topics at 128k4/4 / 4/4

Verdict: Choose Gemma 4 12B when you need a capable AI model that generates tokens fast and offers great performance for the size, or when you need additional GPU headroom. In our test, Gemma 4 12B used less than half the 128k VRAM use of Qwen or Gemma 4 31B.

5. Muse Glimmer 30B

Muse Glimmer 30B is Meta's agentic model with a dedicated perception encoder. Meta built it for tool use, multi-step reasoning and failure recovery on consumer hardware, and lists OpenClaw and Hermes Agent among the scaffolds it works with. Both of those already run on top of Atomic Chat's local endpoint.

Model factValue
DeveloperMeta
ReleaseAugust 2026
Parameters and architectureAbout 29.6B including the perception encoder, dense, 52 layers
Native context131,072 tokens and above
Original-model inputsText and images through the perception encoder
Tested buildAtomicChat/Muse-Glimmer-30B-GGUF then Muse-Glimmer-30B-AD-Q4_K_M.gguf, 17.75 GiB

Muse Glimmer continues to use reasoning tokens even at a low reasoning setting. Give it enough output budget for both thought and final answer. It also found all four facts at 8k, 32k, and 128k. It uses a denser tokenizer, which encoded our 128k-class document in 107,856 tokens against 130,557 for the Gemma models. Here's how this model performs in the benchmarks, according to Meta:

BenchmarkResult
GPQA Diamond (AA)
83.5%
HLE Text (AA)
22.0%
SWE-bench Pro
Harder engineering
51.2%
SWE-bench Verified
Software engineering
76.0%
Terminal-Bench 2.1 with terminus2
51.7%

And here's how it fared in our test on an RTX 5090 graphics card:

MeasurementResult
GPU memory at 32k context18,392 MiB
4k prompt processing4,387 tok/s
128-token generation73.9 tok/s
First token after our long document31.8 s
GPU memory with the long document19,752 MiB
Document facts / summary topics4/4 / 4/4 at all three lengths

Verdict: Muse Glimmer is the pick when you run agents locally. On Meta's own numbers it reaches 76.0% on SWE-bench Verified and 51.7% on Terminal-Bench 2.1, well ahead of the other agentic candidate here. Its denser tokenizer also fits more text into the same context length.

6. NVIDIA Nemotron 3.5 Lightning 30B-A3B

NVIDIA Nemotron 3.5 Lightning 30B-A3B is a great pick to run on an RTX 5090 when you're optimizing for inference speed. Its hybrid Mamba-2, attention, and mixture-of-experts architecture stores about 32.9B parameters but activates roughly 3B per token, and allows the model to output tokens quickly, especially for its parameter count.

Model factValue
DeveloperNVIDIA
ReleaseAugust 2026
Parameters and architectureAbout 32.9B stored, 3B active; hybrid Mamba-2 / attention / MoE
Model context limitUp to 1M tokens
Original-model inputsText
Tested buildAtomicChat/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF then NVIDIA-Nemotron-3.5-Lightning-30B-A3B-AD-IQ4_NL.gguf, 18.30 GiB

The architecture is built to reduce per-token work and handle long sequences. In our tested build, it generated 334.3 tokens/s and returned the first token after the 128k document in 14.9 seconds. Those figures are for ordinary generation without speculative decoding. Here's how the model performs in benchmarks, according to NVIDIA:

BenchmarkResult
GPQA Diamond (no tools)
75.44%
SWE-bench Verified
Software engineering
51.56%
Terminal-Bench 2.1
Terminal agents
24.58%

And here are the results of our RTX 5090 test:

MeasurementResult
GPU memory at 32k context18,612 MiB
4k prompt processing12,047 tok/s
128-token generation334.3 tok/s
First token after our 128k document14.9 s
GPU memory with the 128k document19,367 MiB
Document facts / summary topics at 128k4/4 / 4/4

Verdict: Choose Nemotron 3.5 Lightning for fast output and long-document response time. On published agentic benchmarks it trails Muse Glimmer, so pick it for speed rather than for task success.

How the Six Models Compare

Only GPQA Diamond is published for all six models, so the table below marks which numbers come from the same benchmark and which do not. Each vendor runs its own harness: Qwen scored SWE-bench Pro with the Claude Code harness at a 256k context, while Meta scored it with its own setup. Treat cross-vendor rows as indicative, not as a head-to-head.

ModelGPQA DiamondLiveCodeBench v6SWE-bench ProSWE-bench VerifiedTerminal-Bench 2.1MRCR v2 at 128kSource
Qwen3.8-27B89.290.361.7not published73.0not publishedQwen
Gemma 4 31B it84.380.0not publishednot publishednot published66.4Google
Gemma 4 26B-A4B it82.377.1not publishednot publishednot published44.1Google
Gemma 4 12B it78.872.0not publishednot publishednot published43.4Google
Muse Glimmer 30B83.5not published51.276.051.7not publishedMeta
Nemotron 3.5 Lightning75.44not publishednot published51.5624.58not publishedNVIDIA

Two comparisons here are like for like. On MRCR v2 at 128k the three Gemma models line up 66.4, 44.1 and 43.4, so the dense 31B is far better at finding things in a long document than its faster siblings. On SWE-bench Verified and Terminal-Bench 2.1, Muse Glimmer reaches 76.0 and 51.7 against 51.56 and 24.58 for Nemotron, so the two agentic candidates are not close.

Picking Context Over Parameters

The RTX 5090's 32 GB lets you spend memory on a bigger model or on a longer context, and the two compete for the same card. Our 128k measurements show how wide that spread is.

ModelVRAM with the 128k documentFree on a 32,607 MiB cardFits in 24 GB
Gemma 4 12B it10,447 MiB22,160 MiBYes
Nemotron 3.5 Lightning19,367 MiB13,240 MiBYes
Muse Glimmer 30B19,752 MiB12,855 MiBYes
Gemma 4 26B-A4B it19,831 MiB12,776 MiBYes
Qwen3.8-27B AD-Q424,773 MiB7,834 MiBNo
Gemma 4 31B it30,361 MiB2,246 MiBNo

Gemma 4 12B is the clearest case for trading parameters for room. It held our full 128k document in 10,447 MiB, less than half of what Qwen3.8-27B needed and a third of Gemma 4 31B, and it still answered all four planted-fact questions. That leaves 22 GB for a second model, an image model, or the rest of your desktop.

The trade is real, though, and Google's own numbers price it. On MRCR v2 at 128k the 12B scores 43.4% against 66.4% for the dense 31B. A long context you can load is not the same as a long context the model reads well. Load the biggest model whose answers you trust, then give the remaining memory to context.

Recommended file type for Blackwell architecture

The RTX 5090 uses NVIDIA's Blackwell architecture and has FP4 tensor cores. In our llama.cpp test, the NVFP4 GGUF files processed 4k prompts 28 to 58% faster than the standard GGUF build of the same model. Generation improved too, by 2 to 10%. For a long input, faster prompt processing reduces the wait before generation starts.

ModelOrdinary GGUF: 4k prompt / generationCommunity NVFP4 GGUF: 4k prompt / generationChange
Qwen3.8-27B3,830 / 74.2 tok/s5,100 / 79.4 tok/s+33% prompt; +7% generation
Gemma 4 31B3,698 / 68.5 tok/s5,826 / 69.7 tok/s+58% prompt; +2% generation
Nemotron 3.5 Lightning12,047 / 334.3 tok/s15,374 / 367.2 tok/s+28% prompt; +10% generation

The exact community NVFP4 files tested were Qwen3.8-27B-iMatrix-NVFP4-MTP.gguf, gemma-4-31B-it-NVFP4-turbo-NVFP4.gguf, and Nemotron-3.5-Lightning-30B-A3B-NVFP4.gguf, respectively.

  • The Qwen file has 70.8% of its weights in NVFP4 and the rest in Q5_K/Q6_K/Q8_0.
  • The Gemma and Nemotron files are 95.4% and 95.6% NVFP4 by weight.

These are measured results for these specific files on llama.cpp b10988. We did not isolate how much of the gain comes from the FP4 tensor cores and how much from the smaller weights. NVIDIA's own card for Nemotron lists the GeForce RTX 5090 compute path as still to be confirmed, so treat the numbers as what these builds did on our card rather than as a property of the format.

How we tested these models

The main throughput and document tests used a desktop RTX 5090 with 32,607 MiB of VRAM, Ubuntu 24.04, and the official CUDA 12.8 build of llama.cpp b10988. We used full GPU offload, one parallel slot, fp16 KV cache, and automatic flash-attention selection. Throughput comes from llama-bench -p 512,4096 -n 128 -r 2; the tables show the 4k-prompt and 128-token-generation results. The Gemma files here are AtomicChat GGUF builds, not the QAT files from the earlier run. Qwen AD-Q5 speed and the Muse long-document run used a second pod with a different driver, so those two rows are not strictly comparable with the rest.

One caveat on the throughput table: the six models do not all run the same quantization recipe. Three Gemma files are Q4_K_M, Muse Glimmer is AD-Q4_K_M, Nemotron is AD-IQ4_NL, and Qwen is the mixed cdiamond NVFP4 file. Part of the spread between rows is the build, not the model.

Coding a Self-Playing Snake

We asked each model to build a snake game that plays itself, eats food, and grows. This test used vLLM with NVFP4 builds: AtomicChat/gemma-4-12B-it-NVFP4, AtomicChat/gemma-4-26B-A4B-it-NVFP4 and AtomicChat/gemma-4-31B-it-NVFP4 for the Gemma models, and nvidia/Qwen3.8-27B-NVFP4, nvidia/Muse-Glimmer-30B-NVFP4 and nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 for the rest.

ModelAttemptsTotal generation time
Qwen3.8-27B297.5 s
Gemma 4 31B147.9 s
Gemma 4 26B-A4B127.1 s
Gemma 4 12B3134.0 s
Muse Glimmer 30B2169.8 s
Nemotron 3.5 Lightning545.7 s

Every model eventually completed the test, creating a snake that chases food, grows, and increases its score as the game plays.

Gemma 4 26B-A4B and 31B were able to build the game on the first attempt, and the 26B-A4B was also quickest at 27.1 seconds. Nemotron needed five attempts, but its total generation time was 45.7 seconds. Its repair requests used thinking off after the initial response exhausted the output budget.

Coding Five Bouncing Balls

We also asked each model to make a page with five balls affected by gravity and collisions. This test used the same vLLM and NVFP4 setup as the snake task.

ModelAttemptsTotal generation time
Qwen3.8-27B288.5 s
Gemma 4 31B154.5 s
Gemma 4 26B-A4B132.9 s
Gemma 4 12B4161.7 s
Muse Glimmer 30B2142.1 s
Nemotron 3.5 Lightning331.8 s

As in the snake test, only Gemma 4 26B-A4B and 31B worked on the first attempt. Despite needing three attempts, Nemotron had the shortest generation time at 31.8 seconds.

Finding facts in a 128k-Token Document

We gave each model the same source document at roughly 8k, 32k, and 128k tokens. We planted facts near the beginning, middle, and end, then asked a fourth question requiring facts from two separate places.

We also checked whether a summary retained four specified topics. The document's token count differs by tokenizer, even though the source text is the same.

ModelFacts at 8k / 32k / 128kSummary topics at 128kInput tokens processed at longest lengthTime to first tokenVRAM at longest length
Nemotron 3.5 Lightning4/4 · 4/4 · 4/44/4129,62814.9 s19,367 MiB
Gemma 4 26B-A4B4/4 · 4/4 · 4/44/4130,55724.4 s19,831 MiB
Gemma 4 12B4/4 · 4/4 · 4/44/4130,55730.1 s10,447 MiB
Muse Glimmer 30B4/4 · 4/4 · 4/44/4107,85631.8 s19,752 MiB
Qwen3.8-27B, AtomicChat AD-Q44/4 · 4/4 · 4/44/4126,82062.6 s24,773 MiB
Gemma 4 31B4/4 · 4/4 · 4/44/4130,55785.0 s30,361 MiB

All six passed this particular retrieval and summary test at every length, so it separates them on wait time rather than on accuracy. For a harder read of long-context quality, use the published MRCR v2 scores in the comparison table above. Two rows also did less work than the others: Muse Glimmer processed 107,856 tokens against 130,557 for the Gemma models, and its run used the second pod.

At this length, Gemma 4 31B used 30,361 MiB and Qwen3.8-27B AD-Q4 used 24,773 MiB. Both fit on our 32,607 MiB RTX 5090, but exceed the nominal 24 GB (24,576 MiB) VRAM budget of an RTX 4090 or 3090. The other four models used 10,447 to 19,831 MiB in this test, below 24 GB. These are measured RTX 5090 memory footprints; the table does not show their speed on older GPUs. Qwen's 128k result is for the AtomicChat AD-Q4 file, not the cdiamond NVFP4 file recommended above.

How to Run Offline AI Models in Atomic Chat

1. Install Atomic Chat. When choosing a backend, select Find optimal backend for your GPU.

2. For an AtomicChat-built model in the table, open Models, search the linked AtomicChat repository, choose the exact filename shown in its card under Download Options, and download it.

3. For our recommended Qwen build, download Qwen3.8-27B-iMatrix-NVFP4-MTP.gguf from cdiamond's repository. In Atomic Chat, go to Settings, then Model Providers, then llama.cpp, then Import, and select that local GGUF file.

4. Select Use this model, open a new chat, and start with a 32k context allocation. Increase the context when a task requires it, watching GPU memory use. If the imported NVFP4 file does not load, update or select a compatible llama.cpp backend.

Atomic Chat's own self-hosted LLM guide walks through the app's model download and import controls.

FAQ

What is the best LLM to run on an RTX 5090?

For a balance of speed and performance, Qwen3.8-27B is the best AI model to run on an RTX 5090. If your priority is the shortest task completion time or the highest generation speed, choose Nemotron 3.5 Lightning.

Is 32 GB VRAM enough for a local LLM?

Yes. 32 GB of VRAM is enough to run most 12B to 33B models in 4-bit quantization entirely on the GPU, with room for a long context. Larger models or very long context windows may need a smaller quantization or partial CPU offload.

Can an RTX 5090 run a 70B model?

Not at 4-bit. A 70B Q4 file generally exceeds the RTX 5090's 32 GB VRAM before context memory is added. You can reach a 70B at a lower bit rate or with part of the model on the CPU, but both cost quality or speed.

What is the best coding LLM for RTX 5090?

Qwen3.8-27B holds the best published coding scores in our lineup, at 61.7 on SWE-bench Pro and 90.3 on LiveCodeBench v6. In our own two coding tasks it was slower to a working result than the Gemma models: Gemma 4 26B-A4B built both the snake and the balls page on the first attempt, in 27.1 and 32.9 seconds, against two attempts and 97.5 and 88.5 seconds for Qwen. Pick Qwen for harder problems and Gemma 4 26B-A4B when you want a working answer quickly.

How much power and system RAM does an RTX 5090 setup need?

NVIDIA specifies 575 W total graphics power for the RTX 5090 and a 1,000 W system power requirement for the reference configuration. System RAM needs depend on file size, context, and whether any model layers run on the CPU. Our benchmark host had 117 GB of RAM; that is not a requirement for these GPU-resident files.

Best Local AI Video Generators Compared on Quality, Speed and VRAM

Best Local AI Video Generators Compared on Quality, Speed and VRAM

Compare local AI video generators on quality, speed and VRAM, with RTX 4090 benchmarks, example clips and an Atomic Chat setup guide.

10/9/26

16 min

Best Local LLMs for 12GB VRAM in 2026

Best Local LLMs for 12GB VRAM in 2026

Eight local LLMs for 12GB VRAM: exact GGUF files, RTX 3080 Ti speed and memory figures, plus Snake and physics tests on an RTX 4070.

9/30/26

15 min

Claude Sonnet 5.5 Alternatives Compared

Claude Sonnet 5.5 Alternatives Compared

Compare Claude Sonnet 5.5 with Opus, Fable and GPT-6 Astra, then choose a local Qwen, Ornith or Bonsai model for your hardware.

9/29/26

13 min

Jev 1.13: Can You Run It Locally? Laya Setup Guide

Jev 1.13: Can You Run It Locally? Laya Setup Guide

Can you run Jev 1.13 locally? Learn how it works, see our Jev vs Laya Tetris demo, and set up Laya as an independent local alternative.

9/25/26

12 min