Blog

/

Guides

/

Best Local LLMs for RTX 3090: 6 Models Tested

Best Local LLMs for RTX 3090: 6 Models Tested

We tested six local LLM builds on a 24 GB RTX 3090, comparing generation speed, memory use and working code. Five share the main benchmark; Bonsai 2 uses a separate runtime.

Best Local LLMs for RTX 3090: 6 Models Tested
Alex Shapiro
Alex Shapiro
Calendar icon

September 21, 2026

Table of Contents

Quick Answer

Gemma 4 26B-A4B Q4_K_M is our overall pick among the six builds tested on a desktop RTX 3090. It generated 159.8 tokens/s and used 17,430 MiB at 32k. Start with gemma-4-26B-A4B-it-Q4_K_M.gguf and a 32k context allocation. Its largest tested fp16 allocation was 192k, using 20,789 MiB.

For a straightforward low-memory setup, Gemma 4 12B Q4_K_M used 8,486 MiB at 32k and completed both coding tasks on its first attempt. Ternary Bonsai 2 27B had the smallest measured footprint, 7,064 MiB peak at 16k, but requires the Prism llama.cpp fork running as a separate server. Its 5.54 GiB file was tested separately from the main speed benchmark.

Nemotron 3.5 Lightning produced the fastest output in the main benchmark at 193.7 tokens/s, though it needed five snake attempts. Qwen3.8-27B has strong vendor-reported coding results for its original checkpoint; our tested quant needed two attempts for each browser task.

ModelTested quantFile sizeVRAM at 32k; Bonsai at 16k4k promptGenerationLargest allocation loadedBest for
Gemma 4 26B-A4B itQ4_K_M15.64 GiB17,430 MiB4,506 tok/s159.8 tok/s192kOverall tested pick
Ternary Bonsai 2 27BPTQ1_05.54 GiB7,064 MiB peakNot measuredVisual test only16kSmallest footprint
Qwen3.8-27BAD-Q4_K_M15.94 GiB18,080 MiB1,296 tok/s41.8 tok/s96k fp16; 192k q8_0Coding candidate
Nemotron 3.5 Lightning 30B-A3BAD-IQ4_NL18.30 GiB18,342 MiB4,239 tok/s193.7 tok/s192kFast output and long chats
Gemma 4 12B itQ4_K_M6.87 GiB8,486 MiB2,991 tok/s83.7 tok/s192kLow VRAM use
Muse Glimmer 30BAD-Q4_K_M17.75 GiB18,116 MiB1,492 tok/s40.3 tok/s128k tested limitAgent-focused alternative

Prompt processing measures how quickly a model reads its input; generation measures output after that processing. Five rows report the main speed benchmark and memory at 32k. Bonsai reports peak memory during the separate 16k coding test. A loaded context allocation does not establish answer quality across that window.

For newer cards, see our RTX 4090 comparison and RTX 5090 benchmark. Our measurements use standalone inference runtimes; Atomic Chat performance depends on its selected backend and settings.

How We Benchmarked

We ran the main speed and long-context benchmarks on a harness with an NVIDIA GeForce RTX 3090 with 24,576 MiB of VRAM, NVIDIA driver 580.65.06, and the official CUDA 12.8 build of llama.cpp b10988. The binary contained native sm_86 kernels for the Ampere GPU.

The main configuration used full GPU offload, one parallel slot, and an fp16 KV cache. Qwen's additional q8_0 allocation tests are labeled separately. For speed, llama-bench processed a 4,096-token prompt and generated 128 new tokens. Bonsai 2 joined the separate visual coding test on the Prism b10709 fork.

The long-context test used a document at approximately 8k, 32k, and 128k tokens. Facts were placed near the beginning, middle, and end, with another question requiring information from two locations. Every model that loaded the 128k test found all four answers and retained all four required topics in its summary.

Best Local LLMs for RTX 3090

Gemma 4 26B-A4B it: Best Overall

Gemma 4 26B-A4B it is a 25.2B-parameter mixture-of-experts model that activates about 3.8B parameters for each token. It routes work through eight of 128 experts plus a shared expert, giving it much higher generation speed than a dense model of similar total size. The original model accepts text and images, supports configurable reasoning and function calls, and has a native 256k context window.

The model developer reports these scores for the original checkpoint. They were not measured in our RTX 3090 GGUF test:

BenchmarkScore
MMLU Pro
Academic knowledge
82.6%
GPQA Diamond
Expert science
82.3%
LiveCodeBench v6
Competitive coding
77.1%
MRCR v2, 8 needle at 128k
Context retrieval
44.1%

Source.

That architecture suits the RTX 3090 particularly well. The tested model generated 159.8 tokens/s and processed a 4k prompt at 4,506 tokens/s. It also loaded a 192k fp16 context allocation in 20,789 MiB, leaving about 3.7 GiB of VRAM free on our test card.

Tested fileResult
RepositoryAtomicChat/gemma-4-26B-A4B-it-GGUF
Exact GGUFgemma-4-26B-A4B-it-Q4_K_M.gguf
File size15.64 GiB
VRAM at 32k17,430 MiB
4k prompt / generation4,506 / 159.8 tok/s
128k document59.3 s to first token; facts 4/4; summary topics 4/4
Largest fp16 allocation loaded192k; 20,789 MiB

The main tradeoff is long-prompt latency. A 128k document took 59.3 seconds to produce the first token, so working with very large documents on this card comes with a noticeable delay before the response begins.

Verdict: Start with Gemma 4 26B-A4B for fast output and room for longer documents. It is our overall pick among the tested builds; Gemma 12B leaves more VRAM free.

Ternary Bonsai 2 27B: Smallest Memory Footprint

Ternary Bonsai 2 27B is a heavily compressed version of Qwen3.8-27B that uses ternary weights. Prism describes a 1.72-bit effective ternary representation; the deployed PTQ1_0 packing uses about 1.75 bits per weight.

As a result, the 27B model fits into a file of just 5.54 GiB and in our test used about 7,064 MiB of VRAM on the RTX 3090. Its memory footprint is comparable to much smaller models.

Prism reports these scores using EvalScope and vLLM on H100 with thinking enabled. Our RTX 3090 coding results follow in a separate table:

BenchmarkScore
MMLU-Redux
Academic knowledge
89.09
AIME 2026
Competition math
95.83
LiveCodeBench
Competitive coding
90.07
IFBench, prompt-loose
Instruction following
74.00
BFCL v3
Function calling
74.92

Source.

Tested fileResult
Repositoryprism-ml/Ternary-Bonsai-2-27B-gguf
Exact GGUFTernary-Bonsai-2-27B-PTQ1_0.gguf
File size5.54 GiB
Peak VRAM7,064 MiB
Five bouncing balls1 attempt; 175.7 s
Self-playing snake2 attempts; 237.9 s
Tested context16k

Verdict: Try Ternary Bonsai 2 if the smaller download and measured memory use justify running a separate Prism server. Our coding test does not establish that its reasoning quality matches every larger 27B build. See the Bonsai 2 setup guide for the required runtime.

Qwen3.8-27B: Coding Candidate

Qwen3.8-27B is a dense 27.32B-parameter model built for coding, research, and general assistance. Its architecture combines Gated DeltaNet layers with conventional attention layers. The model accepts text, images, and video.

The model developer reports these scores for the original checkpoint. They were not measured in our RTX 3090 GGUF test:

BenchmarkScore
GPQA Diamond
Expert science
89.2
Humanity's Last Exam
Expert questions
30.8
SWE-bench Pro
Harder engineering
61.7
Terminal-Bench 2.1
Terminal agents
73.0
LiveCodeBench v6
Competitive coding
90.3

Source.

Qwen's original checkpoint has strong published coding results, but those evaluations do not rank these six quantized files under one harness. Our AD-Q4_K_M build generated 41.8 tokens/s and needed two attempts on each browser task. Gemma 12B and Muse completed both tasks on their first attempt.

Tested fileResult
RepositoryAtomicChat/Qwen3.8-27B-GGUF
Exact GGUFQwen3.8-27B-AD-Q4_K_M.gguf
File size15.94 GiB
VRAM at 32k18,080 MiB
4k prompt / generation1,296 / 41.8 tok/s
Largest fp16 allocation loaded96k; 22,239 MiB
Largest q8_0 KV allocation loaded192k; 23,447 MiB

The fp16 KV cache is the other limit. Qwen loaded at 96k but ran out of memory at 128k. Switching both KV cache types to q8_0 reduced memory enough to load 128k in 20,951 MiB and 192k in 23,447 MiB.

Verdict: Consider Qwen for coding at moderate context lengths, then test it on your own codebase. Its published checkpoint evaluations cover larger software tasks than the two browser exercises here.

NVIDIA Nemotron 3.5 Lightning: Best for Long Chats

NVIDIA Nemotron 3.5 Lightning 30B-A3B combines Mamba-2, attention, and mixture-of-experts layers. It stores roughly 30B parameters and activates about 3B per token. NVIDIA designed it for agent workflows, tool calls, structured output, and long-running tasks.

NVIDIA reports these scores for its BF16 baseline. They are separate from our RTX 3090 GGUF measurements:

BenchmarkScore
MMLU Pro
Academic knowledge
81.94
GPQA Diamond, no tools
Expert science
75.44
SWE-bench Verified
Software engineering
51.56
IFBench, loose
Instruction following
71.88
AA-LCR
Long-context reasoning
52.00

Source.

Nemotron produced 193.7 tokens/s on the RTX 3090, the fastest result in our main speed benchmark. Its hybrid architecture also keeps context memory small: moving from 32k to 192k increased measured VRAM by only 1,119 MiB. The 192k allocation used 19,461 MiB, leaving more than 5 GB available.

Tested fileResult
RepositoryAtomicChat/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF
Exact GGUFNVIDIA-Nemotron-3.5-Lightning-30B-A3B-AD-IQ4_NL.gguf
File size18.30 GiB
VRAM at 32k18,342 MiB
4k prompt / generation4,239 / 193.7 tok/s
128k document43.5 s to first token; facts 4/4; summary topics 4/4
Largest fp16 allocation loaded192k; 19,461 MiB

Its measured output speed and compact context allocation make Nemotron worth testing for long text conversations. These experiments did not measure RAG accuracy or autonomous agent reliability.

Verdict: Choose Nemotron when fast output is your priority. It had the shortest summed request time for snake despite requiring five attempts, so allow for the extra feedback between attempts.

Gemma 4 12B it: Best for Low VRAM Use

Gemma 4 12B it is an 11.95B-parameter dense model from the Gemma family. It uses the same mixture of local sliding-window and periodic global attention as the larger Gemma 4 models, with configurable reasoning and a native 256k context window. The original checkpoint accepts text, images, audio, and video and generates text. Our GGUF tests used text only; they do not establish multimodal support in every runtime.

The model developer reports these scores for the original checkpoint. They were not measured in our RTX 3090 GGUF test:

BenchmarkScore
MMLU Pro
Academic knowledge
77.2%
GPQA Diamond
Expert science
78.8%
LiveCodeBench v6
Competitive coding
72.0%
MRCR v2, 8 needle at 128k
Context retrieval
43.4%

Source.

Gemma 4 12B is relatively lightweight, and in our testing on an RTX 3090 the Q4_K_M checkpoint only used 8,486 MiB at 32k tokens, which leaves about 15.7 GiB of the card free for other tasks. Despite the smaller memory footprint, it wasn't the fastest, generating at 83.7 tokens/s, but that experience is still fast enough for fluid local chat.

Tested fileResult
RepositoryAtomicChat/gemma-4-12B-it-GGUF
Exact GGUFgemma-4-12b-it-Q4_K_M.gguf
File size6.87 GiB
VRAM at 32k8,486 MiB
4k prompt / generation2,991 / 83.7 tok/s
128k document91.1 s to first token; facts 4/4; summary topics 4/4
Largest fp16 allocation loaded192k; 11,205 MiB

Gemma 4 12B took longer to read the 128k document than the two MoE models. It left more memory free, which can be useful when the GPU also runs other applications.

Verdict: Choose Gemma 4 12B for a smaller standard-runtime build. It used 8,486 MiB at 32k and completed both coding exercises on the first attempt.

Muse Glimmer 30B: Agent-Focused Alternative

Muse Glimmer 30B is Meta's dense local agent model, distilled from Muse Spark. It is trained for multi-step reasoning, tool use, and recovery after failed actions.

The model developer reports these scores for the original checkpoint. They were not measured in our RTX 3090 GGUF test:

BenchmarkScore
MCP Atlas
MCP tools
75.5
DeepSearch QA
Deep research
74.6
GPQA Diamond
Expert science
83.5
SWE-bench Pro
Harder engineering
51.2
Terminal-Bench 2.1
Terminal agents
51.7

Source.

Tested fileResult
RepositoryAtomicChat/Muse-Glimmer-30B-GGUF
Exact GGUFMuse-Glimmer-30B-AD-Q4_K_M.gguf
File size17.75 GiB
VRAM at 32k18,116 MiB
4k prompt / generation1,492 / 40.3 tok/s
128k document94.3 s to first token; facts 4/4; summary topics 4/4
Largest allocation loaded128k tested limit; 19,459 MiB

Muse is the slowest generator in the roundup by a narrow margin and takes the longest to ingest the full document. The reason to choose it is the model's training for tool-driven, multi-step work rather than throughput.

Verdict: Try Muse if you want to evaluate Meta's agent-focused model locally. It completed both browser tasks on the first attempt; that result does not measure autonomous tool use.

Live Tests: Coding a Snake Game and Bouncing Balls

We gave six models two visual coding tasks: a self-playing snake game where the snake finds food and grows, and a physics simulation containing five balls moving under gravity with elastic collisions.

Each model had to return one self-contained HTML file. When code threw an error, we sent the browser error or a precise description of the broken behavior back to the model.

The coding tests ran September 19 to 21, 2026, on a separate RTX 3090 host with NVIDIA driver 590.48.01. They used a 16k context, temperature 1.0, top-p 0.95, top-k 64, and seed 42. Five models ran on llama.cpp b10988 with an 8,192-token output budget; Bonsai 2 used the Prism b10709 fork with a 16,384-token output budget. Qwen and Bonsai used medium reasoning, Muse high, the Gemmas enabled/default, and Nemotron disabled.

Total generation time sums the model request durations, including prompt processing and reasoning. Browser testing and feedback time are excluded. Bonsai's attempt counts and times start from its final 16,384-token output configuration; preliminary failed 8k-budget trials are excluded. The videos show the final working outputs. These two exercises do not establish an overall coding-quality ranking.

Coding a Self-Playing Snake

Only Gemma 4 12B and Muse Glimmer produced a working snake on the first attempt, but all six models eventually built a game in which the snake found food, grew, and increased its score.

Watch the RTX 3090 snake demo on YouTube

ModelAttemptsTotal generation time
Gemma 4 26B-A4B299.4 s
Nemotron 3.5 Lightning568.4 s
Qwen3.8-27B2185.6 s
Muse Glimmer 30B1150.0 s
Gemma 4 12B171.5 s
Ternary Bonsai 2 27B2237.9 s

Nemotron needed five attempts but still had the shortest total generation time. Bonsai 2's first game moved without errors but never ate the food. After receiving feedback to minimize Manhattan distance, its second version reached a score of 18 within 30 seconds.

Coding Five Bouncing Balls

Gemma 4 26B-A4B, Gemma 4 12B, Muse Glimmer, and Bonsai 2 completed the simulation on their first attempt. Qwen needed two attempts and Nemotron needed four.

Watch the RTX 3090 bouncing-balls demo on YouTube

ModelAttemptsTotal generation timePeak VRAM
Gemma 4 26B-A4B136.1 s17,114 MiB
Nemotron 3.5 Lightning441.0 s18,268 MiB
Qwen3.8-27B2182.4 s17,046 MiB
Muse Glimmer 30B1140.3 s17,898 MiB
Gemma 4 12B157.0 s8,220 MiB
Ternary Bonsai 2 27B1175.7 s7,064 MiB

Gemma 4 26B-A4B finished first at 36.1 seconds. Nemotron's early versions applied the collision impulse with the wrong sign, causing the balls to accelerate into one another and become trapped at the corners. The fourth version corrected the physics.

RTX 3090 vs 4090 vs 5090 for Local LLMs

The desktop RTX 3090 and RTX 4090 both have 24 GB of VRAM, giving them similar model-memory budgets. Actual fit also depends on runtime, context and GPU memory used elsewhere. However, the 4090 uses a newer architecture, allowing it to process large prompts more than twice as fast. The generation speed advantage is much smaller.

The RTX 5090 has 32 GB of VRAM, so it can fit larger models, and it is also faster at both generation and prompt processing.

The rows below compare the same named GGUF builds at the 4,096-token prompt / 128-token generation workload. The 4090 and 5090 values come from our linked GPU comparisons:

Model and metricRTX 3090RTX 4090RTX 5090
Qwen3.8-27B, 4k prompt1,296 tok/s3,004 tok/s3,830 tok/s
Qwen3.8-27B, generation41.8 tok/s48.5 tok/s74.2 tok/s
Gemma 4 26B-A4B, 4k prompt4,506 tok/s9,970 tok/s11,829 tok/s
Gemma 4 26B-A4B, generation159.8 tok/s194.1 tok/s248.6 tok/s

The two examples show the practical tradeoff: on the 3090, generation remained at 41.8 and 159.8 tokens/s, while processing large inputs took substantially longer.

For Qwen and Gemma 26B-A4B, the RTX 4090 processed the 4k prompt about 2.2 to 2.3 times faster. RTX 3090 generation throughput was about 14% and 18% lower, respectively. These percentages describe the two tested configurations, not every local LLM.

How to Run These Models in Atomic Chat

For Gemma, Qwen, Nemotron, and Muse:

  1. Install Atomic Chat and select Find optimal backend.
  2. Open Models and search for the repository shown in the selected model's card.
  3. Open Download Options and choose the exact GGUF filename from the table.
  4. Start with a 32k context allocation and full GPU offload. Increase the context when a specific task needs it.

If the GGUF is already on your computer, import it through Settings → Model Providers → llama.cpp → Import. Read the full self-hosted LLM setup guide.

Qwen needs a compressed KV cache above 96k on this card. The equivalent direct llama.cpp command for 128k is:

llama-server -m /path/to/Qwen3.8-27B-AD-Q4_K_M.gguf -ngl 999 -c 131072 \
--parallel 1 --cache-type-k q8_0 --cache-type-v q8_0

The model loaded at 128k with 20,951 MiB allocated in our test. Use a smaller context allocation with fp16 for ordinary chats where 32k or 64k is enough.

For Ternary Bonsai 2:

Stock llama.cpp does not support this model. It requires the PrismML fork because the PTQ1_0 and PQ2_0 formats are not available in the standard runtime. The model also needs a large output limit. With an 8,192-token budget, our trials did not produce a complete working HTML file. The working limit in our test was 16,384 tokens.

In Atomic Chat, run Prism's llama-server separately and connect to it through the self-hosted llama.cpp server provider. The built-in Atomic Chat engine does not currently load this GGUF.

FAQ

What is the best LLM to run on an RTX 3090?

Gemma 4 26B-A4B Q4_K_M is our overall pick among the six tested builds. It generated 159.8 tokens/s and used 17,430 MiB at 32k. Start at 32k; the separate 192k allocation used 20,789 MiB.

Is 24 GB VRAM enough for a local LLM?

An RTX 3090 with 24 GB can run many current 12B to 32B models entirely on the GPU. The five models in our 32k benchmark used between 8,486 and 18,342 MiB. Ternary Bonsai 2 used 7,064 MiB in the separate 16k coding test.

Can an RTX 3090 run a 70B model?

At exactly four bits per parameter, 70B weights alone would occupy about 35 GB before format overhead and runtime memory. That exceeds 24 GB. Partial CPU offload or a supported multi-GPU setup can run larger files, with performance depending on the runtime and interconnect.

What is the best coding LLM for RTX 3090?

Qwen3.8-27B has strong vendor-reported coding results for its original checkpoint. Our AD-Q4_K_M file used 18,080 MiB at 32k and generated 41.8 tokens/s, but needed two attempts on each browser task. Gemma 12B and Muse passed both first time. These exercises do not establish a winner for a large codebase.

Is a used RTX 3090 still good for local AI in 2026?

The RTX 3090 remains useful for local AI with 24 GB of VRAM. In our matched Qwen and Gemma 26B-A4B examples, the 4090 processed a 4k prompt about 2.2 to 2.3 times faster. The 3090's generation throughput was about 14% and 18% lower, respectively. Those figures describe these two configurations; used-card condition and price require a separate buying decision.

Does dual RTX 3090 double usable VRAM?

Two RTX 3090 cards have 48 GB of physical VRAM in total. A compatible inference runtime can split weights and work across them, but usable capacity depends on its split strategy and overhead. Communication between the cards can limit speed; two cards do not behave like one 48 GB GPU.

Best Local AI Video Generators Compared on Quality, Speed and VRAM

Best Local AI Video Generators Compared on Quality, Speed and VRAM

Compare local AI video generators on quality, speed and VRAM, with RTX 4090 benchmarks, example clips and an Atomic Chat setup guide.

10/9/26

16 min

Best Local LLMs for 12GB VRAM in 2026

Best Local LLMs for 12GB VRAM in 2026

Eight local LLMs for 12GB VRAM: exact GGUF files, RTX 3080 Ti speed and memory figures, plus Snake and physics tests on an RTX 4070.

9/30/26

15 min

Claude Sonnet 5.5 Alternatives Compared

Claude Sonnet 5.5 Alternatives Compared

Compare Claude Sonnet 5.5 with Opus, Fable and GPT-6 Astra, then choose a local Qwen, Ornith or Bonsai model for your hardware.

9/29/26

13 min

Jev 1.13: Can You Run It Locally? Laya Setup Guide

Jev 1.13: Can You Run It Locally? Laya Setup Guide

Can you run Jev 1.13 locally? Learn how it works, see our Jev vs Laya Tetris demo, and set up Laya as an independent local alternative.

9/25/26

12 min