Quick Answer
Gemma 4 26B-A4B Q4_K_M is our overall pick among the six builds tested on a desktop RTX 3090. It generated 159.8 tokens/s and used 17,430 MiB at 32k. Start with gemma-4-26B-A4B-it-Q4_K_M.gguf and a 32k context allocation. Its largest tested fp16 allocation was 192k, using 20,789 MiB.
For a straightforward low-memory setup, Gemma 4 12B Q4_K_M used 8,486 MiB at 32k and completed both coding tasks on its first attempt. Ternary Bonsai 2 27B had the smallest measured footprint, 7,064 MiB peak at 16k, but requires the Prism llama.cpp fork running as a separate server. Its 5.54 GiB file was tested separately from the main speed benchmark.
Nemotron 3.5 Lightning produced the fastest output in the main benchmark at 193.7 tokens/s, though it needed five snake attempts. Qwen3.8-27B has strong vendor-reported coding results for its original checkpoint; our tested quant needed two attempts for each browser task.
| Model | Tested quant | File size | VRAM at 32k; Bonsai at 16k | 4k prompt | Generation | Largest allocation loaded | Best for |
|---|---|---|---|---|---|---|---|
| Gemma 4 26B-A4B it | Q4_K_M | 15.64 GiB | 17,430 MiB | 4,506 tok/s | 159.8 tok/s | 192k | Overall tested pick |
| Ternary Bonsai 2 27B | PTQ1_0 | 5.54 GiB | 7,064 MiB peak | Not measured | Visual test only | 16k | Smallest footprint |
| Qwen3.8-27B | AD-Q4_K_M | 15.94 GiB | 18,080 MiB | 1,296 tok/s | 41.8 tok/s | 96k fp16; 192k q8_0 | Coding candidate |
| Nemotron 3.5 Lightning 30B-A3B | AD-IQ4_NL | 18.30 GiB | 18,342 MiB | 4,239 tok/s | 193.7 tok/s | 192k | Fast output and long chats |
| Gemma 4 12B it | Q4_K_M | 6.87 GiB | 8,486 MiB | 2,991 tok/s | 83.7 tok/s | 192k | Low VRAM use |
| Muse Glimmer 30B | AD-Q4_K_M | 17.75 GiB | 18,116 MiB | 1,492 tok/s | 40.3 tok/s | 128k tested limit | Agent-focused alternative |
Prompt processing measures how quickly a model reads its input; generation measures output after that processing. Five rows report the main speed benchmark and memory at 32k. Bonsai reports peak memory during the separate 16k coding test. A loaded context allocation does not establish answer quality across that window.
For newer cards, see our RTX 4090 comparison and RTX 5090 benchmark. Our measurements use standalone inference runtimes; Atomic Chat performance depends on its selected backend and settings.
How We Benchmarked
We ran the main speed and long-context benchmarks on a harness with an NVIDIA GeForce RTX 3090 with 24,576 MiB of VRAM, NVIDIA driver 580.65.06, and the official CUDA 12.8 build of llama.cpp b10988. The binary contained native sm_86 kernels for the Ampere GPU.
The main configuration used full GPU offload, one parallel slot, and an fp16 KV cache. Qwen's additional q8_0 allocation tests are labeled separately. For speed, llama-bench processed a 4,096-token prompt and generated 128 new tokens. Bonsai 2 joined the separate visual coding test on the Prism b10709 fork.
The long-context test used a document at approximately 8k, 32k, and 128k tokens. Facts were placed near the beginning, middle, and end, with another question requiring information from two locations. Every model that loaded the 128k test found all four answers and retained all four required topics in its summary.
Best Local LLMs for RTX 3090
Gemma 4 26B-A4B it: Best Overall
Gemma 4 26B-A4B it is a 25.2B-parameter mixture-of-experts model that activates about 3.8B parameters for each token. It routes work through eight of 128 experts plus a shared expert, giving it much higher generation speed than a dense model of similar total size. The original model accepts text and images, supports configurable reasoning and function calls, and has a native 256k context window.
The model developer reports these scores for the original checkpoint. They were not measured in our RTX 3090 GGUF test:
| Benchmark | Score |
|---|---|
MMLU Pro Academic knowledge | 82.6% |
GPQA Diamond Expert science | 82.3% |
LiveCodeBench v6 Competitive coding | 77.1% |
MRCR v2, 8 needle at 128k Context retrieval | 44.1% |
That architecture suits the RTX 3090 particularly well. The tested model generated 159.8 tokens/s and processed a 4k prompt at 4,506 tokens/s. It also loaded a 192k fp16 context allocation in 20,789 MiB, leaving about 3.7 GiB of VRAM free on our test card.
| Tested file | Result |
|---|---|
| Repository | AtomicChat/gemma-4-26B-A4B-it-GGUF |
| Exact GGUF | gemma-4-26B-A4B-it-Q4_K_M.gguf |
| File size | 15.64 GiB |
| VRAM at 32k | 17,430 MiB |
| 4k prompt / generation | 4,506 / 159.8 tok/s |
| 128k document | 59.3 s to first token; facts 4/4; summary topics 4/4 |
| Largest fp16 allocation loaded | 192k; 20,789 MiB |
The main tradeoff is long-prompt latency. A 128k document took 59.3 seconds to produce the first token, so working with very large documents on this card comes with a noticeable delay before the response begins.
Verdict: Start with Gemma 4 26B-A4B for fast output and room for longer documents. It is our overall pick among the tested builds; Gemma 12B leaves more VRAM free.
Ternary Bonsai 2 27B: Smallest Memory Footprint
Ternary Bonsai 2 27B is a heavily compressed version of Qwen3.8-27B that uses ternary weights. Prism describes a 1.72-bit effective ternary representation; the deployed PTQ1_0 packing uses about 1.75 bits per weight.
As a result, the 27B model fits into a file of just 5.54 GiB and in our test used about 7,064 MiB of VRAM on the RTX 3090. Its memory footprint is comparable to much smaller models.
Prism reports these scores using EvalScope and vLLM on H100 with thinking enabled. Our RTX 3090 coding results follow in a separate table:
| Benchmark | Score |
|---|---|
MMLU-Redux Academic knowledge | 89.09 |
AIME 2026 Competition math | 95.83 |
LiveCodeBench Competitive coding | 90.07 |
IFBench, prompt-loose Instruction following | 74.00 |
BFCL v3 Function calling | 74.92 |
| Tested file | Result |
|---|---|
| Repository | prism-ml/Ternary-Bonsai-2-27B-gguf |
| Exact GGUF | Ternary-Bonsai-2-27B-PTQ1_0.gguf |
| File size | 5.54 GiB |
| Peak VRAM | 7,064 MiB |
| Five bouncing balls | 1 attempt; 175.7 s |
| Self-playing snake | 2 attempts; 237.9 s |
| Tested context | 16k |
Verdict: Try Ternary Bonsai 2 if the smaller download and measured memory use justify running a separate Prism server. Our coding test does not establish that its reasoning quality matches every larger 27B build. See the Bonsai 2 setup guide for the required runtime.
Qwen3.8-27B: Coding Candidate
Qwen3.8-27B is a dense 27.32B-parameter model built for coding, research, and general assistance. Its architecture combines Gated DeltaNet layers with conventional attention layers. The model accepts text, images, and video.
The model developer reports these scores for the original checkpoint. They were not measured in our RTX 3090 GGUF test:
| Benchmark | Score |
|---|---|
GPQA Diamond Expert science | 89.2 |
Humanity's Last Exam Expert questions | 30.8 |
SWE-bench Pro Harder engineering | 61.7 |
Terminal-Bench 2.1 Terminal agents | 73.0 |
LiveCodeBench v6 Competitive coding | 90.3 |
Qwen's original checkpoint has strong published coding results, but those evaluations do not rank these six quantized files under one harness. Our AD-Q4_K_M build generated 41.8 tokens/s and needed two attempts on each browser task. Gemma 12B and Muse completed both tasks on their first attempt.
| Tested file | Result |
|---|---|
| Repository | AtomicChat/Qwen3.8-27B-GGUF |
| Exact GGUF | Qwen3.8-27B-AD-Q4_K_M.gguf |
| File size | 15.94 GiB |
| VRAM at 32k | 18,080 MiB |
| 4k prompt / generation | 1,296 / 41.8 tok/s |
| Largest fp16 allocation loaded | 96k; 22,239 MiB |
| Largest q8_0 KV allocation loaded | 192k; 23,447 MiB |
The fp16 KV cache is the other limit. Qwen loaded at 96k but ran out of memory at 128k. Switching both KV cache types to q8_0 reduced memory enough to load 128k in 20,951 MiB and 192k in 23,447 MiB.
Verdict: Consider Qwen for coding at moderate context lengths, then test it on your own codebase. Its published checkpoint evaluations cover larger software tasks than the two browser exercises here.
NVIDIA Nemotron 3.5 Lightning: Best for Long Chats
NVIDIA Nemotron 3.5 Lightning 30B-A3B combines Mamba-2, attention, and mixture-of-experts layers. It stores roughly 30B parameters and activates about 3B per token. NVIDIA designed it for agent workflows, tool calls, structured output, and long-running tasks.
NVIDIA reports these scores for its BF16 baseline. They are separate from our RTX 3090 GGUF measurements:
| Benchmark | Score |
|---|---|
MMLU Pro Academic knowledge | 81.94 |
GPQA Diamond, no tools Expert science | 75.44 |
SWE-bench Verified Software engineering | 51.56 |
IFBench, loose Instruction following | 71.88 |
AA-LCR Long-context reasoning | 52.00 |
Nemotron produced 193.7 tokens/s on the RTX 3090, the fastest result in our main speed benchmark. Its hybrid architecture also keeps context memory small: moving from 32k to 192k increased measured VRAM by only 1,119 MiB. The 192k allocation used 19,461 MiB, leaving more than 5 GB available.
| Tested file | Result |
|---|---|
| Repository | AtomicChat/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF |
| Exact GGUF | NVIDIA-Nemotron-3.5-Lightning-30B-A3B-AD-IQ4_NL.gguf |
| File size | 18.30 GiB |
| VRAM at 32k | 18,342 MiB |
| 4k prompt / generation | 4,239 / 193.7 tok/s |
| 128k document | 43.5 s to first token; facts 4/4; summary topics 4/4 |
| Largest fp16 allocation loaded | 192k; 19,461 MiB |
Its measured output speed and compact context allocation make Nemotron worth testing for long text conversations. These experiments did not measure RAG accuracy or autonomous agent reliability.
Verdict: Choose Nemotron when fast output is your priority. It had the shortest summed request time for snake despite requiring five attempts, so allow for the extra feedback between attempts.
Gemma 4 12B it: Best for Low VRAM Use
Gemma 4 12B it is an 11.95B-parameter dense model from the Gemma family. It uses the same mixture of local sliding-window and periodic global attention as the larger Gemma 4 models, with configurable reasoning and a native 256k context window. The original checkpoint accepts text, images, audio, and video and generates text. Our GGUF tests used text only; they do not establish multimodal support in every runtime.
The model developer reports these scores for the original checkpoint. They were not measured in our RTX 3090 GGUF test:
| Benchmark | Score |
|---|---|
MMLU Pro Academic knowledge | 77.2% |
GPQA Diamond Expert science | 78.8% |
LiveCodeBench v6 Competitive coding | 72.0% |
MRCR v2, 8 needle at 128k Context retrieval | 43.4% |
Gemma 4 12B is relatively lightweight, and in our testing on an RTX 3090 the Q4_K_M checkpoint only used 8,486 MiB at 32k tokens, which leaves about 15.7 GiB of the card free for other tasks. Despite the smaller memory footprint, it wasn't the fastest, generating at 83.7 tokens/s, but that experience is still fast enough for fluid local chat.
| Tested file | Result |
|---|---|
| Repository | AtomicChat/gemma-4-12B-it-GGUF |
| Exact GGUF | gemma-4-12b-it-Q4_K_M.gguf |
| File size | 6.87 GiB |
| VRAM at 32k | 8,486 MiB |
| 4k prompt / generation | 2,991 / 83.7 tok/s |
| 128k document | 91.1 s to first token; facts 4/4; summary topics 4/4 |
| Largest fp16 allocation loaded | 192k; 11,205 MiB |
Gemma 4 12B took longer to read the 128k document than the two MoE models. It left more memory free, which can be useful when the GPU also runs other applications.
Verdict: Choose Gemma 4 12B for a smaller standard-runtime build. It used 8,486 MiB at 32k and completed both coding exercises on the first attempt.
Muse Glimmer 30B: Agent-Focused Alternative
Muse Glimmer 30B is Meta's dense local agent model, distilled from Muse Spark. It is trained for multi-step reasoning, tool use, and recovery after failed actions.
The model developer reports these scores for the original checkpoint. They were not measured in our RTX 3090 GGUF test:
| Benchmark | Score |
|---|---|
MCP Atlas MCP tools | 75.5 |
DeepSearch QA Deep research | 74.6 |
GPQA Diamond Expert science | 83.5 |
SWE-bench Pro Harder engineering | 51.2 |
Terminal-Bench 2.1 Terminal agents | 51.7 |
| Tested file | Result |
|---|---|
| Repository | AtomicChat/Muse-Glimmer-30B-GGUF |
| Exact GGUF | Muse-Glimmer-30B-AD-Q4_K_M.gguf |
| File size | 17.75 GiB |
| VRAM at 32k | 18,116 MiB |
| 4k prompt / generation | 1,492 / 40.3 tok/s |
| 128k document | 94.3 s to first token; facts 4/4; summary topics 4/4 |
| Largest allocation loaded | 128k tested limit; 19,459 MiB |
Muse is the slowest generator in the roundup by a narrow margin and takes the longest to ingest the full document. The reason to choose it is the model's training for tool-driven, multi-step work rather than throughput.
Verdict: Try Muse if you want to evaluate Meta's agent-focused model locally. It completed both browser tasks on the first attempt; that result does not measure autonomous tool use.
Live Tests: Coding a Snake Game and Bouncing Balls
We gave six models two visual coding tasks: a self-playing snake game where the snake finds food and grows, and a physics simulation containing five balls moving under gravity with elastic collisions.
Each model had to return one self-contained HTML file. When code threw an error, we sent the browser error or a precise description of the broken behavior back to the model.
The coding tests ran September 19 to 21, 2026, on a separate RTX 3090 host with NVIDIA driver 590.48.01. They used a 16k context, temperature 1.0, top-p 0.95, top-k 64, and seed 42. Five models ran on llama.cpp b10988 with an 8,192-token output budget; Bonsai 2 used the Prism b10709 fork with a 16,384-token output budget. Qwen and Bonsai used medium reasoning, Muse high, the Gemmas enabled/default, and Nemotron disabled.
Total generation time sums the model request durations, including prompt processing and reasoning. Browser testing and feedback time are excluded. Bonsai's attempt counts and times start from its final 16,384-token output configuration; preliminary failed 8k-budget trials are excluded. The videos show the final working outputs. These two exercises do not establish an overall coding-quality ranking.
Coding a Self-Playing Snake
Only Gemma 4 12B and Muse Glimmer produced a working snake on the first attempt, but all six models eventually built a game in which the snake found food, grew, and increased its score.
Watch the RTX 3090 snake demo on YouTube
| Model | Attempts | Total generation time |
|---|---|---|
| Gemma 4 26B-A4B | 2 | 99.4 s |
| Nemotron 3.5 Lightning | 5 | 68.4 s |
| Qwen3.8-27B | 2 | 185.6 s |
| Muse Glimmer 30B | 1 | 150.0 s |
| Gemma 4 12B | 1 | 71.5 s |
| Ternary Bonsai 2 27B | 2 | 237.9 s |
Nemotron needed five attempts but still had the shortest total generation time. Bonsai 2's first game moved without errors but never ate the food. After receiving feedback to minimize Manhattan distance, its second version reached a score of 18 within 30 seconds.
Coding Five Bouncing Balls
Gemma 4 26B-A4B, Gemma 4 12B, Muse Glimmer, and Bonsai 2 completed the simulation on their first attempt. Qwen needed two attempts and Nemotron needed four.
Watch the RTX 3090 bouncing-balls demo on YouTube
| Model | Attempts | Total generation time | Peak VRAM |
|---|---|---|---|
| Gemma 4 26B-A4B | 1 | 36.1 s | 17,114 MiB |
| Nemotron 3.5 Lightning | 4 | 41.0 s | 18,268 MiB |
| Qwen3.8-27B | 2 | 182.4 s | 17,046 MiB |
| Muse Glimmer 30B | 1 | 140.3 s | 17,898 MiB |
| Gemma 4 12B | 1 | 57.0 s | 8,220 MiB |
| Ternary Bonsai 2 27B | 1 | 175.7 s | 7,064 MiB |
Gemma 4 26B-A4B finished first at 36.1 seconds. Nemotron's early versions applied the collision impulse with the wrong sign, causing the balls to accelerate into one another and become trapped at the corners. The fourth version corrected the physics.
RTX 3090 vs 4090 vs 5090 for Local LLMs
The desktop RTX 3090 and RTX 4090 both have 24 GB of VRAM, giving them similar model-memory budgets. Actual fit also depends on runtime, context and GPU memory used elsewhere. However, the 4090 uses a newer architecture, allowing it to process large prompts more than twice as fast. The generation speed advantage is much smaller.
The RTX 5090 has 32 GB of VRAM, so it can fit larger models, and it is also faster at both generation and prompt processing.
The rows below compare the same named GGUF builds at the 4,096-token prompt / 128-token generation workload. The 4090 and 5090 values come from our linked GPU comparisons:
| Model and metric | RTX 3090 | RTX 4090 | RTX 5090 |
|---|---|---|---|
| Qwen3.8-27B, 4k prompt | 1,296 tok/s | 3,004 tok/s | 3,830 tok/s |
| Qwen3.8-27B, generation | 41.8 tok/s | 48.5 tok/s | 74.2 tok/s |
| Gemma 4 26B-A4B, 4k prompt | 4,506 tok/s | 9,970 tok/s | 11,829 tok/s |
| Gemma 4 26B-A4B, generation | 159.8 tok/s | 194.1 tok/s | 248.6 tok/s |
The two examples show the practical tradeoff: on the 3090, generation remained at 41.8 and 159.8 tokens/s, while processing large inputs took substantially longer.
For Qwen and Gemma 26B-A4B, the RTX 4090 processed the 4k prompt about 2.2 to 2.3 times faster. RTX 3090 generation throughput was about 14% and 18% lower, respectively. These percentages describe the two tested configurations, not every local LLM.
How to Run These Models in Atomic Chat
For Gemma, Qwen, Nemotron, and Muse:
- Install Atomic Chat and select Find optimal backend.
- Open Models and search for the repository shown in the selected model's card.
- Open Download Options and choose the exact GGUF filename from the table.
- Start with a 32k context allocation and full GPU offload. Increase the context when a specific task needs it.
If the GGUF is already on your computer, import it through Settings → Model Providers → llama.cpp → Import. Read the full self-hosted LLM setup guide.
Qwen needs a compressed KV cache above 96k on this card. The equivalent direct llama.cpp command for 128k is:
llama-server -m /path/to/Qwen3.8-27B-AD-Q4_K_M.gguf -ngl 999 -c 131072 \ --parallel 1 --cache-type-k q8_0 --cache-type-v q8_0
The model loaded at 128k with 20,951 MiB allocated in our test. Use a smaller context allocation with fp16 for ordinary chats where 32k or 64k is enough.
For Ternary Bonsai 2:
Stock llama.cpp does not support this model. It requires the PrismML fork because the PTQ1_0 and PQ2_0 formats are not available in the standard runtime. The model also needs a large output limit. With an 8,192-token budget, our trials did not produce a complete working HTML file. The working limit in our test was 16,384 tokens.
In Atomic Chat, run Prism's llama-server separately and connect to it through the self-hosted llama.cpp server provider. The built-in Atomic Chat engine does not currently load this GGUF.
FAQ
What is the best LLM to run on an RTX 3090?
Gemma 4 26B-A4B Q4_K_M is our overall pick among the six tested builds. It generated 159.8 tokens/s and used 17,430 MiB at 32k. Start at 32k; the separate 192k allocation used 20,789 MiB.
Is 24 GB VRAM enough for a local LLM?
An RTX 3090 with 24 GB can run many current 12B to 32B models entirely on the GPU. The five models in our 32k benchmark used between 8,486 and 18,342 MiB. Ternary Bonsai 2 used 7,064 MiB in the separate 16k coding test.
Can an RTX 3090 run a 70B model?
At exactly four bits per parameter, 70B weights alone would occupy about 35 GB before format overhead and runtime memory. That exceeds 24 GB. Partial CPU offload or a supported multi-GPU setup can run larger files, with performance depending on the runtime and interconnect.
What is the best coding LLM for RTX 3090?
Qwen3.8-27B has strong vendor-reported coding results for its original checkpoint. Our AD-Q4_K_M file used 18,080 MiB at 32k and generated 41.8 tokens/s, but needed two attempts on each browser task. Gemma 12B and Muse passed both first time. These exercises do not establish a winner for a large codebase.
Is a used RTX 3090 still good for local AI in 2026?
The RTX 3090 remains useful for local AI with 24 GB of VRAM. In our matched Qwen and Gemma 26B-A4B examples, the 4090 processed a 4k prompt about 2.2 to 2.3 times faster. The 3090's generation throughput was about 14% and 18% lower, respectively. Those figures describe these two configurations; used-card condition and price require a separate buying decision.
Does dual RTX 3090 double usable VRAM?
Two RTX 3090 cards have 48 GB of physical VRAM in total. A compatible inference runtime can split weights and work across them, but usable capacity depends on its split strategy and overhead. Communication between the cards can limit speed; two cards do not behave like one 48 GB GPU.

