This is a guide to the desktop RTX 5090. The RTX 5090 Laptop GPU has a different memory budget and is not part of this ranking.
Quick Answer
Qwen3.8-27B is our general-purpose pick for the RTX 5090. It generated 79.4 tokens/s and processed a 4k prompt at 5,100 tokens/s. If output speed matters most, Nemotron 3.5 Lightning reached 334.3 tokens/s in the tested AtomicChat build. Gemma 4 26B-A4B offers a strong balance of speed and long-document work: it generated 248.6 tokens/s and took 24.4 seconds to answer the first question after our 128k-token document. All six models answered the document questions correctly, so the useful differences were wait time and VRAM use. Every number on this page came off a real RTX 5090 running the model, not from a memory-bandwidth estimate.
Best Local LLMs for RTX 5090
| Model | Tested GGUF file | File (GiB) | VRAM at 32k (MiB) | 4k prompt (tok/s) | Generation (tok/s) | Best for |
|---|---|---|---|---|---|---|
| Qwen3.8-27B | Qwen3.8-27B-iMatrix-NVFP4-MTP.gguf | 15.95 | 18,048 | 5,100 | 79.4 | General-purpose use |
| Gemma 4 31B it | gemma-4-31B-it-Q4_K_M.gguf | 17.40 | 22,372 | 3,698 | 68.5 | Dense Gemma alternative |
| Gemma 4 26B-A4B it | gemma-4-26B-A4B-it-Q4_K_M.gguf | 15.64 | 17,702 | 11,829 | 248.6 | Fast long-document work |
| Gemma 4 12B it | gemma-4-12b-it-Q4_K_M.gguf | 6.87 | 8,758 | 8,659 | 145.4 | Lowest VRAM use |
| Muse Glimmer 30B | Muse-Glimmer-30B-AD-Q4_K_M.gguf | 17.75 | 18,392 | 4,387 | 73.9 | Reasoning-first assistant |
| Nemotron 3.5 Lightning 30B-A3B | NVIDIA-Nemotron-3.5-Lightning-30B-A3B-AD-IQ4_NL.gguf | 18.30 | 18,612 | 12,047 | 334.3 | Fast output and first token |
1. Qwen3.8-27B
Qwen3.8-27B is our all-round pick for a desktop RTX 5090. The dense model combines conventional attention with Gated DeltaNet layers and supports a native 262k-token context. Its strong performance in coding benchmarks makes it a great starting point for development work.
| Model fact | Value |
|---|---|
| Developer | Qwen / Alibaba |
| Release | August 14, 2026 |
| Parameters and architecture | 27.32B, dense, 64 hybrid Gated DeltaNet / attention layers |
| Native context | 262,144 tokens |
| Base-model inputs | Text, images, and video |
| Thinking | Switchable, reasoning depth can be adjusted |
| Recommended tested build | cdiamond/Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF then Qwen3.8-27B-iMatrix-NVFP4-MTP.gguf, 15.95 GiB |
Released in August 2026, it is the 27B dense member of the Qwen3.8 family. The original model accepts text, images, and video and has switchable thinking. Our GGUF tests used text with thinking off. For this card we recommend the cdiamond community file: 70.8% of its weights are NVFP4, with the rest in Q5_K, Q6_K, and Q8_0.
| Benchmark | Result |
|---|---|
GPQA Diamond Expert science | 89.2 |
SWE-bench Pro Harder engineering | 61.7 |
Terminal-Bench 2.1 Terminal agents | 73.0 |
LiveCodeBench v6 Competitive coding | 90.3 |
The results above are for the original Qwen model.
| Measurement | AtomicChat AD-Q4_K_M | AtomicChat AD-Q5_K_M | cdiamond NVFP4 |
|---|---|---|---|
| Exact filename | Qwen3.8-27B-AD-Q4_K_M.gguf | Qwen3.8-27B-AD-Q5_K_M.gguf | Qwen3.8-27B-iMatrix-NVFP4-MTP.gguf |
| File size | 15.94 GiB | 18.84 GiB | 15.95 GiB |
| GPU memory at 32k context | 18,356 MiB | not measured | 18,048 MiB |
| 4k prompt processing | 3,830 tok/s | 3,540 tok/s | 5,100 tok/s |
| 128-token generation | 74.2 tok/s | 69.2 tok/s | 79.4 tok/s |
Verdict: Choose the exact cdiamond NVFP4 GGUF for Qwen3.8-27B on Blackwell. It matched the AD-Q4 file size and processed the prompt faster. The 128k document took 62.6 seconds to produce its first token and used 24,773 MiB with the tested AtomicChat AD-Q4 build; do not assume those long-context figures transfer unchanged to the cdiamond file.
2. Gemma 4 31B it
Gemma 4 31B it is Google's largest dense Gemma 4 model. It is a useful alternative if you prefer Gemma's output style, but it is also the most memory-hungry of our six at long context.
| Model fact | Value |
|---|---|
| Developer | Google DeepMind |
| Release | April 2026 |
| Parameters and architecture | 30.7B, dense, 60 layers |
| Native context | 262,144 tokens |
| Original-model inputs | Text and images |
| Tested build | AtomicChat/gemma-4-31B-it-GGUF then gemma-4-31B-it-Q4_K_M.gguf, 17.40 GiB |
Gemma 4 31B uses local sliding-window and global attention across its layers. Unlike the 26B-A4B model below, it uses all its language-model weights to produce each token. The table below shows how it performs in benchmarks. These are official Google numbers, and Google doesn't specify a per-row thinking setting:
| Benchmark | Result |
|---|---|
GPQA Diamond Expert science | 84.3% |
LiveCodeBench v6 Competitive coding | 80.0% |
MRCR v2, eight needles at 128k | 66.4% |
Our RTX 5090 test:
| Measurement | Result |
|---|---|
| GPU memory at 32k context | 22,372 MiB |
| 4k prompt processing | 3,698 tok/s |
| 128-token generation | 68.5 tok/s |
| First token after our 128k document | 85.0 s |
| GPU memory with the 128k document | 30,361 MiB |
| Document facts / summary topics at 128k | 4/4 / 4/4 |
Verdict: The dense Gemma alternative works at 128k on the RTX 5090, but it leaves only 2,246 MiB of the card's 32,607 MiB free in our test. Choose a lighter model if long-document response time or room for other GPU tasks matters more. Among our six it also holds the best published long-context score, 66.4% on MRCR v2 at 128k.
3. Gemma 4 26B-A4B it
If you need more generation speed at the small cost of reasoning performance, and want to stick to the Gemma family, choose Gemma 4 26B-A4B it. Its mixture-of-experts design stores about 25.2B parameters but activates roughly 4B for each token, helping it generate much faster than the dense 31B model.
| Model fact | Value |
|---|---|
| Developer | Google DeepMind |
| Release | April 2026 |
| Parameters and architecture | 25.2B stored, about 4B active per token; 30-layer MoE |
| Native context | 262,144 tokens |
| Original-model inputs | Text and images |
| Tested build | AtomicChat/gemma-4-26B-A4B-it-GGUF then gemma-4-26B-A4B-it-Q4_K_M.gguf, 15.64 GiB |
The smaller active parameter count reduces computation per generated token, which is what makes generation fast here. Its lighter memory use at long context comes from a different place: a smaller file plus the way its attention layers store the KV cache. In our setup, this model used 19,831 MiB at 128k against 30,361 MiB for Gemma 4 31B. And here's how this model performs in the benchmarks, according to Google's official measurements for the original full precision file:
| Benchmark | Result |
|---|---|
GPQA Diamond Expert science | 82.3% |
LiveCodeBench v6 Competitive coding | 77.1% |
MRCR v2, eight needles at 128k | 44.1% |
The table below displays results during our RTX 5090 test:
| Measurement | Result |
|---|---|
| GPU memory at 32k context | 17,702 MiB |
| 4k prompt processing | 11,829 tok/s |
| 128-token generation | 248.6 tok/s |
| First token after our 128k document | 24.4 s |
| GPU memory with the 128k document | 19,831 MiB |
| Document facts / summary topics at 128k | 4/4 / 4/4 |
Verdict: Choose Gemma 4 26B-A4B when you want fast generation and a 128k document without consuming most of the RTX 5090's VRAM. On Google's own MRCR v2 test at 128k it scores 44.1% against 66.4% for the dense 31B, so pick the 31B when the answer matters more than the wait.
4. Gemma 4 12B it
Gemma 4 12B it is the smallest model in our lineup of the best models to run on an RTX 5090. As the smallest file, it makes the most sense when you need an AI model to run alongside other apps that also need GPU memory. The GGUF file we tested occupied 8,758 MiB at a 32k allocation and 10,447 MiB with our 128k document.
| Model fact | Value |
|---|---|
| Developer | Google DeepMind |
| Release | April 2026 |
| Parameters and architecture | About 11.9B, dense |
| Native context | 262,144 tokens |
| Original-model inputs | Text and images |
| Tested build | AtomicChat/gemma-4-12B-it-GGUF then gemma-4-12b-it-Q4_K_M.gguf, 6.87 GiB |
This is a much lighter model than the two larger Gemmas we've tested, yet it completed all of the planted-fact questions and preserved all four required topics in our 128k summary.
| Benchmark | Result |
|---|---|
GPQA Diamond Expert science | 78.8% |
LiveCodeBench v6 Competitive coding | 72.0% |
MRCR v2, eight needles at 128k | 43.4% |
And here are the results from our test run:
| Measurement | Result |
|---|---|
| GPU memory at 32k context | 8,758 MiB |
| 4k prompt processing | 8,659 tok/s |
| 128-token generation | 145.4 tok/s |
| First token after our 128k document | 30.1 s |
| GPU memory with the 128k document | 10,447 MiB |
| Document facts / summary topics at 128k | 4/4 / 4/4 |
Verdict: Choose Gemma 4 12B when you need a capable AI model that generates tokens fast and offers great performance for the size, or when you need additional GPU headroom. In our test, Gemma 4 12B used less than half the 128k VRAM use of Qwen or Gemma 4 31B.
5. Muse Glimmer 30B
Muse Glimmer 30B is Meta's agentic model with a dedicated perception encoder. Meta built it for tool use, multi-step reasoning and failure recovery on consumer hardware, and lists OpenClaw and Hermes Agent among the scaffolds it works with. Both of those already run on top of Atomic Chat's local endpoint.
| Model fact | Value |
|---|---|
| Developer | Meta |
| Release | August 2026 |
| Parameters and architecture | About 29.6B including the perception encoder, dense, 52 layers |
| Native context | 131,072 tokens and above |
| Original-model inputs | Text and images through the perception encoder |
| Tested build | AtomicChat/Muse-Glimmer-30B-GGUF then Muse-Glimmer-30B-AD-Q4_K_M.gguf, 17.75 GiB |
Muse Glimmer continues to use reasoning tokens even at a low reasoning setting. Give it enough output budget for both thought and final answer. It also found all four facts at 8k, 32k, and 128k. It uses a denser tokenizer, which encoded our 128k-class document in 107,856 tokens against 130,557 for the Gemma models. Here's how this model performs in the benchmarks, according to Meta:
| Benchmark | Result |
|---|---|
GPQA Diamond (AA) | 83.5% |
HLE Text (AA) | 22.0% |
SWE-bench Pro Harder engineering | 51.2% |
SWE-bench Verified Software engineering | 76.0% |
Terminal-Bench 2.1 with terminus2 | 51.7% |
And here's how it fared in our test on an RTX 5090 graphics card:
| Measurement | Result |
|---|---|
| GPU memory at 32k context | 18,392 MiB |
| 4k prompt processing | 4,387 tok/s |
| 128-token generation | 73.9 tok/s |
| First token after our long document | 31.8 s |
| GPU memory with the long document | 19,752 MiB |
| Document facts / summary topics | 4/4 / 4/4 at all three lengths |
Verdict: Muse Glimmer is the pick when you run agents locally. On Meta's own numbers it reaches 76.0% on SWE-bench Verified and 51.7% on Terminal-Bench 2.1, well ahead of the other agentic candidate here. Its denser tokenizer also fits more text into the same context length.
6. NVIDIA Nemotron 3.5 Lightning 30B-A3B
NVIDIA Nemotron 3.5 Lightning 30B-A3B is a great pick to run on an RTX 5090 when you're optimizing for inference speed. Its hybrid Mamba-2, attention, and mixture-of-experts architecture stores about 32.9B parameters but activates roughly 3B per token, and allows the model to output tokens quickly, especially for its parameter count.
| Model fact | Value |
|---|---|
| Developer | NVIDIA |
| Release | August 2026 |
| Parameters and architecture | About 32.9B stored, 3B active; hybrid Mamba-2 / attention / MoE |
| Model context limit | Up to 1M tokens |
| Original-model inputs | Text |
| Tested build | AtomicChat/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF then NVIDIA-Nemotron-3.5-Lightning-30B-A3B-AD-IQ4_NL.gguf, 18.30 GiB |
The architecture is built to reduce per-token work and handle long sequences. In our tested build, it generated 334.3 tokens/s and returned the first token after the 128k document in 14.9 seconds. Those figures are for ordinary generation without speculative decoding. Here's how the model performs in benchmarks, according to NVIDIA:
| Benchmark | Result |
|---|---|
GPQA Diamond (no tools) | 75.44% |
SWE-bench Verified Software engineering | 51.56% |
Terminal-Bench 2.1 Terminal agents | 24.58% |
And here are the results of our RTX 5090 test:
| Measurement | Result |
|---|---|
| GPU memory at 32k context | 18,612 MiB |
| 4k prompt processing | 12,047 tok/s |
| 128-token generation | 334.3 tok/s |
| First token after our 128k document | 14.9 s |
| GPU memory with the 128k document | 19,367 MiB |
| Document facts / summary topics at 128k | 4/4 / 4/4 |
Verdict: Choose Nemotron 3.5 Lightning for fast output and long-document response time. On published agentic benchmarks it trails Muse Glimmer, so pick it for speed rather than for task success.
How the Six Models Compare
Only GPQA Diamond is published for all six models, so the table below marks which numbers come from the same benchmark and which do not. Each vendor runs its own harness: Qwen scored SWE-bench Pro with the Claude Code harness at a 256k context, while Meta scored it with its own setup. Treat cross-vendor rows as indicative, not as a head-to-head.
| Model | GPQA Diamond | LiveCodeBench v6 | SWE-bench Pro | SWE-bench Verified | Terminal-Bench 2.1 | MRCR v2 at 128k | Source |
|---|---|---|---|---|---|---|---|
| Qwen3.8-27B | 89.2 | 90.3 | 61.7 | not published | 73.0 | not published | Qwen |
| Gemma 4 31B it | 84.3 | 80.0 | not published | not published | not published | 66.4 | |
| Gemma 4 26B-A4B it | 82.3 | 77.1 | not published | not published | not published | 44.1 | |
| Gemma 4 12B it | 78.8 | 72.0 | not published | not published | not published | 43.4 | |
| Muse Glimmer 30B | 83.5 | not published | 51.2 | 76.0 | 51.7 | not published | Meta |
| Nemotron 3.5 Lightning | 75.44 | not published | not published | 51.56 | 24.58 | not published | NVIDIA |
Two comparisons here are like for like. On MRCR v2 at 128k the three Gemma models line up 66.4, 44.1 and 43.4, so the dense 31B is far better at finding things in a long document than its faster siblings. On SWE-bench Verified and Terminal-Bench 2.1, Muse Glimmer reaches 76.0 and 51.7 against 51.56 and 24.58 for Nemotron, so the two agentic candidates are not close.
Picking Context Over Parameters
The RTX 5090's 32 GB lets you spend memory on a bigger model or on a longer context, and the two compete for the same card. Our 128k measurements show how wide that spread is.
| Model | VRAM with the 128k document | Free on a 32,607 MiB card | Fits in 24 GB |
|---|---|---|---|
| Gemma 4 12B it | 10,447 MiB | 22,160 MiB | Yes |
| Nemotron 3.5 Lightning | 19,367 MiB | 13,240 MiB | Yes |
| Muse Glimmer 30B | 19,752 MiB | 12,855 MiB | Yes |
| Gemma 4 26B-A4B it | 19,831 MiB | 12,776 MiB | Yes |
| Qwen3.8-27B AD-Q4 | 24,773 MiB | 7,834 MiB | No |
| Gemma 4 31B it | 30,361 MiB | 2,246 MiB | No |
Gemma 4 12B is the clearest case for trading parameters for room. It held our full 128k document in 10,447 MiB, less than half of what Qwen3.8-27B needed and a third of Gemma 4 31B, and it still answered all four planted-fact questions. That leaves 22 GB for a second model, an image model, or the rest of your desktop.
The trade is real, though, and Google's own numbers price it. On MRCR v2 at 128k the 12B scores 43.4% against 66.4% for the dense 31B. A long context you can load is not the same as a long context the model reads well. Load the biggest model whose answers you trust, then give the remaining memory to context.
Recommended file type for Blackwell architecture
The RTX 5090 uses NVIDIA's Blackwell architecture and has FP4 tensor cores. In our llama.cpp test, the NVFP4 GGUF files processed 4k prompts 28 to 58% faster than the standard GGUF build of the same model. Generation improved too, by 2 to 10%. For a long input, faster prompt processing reduces the wait before generation starts.
| Model | Ordinary GGUF: 4k prompt / generation | Community NVFP4 GGUF: 4k prompt / generation | Change |
|---|---|---|---|
| Qwen3.8-27B | 3,830 / 74.2 tok/s | 5,100 / 79.4 tok/s | +33% prompt; +7% generation |
| Gemma 4 31B | 3,698 / 68.5 tok/s | 5,826 / 69.7 tok/s | +58% prompt; +2% generation |
| Nemotron 3.5 Lightning | 12,047 / 334.3 tok/s | 15,374 / 367.2 tok/s | +28% prompt; +10% generation |
The exact community NVFP4 files tested were Qwen3.8-27B-iMatrix-NVFP4-MTP.gguf, gemma-4-31B-it-NVFP4-turbo-NVFP4.gguf, and Nemotron-3.5-Lightning-30B-A3B-NVFP4.gguf, respectively.
- The Qwen file has 70.8% of its weights in NVFP4 and the rest in Q5_K/Q6_K/Q8_0.
- The Gemma and Nemotron files are 95.4% and 95.6% NVFP4 by weight.
These are measured results for these specific files on llama.cpp b10988. We did not isolate how much of the gain comes from the FP4 tensor cores and how much from the smaller weights. NVIDIA's own card for Nemotron lists the GeForce RTX 5090 compute path as still to be confirmed, so treat the numbers as what these builds did on our card rather than as a property of the format.
How we tested these models
The main throughput and document tests used a desktop RTX 5090 with 32,607 MiB of VRAM, Ubuntu 24.04, and the official CUDA 12.8 build of llama.cpp b10988. We used full GPU offload, one parallel slot, fp16 KV cache, and automatic flash-attention selection. Throughput comes from llama-bench -p 512,4096 -n 128 -r 2; the tables show the 4k-prompt and 128-token-generation results. The Gemma files here are AtomicChat GGUF builds, not the QAT files from the earlier run. Qwen AD-Q5 speed and the Muse long-document run used a second pod with a different driver, so those two rows are not strictly comparable with the rest.
One caveat on the throughput table: the six models do not all run the same quantization recipe. Three Gemma files are Q4_K_M, Muse Glimmer is AD-Q4_K_M, Nemotron is AD-IQ4_NL, and Qwen is the mixed cdiamond NVFP4 file. Part of the spread between rows is the build, not the model.
Coding a Self-Playing Snake
We asked each model to build a snake game that plays itself, eats food, and grows. This test used vLLM with NVFP4 builds: AtomicChat/gemma-4-12B-it-NVFP4, AtomicChat/gemma-4-26B-A4B-it-NVFP4 and AtomicChat/gemma-4-31B-it-NVFP4 for the Gemma models, and nvidia/Qwen3.8-27B-NVFP4, nvidia/Muse-Glimmer-30B-NVFP4 and nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 for the rest.
| Model | Attempts | Total generation time |
|---|---|---|
| Qwen3.8-27B | 2 | 97.5 s |
| Gemma 4 31B | 1 | 47.9 s |
| Gemma 4 26B-A4B | 1 | 27.1 s |
| Gemma 4 12B | 3 | 134.0 s |
| Muse Glimmer 30B | 2 | 169.8 s |
| Nemotron 3.5 Lightning | 5 | 45.7 s |
Every model eventually completed the test, creating a snake that chases food, grows, and increases its score as the game plays.
Gemma 4 26B-A4B and 31B were able to build the game on the first attempt, and the 26B-A4B was also quickest at 27.1 seconds. Nemotron needed five attempts, but its total generation time was 45.7 seconds. Its repair requests used thinking off after the initial response exhausted the output budget.
Coding Five Bouncing Balls
We also asked each model to make a page with five balls affected by gravity and collisions. This test used the same vLLM and NVFP4 setup as the snake task.
| Model | Attempts | Total generation time |
|---|---|---|
| Qwen3.8-27B | 2 | 88.5 s |
| Gemma 4 31B | 1 | 54.5 s |
| Gemma 4 26B-A4B | 1 | 32.9 s |
| Gemma 4 12B | 4 | 161.7 s |
| Muse Glimmer 30B | 2 | 142.1 s |
| Nemotron 3.5 Lightning | 3 | 31.8 s |
As in the snake test, only Gemma 4 26B-A4B and 31B worked on the first attempt. Despite needing three attempts, Nemotron had the shortest generation time at 31.8 seconds.
Finding facts in a 128k-Token Document
We gave each model the same source document at roughly 8k, 32k, and 128k tokens. We planted facts near the beginning, middle, and end, then asked a fourth question requiring facts from two separate places.
We also checked whether a summary retained four specified topics. The document's token count differs by tokenizer, even though the source text is the same.
| Model | Facts at 8k / 32k / 128k | Summary topics at 128k | Input tokens processed at longest length | Time to first token | VRAM at longest length |
|---|---|---|---|---|---|
| Nemotron 3.5 Lightning | 4/4 · 4/4 · 4/4 | 4/4 | 129,628 | 14.9 s | 19,367 MiB |
| Gemma 4 26B-A4B | 4/4 · 4/4 · 4/4 | 4/4 | 130,557 | 24.4 s | 19,831 MiB |
| Gemma 4 12B | 4/4 · 4/4 · 4/4 | 4/4 | 130,557 | 30.1 s | 10,447 MiB |
| Muse Glimmer 30B | 4/4 · 4/4 · 4/4 | 4/4 | 107,856 | 31.8 s | 19,752 MiB |
| Qwen3.8-27B, AtomicChat AD-Q4 | 4/4 · 4/4 · 4/4 | 4/4 | 126,820 | 62.6 s | 24,773 MiB |
| Gemma 4 31B | 4/4 · 4/4 · 4/4 | 4/4 | 130,557 | 85.0 s | 30,361 MiB |
All six passed this particular retrieval and summary test at every length, so it separates them on wait time rather than on accuracy. For a harder read of long-context quality, use the published MRCR v2 scores in the comparison table above. Two rows also did less work than the others: Muse Glimmer processed 107,856 tokens against 130,557 for the Gemma models, and its run used the second pod.
At this length, Gemma 4 31B used 30,361 MiB and Qwen3.8-27B AD-Q4 used 24,773 MiB. Both fit on our 32,607 MiB RTX 5090, but exceed the nominal 24 GB (24,576 MiB) VRAM budget of an RTX 4090 or 3090. The other four models used 10,447 to 19,831 MiB in this test, below 24 GB. These are measured RTX 5090 memory footprints; the table does not show their speed on older GPUs. Qwen's 128k result is for the AtomicChat AD-Q4 file, not the cdiamond NVFP4 file recommended above.
How to Run Offline AI Models in Atomic Chat
1. Install Atomic Chat. When choosing a backend, select Find optimal backend for your GPU.
2. For an AtomicChat-built model in the table, open Models, search the linked AtomicChat repository, choose the exact filename shown in its card under Download Options, and download it.
3. For our recommended Qwen build, download Qwen3.8-27B-iMatrix-NVFP4-MTP.gguf from cdiamond's repository. In Atomic Chat, go to Settings, then Model Providers, then llama.cpp, then Import, and select that local GGUF file.
4. Select Use this model, open a new chat, and start with a 32k context allocation. Increase the context when a task requires it, watching GPU memory use. If the imported NVFP4 file does not load, update or select a compatible llama.cpp backend.
Atomic Chat's own self-hosted LLM guide walks through the app's model download and import controls.
FAQ
What is the best LLM to run on an RTX 5090?
For a balance of speed and performance, Qwen3.8-27B is the best AI model to run on an RTX 5090. If your priority is the shortest task completion time or the highest generation speed, choose Nemotron 3.5 Lightning.
Is 32 GB VRAM enough for a local LLM?
Yes. 32 GB of VRAM is enough to run most 12B to 33B models in 4-bit quantization entirely on the GPU, with room for a long context. Larger models or very long context windows may need a smaller quantization or partial CPU offload.
Can an RTX 5090 run a 70B model?
Not at 4-bit. A 70B Q4 file generally exceeds the RTX 5090's 32 GB VRAM before context memory is added. You can reach a 70B at a lower bit rate or with part of the model on the CPU, but both cost quality or speed.
What is the best coding LLM for RTX 5090?
Qwen3.8-27B holds the best published coding scores in our lineup, at 61.7 on SWE-bench Pro and 90.3 on LiveCodeBench v6. In our own two coding tasks it was slower to a working result than the Gemma models: Gemma 4 26B-A4B built both the snake and the balls page on the first attempt, in 27.1 and 32.9 seconds, against two attempts and 97.5 and 88.5 seconds for Qwen. Pick Qwen for harder problems and Gemma 4 26B-A4B when you want a working answer quickly.
How much power and system RAM does an RTX 5090 setup need?
NVIDIA specifies 575 W total graphics power for the RTX 5090 and a 1,000 W system power requirement for the reference configuration. System RAM needs depend on file size, context, and whether any model layers run on the CPU. Our benchmark host had 117 GB of RAM; that is not a requirement for these GPU-resident files.

