Qwen logo
Qwen3.6 27B MTP
27B
No models found
NewestParameters
Reasoning
Code
Multilingual
Web
Vision
Thinking
Tools
Audio
Visuals
Embedding
K2 Horizon 375B-A23B

K2 Horizon 375B-A23B from IFM: 512K context, sparse weights and server setup. Atomic Chat support is not yet verified.

Thinking
Tools
Reasoning
Code
K2 Horizon MoVA-36B-A4B

K2 Horizon MoVA-36B-A4B from IFM: 512K context, sparse weights and server setup. Atomic Chat support is not yet verified.

Thinking
Tools
Reasoning
Code
K2 Horizon 32B

K2 Horizon 32B from IFM: 512K context, Stage 1 dense weights and server setup. Atomic Chat support is not yet verified.

Thinking
Tools
Reasoning
Code
K2 Horizon 7B

K2 Horizon 7B from IFM: 512K context, dense weights and server setup. Atomic Chat support is not yet verified.

Thinking
Tools
Reasoning
Code
K2 Horizon 3.7B

K2 Horizon 3.7B from IFM: 512K context, dense weights and server setup. Atomic Chat support is not yet verified.

Thinking
Tools
Reasoning
Code
K2 Horizon 0.9B

K2 Horizon 0.9B from IFM: 128K context, dense weights and server setup. Atomic Chat support is not yet verified.

Thinking
Tools
Reasoning
Code
Atria Dawn Preview

Atria Dawn Preview: 256K text-only agent model from Shanghai AI Lab. Compare weights and server deployment; Atomic Chat support is not yet verified.

Thinking
Tools
Reasoning
Code
GLM-5.3-Flash

GLM-5.3-Flash is a 320B MoE with 18B active parameters, the first natively multimodal model in the GLM-5 series, released under MIT.

Thinking
Tools
Vision
Reasoning
Code
Qwen3.8-Flash-Next

Qwen3.8-Flash-Next carries 125B parameters with 6B active, plus a 51B n-gram embedding. An experimental preview of the Qwen4 architecture.

Thinking
Tools
Vision
Reasoning
Code
Ornith-1.5-35B-A3B

Ornith-1.5-35B-A3B is a 36B MoE that activates about 3B parameters per token and beats models ten times its size on coding.

Thinking
Tools
Vision
Reasoning
Code
Ornith-1.5-9B

Ornith-1.5-9B is a 9B agentic coding model with a 256K context that runs on 8 GB. We quantized the GGUF builds ourselves.

Thinking
Tools
Vision
Reasoning
Code
Qwen3.8-27B

Qwen3.8-27B is a dense 27.8B multimodal model with a 256K context that fits on a 24 GB GPU at 4-bit. Run it locally in Atomic Chat.

Thinking
Tools
Vision
Reasoning
Code
NVIDIA-Nemotron-3.5-Lightning-30B-A3B

NVIDIA Nemotron 3.5 Lightning is a 30B Mamba-2 hybrid MoE with 3B active parameters and a context window that reaches 1M tokens.

Thinking
Tools
Reasoning
Code
Multilingual
Muse-Glimmer-30B

Muse Glimmer 30B is Meta’s agent model for consumer hardware: 30B dense with a vision encoder, quantized to fit a 24 GB card.

Thinking
Tools
Vision
Reasoning
Code
Ling-3.0-flash

Ling-3.0-flash is a 124B hybrid-linear MoE that activates only 5.1B parameters per token and keeps pace with far larger models on software engineering.

Thinking
Tools
Reasoning
Code
Multilingual
DeepSeek-V4-Flash-0731

DeepSeek V4 Flash 0731 is the official V4 Flash release: 284B MoE with 13B active, a 1M context and MIT weights.

Thinking
Tools
Reasoning
Code
Multilingual
Qwen3.8-2.4T-A95B

Qwen3.8-2.4T-A95B is the downloadable version of Qwen 3.8 Max: a 2.4T MoE with 95B active parameters and a 256K context.

Thinking
Tools
Reasoning
Code
Multilingual
OpenELM-1_1B-Instruct

Apple’s 1.1B instruction-tuned OpenELM model, built with layer-wise scaling for efficient on-device English text generation.

Reasoning
SmolLM2-135M-Instruct

A 135M-parameter instruction-tuned LLM from Hugging Face’s SmolLM2 family, small enough to run on CPU and on-device.

Reasoning
MiniMax-M2.5

A 229B-parameter (10B active) MoE model from MiniMax built for agentic coding, tool use, and search, with a 200K context window.

Thinking
Tools
Reasoning
Code
Web
Granite-4.0-H-Small

IBM’s 32B (9B active) hybrid Mamba-2/MoE instruct model with 128K context, strong tool-calling and multilingual support, under Apache 2.0.

Reasoning
Code
Multilingual
Tools
NVIDIA-Nemotron-Nano-9B-v2

A 9B hybrid Mamba2-Transformer reasoning model from NVIDIA with toggleable thinking, 128K context, and tool calling.

Thinking
Reasoning
Code
Tools
Multilingual
SmolLM3-3B

A fully open 3B reasoning model from Hugging Face with dual-mode thinking, six native languages, tool calling, and 128K context.

Reasoning
Multilingual
Tools
Thinking
Kimi-K2-Instruct

A 1T-parameter MoE chat model from Moonshot AI with 32B active parameters, built for agentic tool use and strong coding.

Tools
Code
Reasoning
Multilingual
Phi-4-mini-instruct

A 3.8B-parameter open instruct model from Microsoft’s Phi-4 family with 128K context, strong math and reasoning, and function calling.

Reasoning
Code
Multilingual
Tools
Llama-3.3-Nemotron-Super-49B-v1.5

A 49B reasoning and chat LLM from NVIDIA, distilled from Llama-3.3-70B via Neural Architecture Search with a 128K context.

Thinking
Reasoning
Code
Tools
Multilingual
DeepSeek-R1-0528

A 671B-parameter MoE reasoning model from DeepSeek with 37B active params, MIT-licensed, strong at math, code, and long chain-of-thought.

Thinking
Reasoning
Code
Tools
Qwen3-30B-A3B-Instruct-2507

A 30.5B-parameter (3.3B active) MoE instruct model from Alibaba’s Qwen3 series with 256K context and strong reasoning, coding, and tool use.

Reasoning
Code
Multilingual
Tools
Phi-4

A 14B open model from Microsoft Research tuned for math, reasoning, and code, competitive with much larger LLMs.

Reasoning
Code
DeepSeek-Coder-V2-Lite-Instruct

A 16B Mixture-of-Experts code model from DeepSeek AI with 2.4B active params, 128K context, and support for 338 programming languages.

Code
Reasoning
Tools
Mistral-7B-Instruct-v0.2

Mistral-7B-Instruct-v0.2 is a 7B instruction-tuned model with a 32K context window. Explore its hardware needs, GGUF builds and local setup.

Reasoning
Code
DeepSeek-V3-0324

A 671B-parameter Mixture-of-Experts LLM from DeepSeek-AI (37B active) with a 128K context, strong coding and improved function calling.

Reasoning
Code
Multilingual
Tools
Qwen3-4B-Thinking-2507

A 4B reasoning-focused LLM from Alibaba’s Qwen3 series that always thinks step by step, with a 256K context and strong math, coding, and agentic scores.

Thinking
Reasoning
Code
Tools
Multilingual
Phi-3.5-mini-instruct

A 3.8B dense instruction-tuned LLM from Microsoft’s Phi-3.5 family with a 128K context window and multilingual support.

Reasoning
Code
Multilingual
Qwen2.5-72B-Instruct

A 72.7B instruction-tuned LLM from Alibaba’s Qwen2.5 series with strong coding, math and multilingual ability across 29+ languages.

Reasoning
Code
Multilingual
Tools
Qwen3-8B

An 8.2B dense LLM from Alibaba’s Qwen3 series with switchable thinking mode, strong reasoning, coding, and 100+ language support.

Thinking
Tools
Reasoning
Code
Multilingual
Qwen3-235B-A22B

A 235B mixture-of-experts LLM from Alibaba’s Qwen3 series that activates 22B parameters and switches between thinking and non-thinking modes.

Thinking
Reasoning
Code
Multilingual
Tools
Qwen2.5-Coder-7B-Instruct

A 7.6B code-specialized LLM from Alibaba’s Qwen2.5-Coder series, tuned for code generation, reasoning, and fixing.

Code
Reasoning
Tools
Qwen2.5-32B-Instruct

A 32.5B instruction-tuned LLM from Alibaba’s Qwen2.5 series with strong coding, math, and 29+ language support.

Reasoning
Code
Multilingual
Tools
Qwen2.5-14B-Instruct

A 14.7B instruction-tuned LLM from Alibaba’s Qwen2.5 series with strong coding, math, and 29+ language support.

Reasoning
Code
Multilingual
Tools
Qwen2.5-7B-Instruct

A 7.61B instruction-tuned LLM from Alibaba’s Qwen2.5 series with strong coding, math, and multilingual ability across 29+ languages.

Reasoning
Code
Multilingual
Tools
Qwen2.5-Coder-32B-Instruct

A 32.5B code-specialized LLM from Alibaba’s Qwen2.5-Coder series with open-model state-of-the-art coding ability and 128K context.

Code
Reasoning
Tools
Multilingual
Kimi-K2-Instruct-0905

Run Kimi-K2-Instruct-0905, a 1T-parameter MoE model, locally in Atomic Chat. Private, offline agentic tool use, no API keys, no cloud. Free.

Tools
Reasoning
Code
GLM-4.7-Flash

GLM-4.7-Flash is Z.ai’s 30B-A3B MoE under MIT: 31.2B parameters, about 3B active per token, and 2-bit GGUF builds fit in 12 GB.

Tools
Thinking
Reasoning
Code
MiniMax-M2.7

Run MiniMax-M2.7, a 229B MoE model, locally in Atomic Chat. Private, offline reasoning and coding, no API keys, no cloud, no limits. Free.

Thinking
Tools
Reasoning
Code
gpt-oss-120b

gpt-oss-120b is OpenAI’s open-weight MoE: 117B parameters, 5.1B active, sized for a single 80 GB GPU. Run it locally, free, with Atomic Chat.

Tools
Thinking
Reasoning
Code
gpt-oss-20b

gpt-oss-20b is OpenAI’s open-weight 21B MoE model with 3.6B active parameters, with its MoE weights post-trained in MXFP4 to run within 16 GB of memory.

Tools
Thinking
Reasoning
Code
gemma-3-270m

Gemma 3 270M is the smallest Gemma 3 model: 268M parameters, 32K context, trained on 6 trillion tokens. The Q8_0 GGUF is a 0.29 GB file.

Reasoning
Multilingual
gemma-3-1b-it

Gemma 3 1B is Google’s smallest Gemma 3 model: 1B parameters, a 32K context window and 2 trillion training tokens. The Q4_K_M GGUF build is 0.81 GB.

Reasoning
Code
Multilingual
DeepSeek-V3.2

Run DeepSeek-V3.2, a 685B MoE model, locally in Atomic Chat. Private, offline reasoning and coding, no API keys, no cloud, no limits. Free.

Thinking
Tools
Reasoning
Code
Multilingual
DeepSeek-R1

Run DeepSeek-R1, a 671B reasoning model, locally with Atomic Chat. Private, offline chain-of-thought with no cloud and no limits. Free.

Thinking
Reasoning
Code
Multilingual
Llama-3.2-3B-Instruct

Llama 3.2 3B Instruct is Meta’s 3.21B multilingual chat model with a 128K context. The 4-bit GGUF build is a 2 GB file.

Tools
Reasoning
Code
Multilingual
Llama-3.1-8B-Instruct

Llama-3.1-8B-Instruct is Meta’s 8B multilingual chat model with a 128K context and tool calling. The Q4_K_M build is 4.92 GB.

Tools
Reasoning
Code
Multilingual
Qwen3-Coder-30B-A3B-Instruct

Qwen3-Coder-30B-A3B-Instruct is Qwen’s agentic coding MoE: 30.5B parameters, 3.3B active per token, 262K native context. Apache 2.0.

Tools
Reasoning
Code
Multilingual
Qwen3-30B-A3B

Qwen3-30B-A3B is a 30.5B MoE model that activates 3.3B params per token, with a switchable thinking mode and 100+ languages.

Tools
Thinking
Reasoning
Code
Multilingual
Qwen3-14B

Qwen3-14B is a 14.8B dense model from the Qwen team with a switchable thinking mode and a 131K YaRN context window. The Q4_K_M build is 9.00 GB.

Tools
Thinking
Reasoning
Code
Multilingual
Qwen3-32B

Qwen3-32B is a 32.8B-parameter Apache 2.0 model from Alibaba’s Qwen team, with a switchable thinking mode and 100+ languages. Q4_K_M fits a 24 GB card.

Tools
Thinking
Reasoning
Code
Multilingual
Nex-N2-mini

Nex-N2-mini is a 35.1B agent model from Nex-AGI, post-trained on Qwen3.5-35B-A3B-Base for coding and tool use. Apache 2.0, GGUF builds from 9.78 GB.

Tools
Thinking
Vision
Reasoning
Code
LocateAnything-3B

LocateAnything-3B is NVIDIA’s visual grounding model for locating objects, GUI elements and text with bounding boxes or points.

Thinking
Vision
Reasoning
Code
diffusiongemma-26B-A4B-it

DiffusionGemma 26B A4B is Google DeepMind’s diffusion text model: a Gemma 4 MoE that denoises 256-token blocks in parallel, 15-20 tokens per pass.

Thinking
Vision
Audio
Reasoning
Code
GLM-5.1

GLM-5.1 is Z.ai’s 753.9B-parameter MIT-licensed flagship for agentic engineering, topping SWE-Bench Pro at 58.4. Run it locally in Atomic Chat.

Tools
Thinking
Reasoning
Code
gemma-4-E2B-it

Gemma 4 E2B is Google DeepMind’s smallest Gemma 4 model: 2.3B effective parameters, text, image and audio input, and a 128K context window.

Thinking
Embedding
Vision
Audio
Reasoning
gemma-4-12B-it

Gemma 4 12B is Google DeepMind’s encoder-free multimodal model: text, image, audio and video input, a 256K context window, Apache 2.0.

Thinking
Embedding
Vision
Audio
Reasoning
FastContext-1.0-4B-RL

Run FastContext-1.0-4B-RL, a compact long-context model, locally in Atomic Chat. Offline and private, no API keys. Free.

Tools
Code
Multilingual
VibeThinker-3B

VibeThinker-3B is WeiboAI’s 3.1B reasoning model: 76.4 on IMO-AnswerBench, in range of 671B to 1T models, and 96.1% on unseen LeetCode contests.

Thinking
Reasoning
Code
Nex-N2-Pro

Nex-N2-Pro is a 396.8B mixture-of-experts agent model from Nex-AGI, post-trained on Qwen3.5-397B-A17B. Apache 2.0, GGUF builds from 82 GB.

Tools
Thinking
Vision
Reasoning
Code
North-Mini-Code-1.0

North-Mini-Code-1.0 is Cohere’s 30B-A3B MoE for agentic coding and terminal tasks: 256K context, Apache 2.0.

Tools
Thinking
Embedding
Reasoning
Code
MiMo-V2.5-Pro

MiMo-V2.5-Pro is Xiaomi’s 1.02T-parameter MoE with 42B active and a 1M-token context, MIT licensed. The smallest GGUF build is 304 GB.

Tools
Thinking
Reasoning
Code
Multilingual
GLM-5.2

GLM-5.2 is Z.ai’s 753.3B-parameter flagship with a solid 1M-token context, released under MIT with no regional limits.

Thinking
Reasoning
Code
MiniMax-M3

Run MiniMax-M3, a 427B MoE model, locally in Atomic Chat. Private, offline long-context reasoning, no API keys, no limits. Free.

Thinking
Vision
Reasoning
Code
Kimi-K2.7-Code

Run Kimi-K2.7-Code, a 1T-parameter coding model, locally with Atomic Chat. Offline agentic coding, fully private. Download free.

Tools
Thinking
Vision
Reasoning
Code
DeepSeek-V4-Flash

DeepSeek-V4-Flash is a 284B MoE model with 13B active parameters and a 1M token context, released by DeepSeek under the MIT license.

Thinking
Reasoning
Code
Multilingual
Qwen3.6-35B-A3B

Qwen3.6-35B-A3B is Qwen’s open MoE vision model: 35B parameters, 3B active, 262K context, Apache 2.0. Low-bit GGUF builds fit in 12 GB.

Tools
Thinking
Embedding
Vision
Reasoning
Qwen3.6-27B

Qwen3.6-27B is a dense 27.8B vision-language model with a 262K context and thinking on by default. GGUF builds run from 9.39 GB on a 12 GB machine.

Tools
Thinking
Embedding
Vision
Reasoning
gemma-4-31B-it

Gemma 4 31B is Google DeepMind’s dense 30.7B open model: text and image input, a 256K-token context window, and an Apache 2.0 license.

Thinking
Embedding
Vision
Audio
Reasoning
gemma-4-26B-A4B-it

Google’s Gemma 4 MoE: 25.2B parameters with 3.8B active, 256K context and image input, Apache 2.0. Runs at the speed of a model near 4B.

Thinking
Embedding
Vision
Audio
Reasoning
WebWorld-8B

WebWorld-8B is a text-only Qwen world model that predicts the next web page state from the current state and an action. Apache 2.0.

Thinking
Reasoning
Web
DeepSeek-V4-Pro

DeepSeek-V4-Pro is a 1.6T-parameter MoE model, 49B active per token, with a 1M-token context under MIT. Run it locally, free, with Atomic Chat.

Thinking
Reasoning
Code
Multilingual
Qwen3.6-27B-MTP

Qwen3.6-27B-MTP is Qwen3.6-27B served with its Multi-Token Prediction head: the model drafts several tokens per step for faster local generation.

Thinking
Tools
Vision
Supertonic-3

Supertonic 3 is a 99M-parameter on-device TTS from Supertone that speaks 31 languages and runs fast on CPU, with about 400 MB of ONNX assets.

Audio
Multilingual
Sulphur-2-base

Sulphur 2 base is an uncensored video model built on LTX 2.3 that handles text-to-video and image-to-video natively, with ComfyUI workflows included.

Visuals
Dramabox

Dramabox is Resemble AI’s expressive TTS built on the LTX-2.3 audio branch: the prompt controls voice, emotion and delivery, with 10-second voice cloning.

Audio
MiniCPM-V 4.6

MiniCPM-V 4.6 is a 1.3B vision-language model from OpenBMB that reads images and video on-device, with vendor GGUF builds starting at 0.5 GB.

Vision
Anima

Anima is a 2B text-to-image model from CircleStone Labs and Comfy Org, trained on several million anime images for illustration and artistic work.

Visuals

No models match your search. Try removing a filter or widening the parameter range.

Choosing an open-source model to run locally

Every model in this catalog runs entirely on your own hardware — no API keys, no per-token billing and no data leaving your machine. That makes local models a good fit for private workloads, offline environments and high-volume tasks where a metered cloud API would get expensive. The trade-off is that you pick the hardware, so the right model depends as much on your machine as on the task.

The two numbers that matter most are parameter count and VRAM required. Parameter count is a rough proxy for capability — larger models reason and write better, but need more memory and run slower. VRAM required tells you whether a model fits on your GPU at all; a quantized build lowers that number, at a small cost to quality. Use the sidebar filters to narrow the list to models your machine can actually run before comparing anything else.

Match the model to the task

A model tuned for code completion behaves differently from one tuned for chat or vision, even at the same size. The Tasks filter groups models by what they were trained to do — text generation, image-to-text, text-to-image and more — so start there, then sort within the task by size or how recently the weights were updated.

Best models for general chat and reasoning

These general-purpose models balance answer quality against hardware cost. Each one runs comfortably on a recent laptop or a mid-range GPU.

ModelParametersVRAM requiredBest for
Qwen 3 8B
8B~8 GBEveryday chat on a laptop
Llama 3.1 8B
8B~8 GBBalanced reasoning and writing
Mistral Small 24B
24B~16 GBStronger reasoning, mid-range GPU
Qwen 3 32B
32B~24 GBHighest quality, desktop GPU

Best models for low-end hardware

If you're working on a CPU-only machine or a GPU with limited memory, these compact and quantized models stay responsive without a discrete graphics card.

ModelParametersVRAM requiredBest for
Qwen 3 1.7B
1.7BCPU onlyFast replies on any laptop
Llama 3.2 3B
3B< 8 GBLightweight assistant tasks
Mistral 7B (Q4)
7B~6 GBQuantized build for older GPUs

These picks are a starting point, not a ranking — the right model is the largest one that runs smoothly on your hardware for the task you care about. Use the filters above to explore the full catalog.

Frequently asked questions

Local models run on your own hardware, so there are no API keys, no usage limits and no data sent to a third party. Cloud APIs are easier to scale, but bill per token and require an internet connection — local models trade that convenience for privacy and predictable cost.

Check the "VRAM required" value against your GPU memory, or filter the catalog by it. If a model is larger than your GPU, a quantized build or a CPU-only model will still work — just more slowly.

Parameters are the learned weights inside a model. More parameters generally means better reasoning and writing, at the cost of more memory and slower generation. It's a rough guide to capability, not an exact score.

It depends on each model's license. Many are released under permissive licenses such as Apache 2.0, while others restrict commercial use. Always check the license listed on the model's page before shipping it in a product.

Yes. Once the weights are downloaded, a local model runs without any network connection — useful for air-gapped setups and for keeping sensitive data on-device.