Nex-N2-mini

Updated
24.08.2026
Tools
Thinking
Vision
Reasoning
Code

Nex-N2-mini is a 35.1B agent model from Nex-AGI, post-trained on Qwen3.5-35B-A3B-Base for coding and tool use. Apache 2.0, GGUF builds from 9.78 GB.

At a glance

  • License: Apache 2.0
  • Parameters: 35.1B total
  • Context length: Not published on the Nex-AGI model card
  • Modalities: Not stated by Nex-AGI; the GGUF repo ships mmproj projector files next to the text weights
  • Minimum hardware: 12 GB memory with the IQ2_XXS GGUF, the smallest build at 9.78 GB

What is Nex-N2-mini?

Nex-N2-mini is a 35.1B-parameter open-weight agent model from Nex-AGI and the smaller of the two models in the Nex-N2 release. It is post-trained on Qwen3.5-35B-A3B-Base and built for real productivity work: agentic coding, deep research, tool calling and terminal execution. Nex-AGI published the weights on Hugging Face on June 4, 2026 under Apache 2.0, alongside the larger Nex-N2-Pro, which is post-trained on Qwen3.5-397B-A17B.

SpecificationNex-N2-mini
Total parameters35.1B
Base modelQwen3.5-35B-A3B-Base, post-trained
Context windowNot stated on the Nex-AGI model card
ModalitiesNot stated by Nex-AGI; the GGUF repo ships mmproj projector files next to the weights
Larger siblingNex-N2-Pro, post-trained on Qwen3.5-397B-A17B
Training focusAgentic Thinking: agentic coding, deep research, tool calling, terminal execution
ReasoningExplicit reasoning traces, parsed with the qwen3 reasoning parser
Function callingSupported, qwen3_coder tool-call parser
Recommended samplingtemperature 0.7, top_p 0.95, top_k 40
Vendor serving stackNex-AGI SGLang fork; the launch example for the mini is one server with two H100s
Release dateJune 4, 2026
LicenseApache 2.0

The idea behind the release is what Nex-AGI calls Agentic Thinking, a framework that connects requirement understanding, task planning, code implementation, environmental feedback, evaluation and debugging, and continuous iteration into a single closed loop. Adaptive Thinking lets the model decide when to think and how deeply, so it executes simple actions quickly and reasons thoroughly on critical decisions. Coherent Thinking keeps one reasoning paradigm across general reasoning and agentic tasks. In practice the model emits explicit reasoning traces and calls functions, and Nex-AGI's reference stack parses both with the qwen3 reasoning parser and the qwen3_coder tool-call parser in its customized SGLang fork.

Nex-N2-mini benchmarks

Nex-AGI published the launch numbers on the model card, scoring both N2 variants against GPT-5.5, Opus 4.7, Kimi-K2.6, GLM-5.1, MiniMax M3 and DeepSeek-V4-Pro. The rows below put the mini next to its Pro sibling and to GPT-5.5 and Opus 4.7.

BenchmarkNex-N2-miniNex-N2-ProGPT-5.5Opus 4.7
BrowseComp
Web browsing
74.183.784.479.8
Toolathlon
Long-horizon tools
33.351.955.652.8
SWE-Bench Verified
Software engineering
74.480.882.987.6
SWE-Bench Pro
Harder engineering
50.258.858.664.3
Terminal-Bench 2.1
Terminal agents
60.775.383.469.7
DeepSWE
Agentic coding
8.033.67054
GPQA Diamond
Expert science
82.690.793.694.2

The mini takes no row here, and the vendor table does not claim it does: these are frontier-scale comparisons and the Pro variant is the one that keeps pace. The widest gap in the rows above is DeepSWE, where the mini scores 8.0 against 70 for GPT-5.5 and 54 for Opus 4.7. On SWE-Bench Verified it scores 74.4 against 82.9 for GPT-5.5 and 87.6 for Opus 4.7, and on GPQA Diamond 82.6 against 93.6 and 94.2. Outside the rows above, the vendor table gives the mini 1402 on GDPval, 47.7 on WildClawBench, 31.5 on SWE Atlas QnA, 30.0 on SWE Atlas RF, 23.3 on SWE Atlas TW and 9.4 on Apex. On rows the vendor left blank for the frontier models, the mini still posts numbers: 62.0 on WideSearch, 65.9 on TAU3 and 89.1 on IFEval.

Nex-N2-mini hardware requirements

The system requirement to check is memory. Nex-AGI ships the original safetensors weights and recommends its own SGLang fork, with a launch example for the mini on one server with two H100s, so for a local machine the practical route is a GGUF build. The ladder below uses the real file sizes from the community repo bartowski/nex-agi_Nex-N2-mini-GGUF on Hugging Face, whose smallest weight file is IQ2_XXS at 9.78 GB.

MemoryBuild to pickFile size
12 GBIQ2_XXS9.78 GB
16 GBQ2_K_L13.11 GB
24 GBIQ4_XS18.81 GB
32 GBQ4_K_M21.39 GB
48 GBQ6_K30.05 GB
64 GB and upQ8_036.91 GB

When two builds both fit, take the larger one: most neighbouring builds in this repo are less than a gigabyte apart, and a few sit almost on top of each other, such as IQ3_XS at 16.22 GB and Q3_K_M at 16.23 GB. The low end runs IQ2_XXS at 9.78 GB, IQ2_XS at 10.80 GB and IQ2_S at 11.01 GB, and above that the repo lists IQ3_XXS at 14.87 GB, IQ4_NL at 19.86 GB and Q5_K_M at 25.02 GB. It also carries two mmproj projector files of 0.90 GB each, in bf16 and f16, separate from the weight builds. If the format is new to you, start with what GGUF is.

How to run Nex-N2-mini in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Nex-N2-mini in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the bigger variant, see Nex-N2-Pro, or browse every Nex model you can run locally.

Nex-N2-mini license

Nex-N2-mini is released under Apache 2.0, the license Nex-AGI set on the model repository. That permits commercial use, modification and redistribution with no royalties, so you can fine-tune the model, ship it inside your own products, and run it on your own hardware without a usage fee.

Get the weights from Hugging Face

huggingface-cli download nex-agi/Nex-N2-mini
from transformers import AutoModel
model = AutoModel.from_pretrained("nex-agi/Nex-N2-mini")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Nex-N2-mini is a 35.1B-parameter open-weight model from nex-agi, built on Qwen3.5-35B-A3B-Base. It is a Mixture-of-Experts model with about 3B active parameters per token, tuned as an agent for coding, tool calling, reasoning, and long-horizon tasks. It also accepts image input and supports a 256K-token context.

At a Q4 quantization the weights are roughly 21 GB, so a 24 GB GPU such as an RTX 3090, 4090, or 5090 handles it, as does a Mac with 32 GB or more of unified memory. Higher-precision Q8 weights need closer to 37 GB. With CPU offloading through llama.cpp it can run on smaller cards at reduced speed.

Yes. The model is released under the Apache-2.0 license, which allows free use including commercial deployment, modification, and redistribution. Running it locally in Atomic Chat has no usage fees, since the model executes on your own hardware rather than a paid API.

Yes. After you download the weights once, the model runs entirely on your machine with no internet connection. Prompts, code, and files stay on-device, which is the reason it is offered for local use in Atomic Chat.

Its strongest area is agentic coding and tool use. It scores 74.4 on SWE-Bench Verified and 60.7 on Terminal-Bench 2.1, so it can edit code across a repo and run terminal commands over multi-step tasks. It also reaches 82.6 on GPQA Diamond for hard reasoning and handles vision input and long documents.