Nex-N2-Pro

Updated
24.08.2026
Tools
Thinking
Vision
Reasoning
Code

Nex-N2-Pro is a 396.8B mixture-of-experts agent model from Nex-AGI, post-trained on Qwen3.5-397B-A17B. Apache 2.0, GGUF builds from 82 GB.

At a glance

  • License: Apache 2.0
  • Parameters: 396.8B total, about 17B active
  • Context length: 262,144 tokens
  • Modalities: Text and image in, text out
  • Minimum hardware: 96 GB memory with the IQ1_S GGUF (81.8 GB)

What is Nex-N2-Pro?

Nex-N2-Pro is a 396.8B-parameter mixture-of-experts model from Nex-AGI and the larger of the two models in the Nex-N2 release, both of which Nex-AGI ships as open source. It is post-trained on Qwen3.5-397B-A17B, so about 17B parameters are active per token, and it targets agentic work: agentic coding, deep research, tool calling and terminal execution. Nex-AGI published the weights on Hugging Face on June 3, 2026 under Apache 2.0, alongside the smaller Nex-N2-mini. The vendor's own launch example serves the Pro across two 8x H100 nodes, so running it locally is a workstation job rather than a laptop one.

SpecificationNex-N2-Pro
Total parameters396.8B
Active parametersAbout 17B per token
ArchitectureMixture of experts, 512 experts, 10 routed per token, plus a shared expert
Base modelQwen3.5-397B-A17B (post-trained)
Layers60
AttentionHybrid: full attention every fourth layer, gated-delta linear attention in the rest
Context window262,144 tokens
ModalitiesText and image input, text output
Multi-token predictionOne MTP layer in the checkpoint
ReasoningExplicit reasoning traces, parsed with the qwen3 reasoning parser
Function callingSupported, qwen3_coder tool-call parser
Recommended samplingtemperature 0.7, top_p 0.95, top_k 40
Release dateJune 3, 2026
LicenseApache 2.0

The idea Nex-AGI builds the release around is what it calls Agentic Thinking: one loop that ties requirement understanding, task planning, code implementation, environmental feedback, evaluation and debugging together. Adaptive Thinking lets the model decide when to think and how deeply, so it executes simple actions quickly and reasons thoroughly on critical decisions. Coherent Thinking carries one reasoning paradigm across general reasoning and agentic tasks, which is what Nex-AGI credits for stable capability transfer. The config matches that agent framing: one layer in four runs full attention and the rest run gated-delta linear attention, with 64 linear value heads and 16 linear key heads, over a 262,144-token position limit. A vision tower of 27 layers lets the model read images alongside text, and the checkpoint carries one multi-token prediction layer.

Nex-N2-Pro benchmarks

Nex-AGI published the launch table on its own model card, scoring both N2 variants against GPT-5.5, Opus 4.7, Kimi-K2.6, GLM-5.1, MiniMax M3 and DeepSeek-V4-Pro. These are the headline rows with the most complete columns:

BenchmarkNex-N2-ProGPT-5.5Opus 4.7Kimi-K2.6MiniMax M3DeepSeek-V4-Pro
BrowseComp
Web browsing
83.784.479.883.283.583.4
GDPval
Professional tasks
1585176917531481-1554
Toolathlon
Long-horizon tools
51.955.652.850.0-51.8
Terminal-Bench 2.1
Terminal agents
75.383.469.7-66.072.0
SWE-Bench Verified
Software engineering
80.882.987.680.280.580.6
SWE-Bench Pro
Harder engineering
58.858.664.358.659.055.4
GPQA Diamond
Expert science
90.793.694.290.5-90.1

Nex-N2-Pro tops none of these rows: GPT-5.5 takes the four agent and terminal rows, Opus 4.7 takes the two SWE-Bench rows and GPQA Diamond. It is ahead of Kimi-K2.6 and DeepSeek-V4-Pro in every row here that those two columns report, and ahead of MiniMax M3 in every row except SWE-Bench Pro, where MiniMax M3 posts 59.0 against 58.8. Nex-AGI's own summary, that the Pro keeps pace with top-tier models such as GPT-5.5 and Opus 4.7, matches the numbers.

Nex-N2-Pro hardware requirements

The system requirement to check is memory, and at this size there is a lot of it to check. Nex-AGI ships only the original safetensors weights, roughly 794 GB in BF16; the ladder below uses the real file sizes from the community repo bartowski/nex-agi_Nex-N2-Pro-GGUF, where every build is split across several shards.

MemoryBuild to pickFile size
96 GBIQ1_S81.8 GB
128 GBIQ2_XXS106.3 GB
192 GBIQ3_XXS166.1 GB
256 GBIQ4_XS212.2 GB
384 GBQ5_K_M283.3 GB
512 GB and upQ8_0421.5 GB

When two builds both fit, take the larger one: the steps on this ladder are tens of gigabytes apart. Add the 0.9 GB mmproj file from the same repo if you want the model to read images. If the format is new to you, start with what GGUF is.

How to run Nex-N2-Pro in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Nex-N2-Pro in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

The smaller sibling, Nex-N2-mini, is post-trained on Qwen3.5-35B-A3B-Base, and Nex-AGI describes the two variants as covering different latency and quality trade-offs; both sit on the Nex models you can run locally page.

Nex-N2-Pro license

Nex-N2-Pro is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can fine-tune the model, ship it inside your own products, and serve it on your own hardware without a usage fee or a per-token bill.

Get the weights from Hugging Face

huggingface-cli download nex-agi/Nex-N2-Pro
from transformers import AutoModel
model = AutoModel.from_pretrained("nex-agi/Nex-N2-Pro")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Nex-N2-Pro is an open-weight mixture-of-experts language model from nex-agi, released in June 2026 and built on the Qwen3.5 architecture. It has 396.8B total parameters with about 17B active per token, accepts text and images, and is tuned for coding, tool use, and long agentic workflows. It runs under the Apache 2.0 license, so you can download and use it for free.

A 4-bit quantized build needs roughly 214-256 GB of combined memory. That fits a Mac Studio with 256 GB of unified memory, or a PC with a 24 GB GPU and large system RAM using llama.cpp MoE offloading, where reports show 25+ tokens per second. Because only 17B parameters activate per token, you do not need a multi-GPU server to get usable speeds.

Yes. The weights are released under the Apache 2.0 license, which permits free personal and commercial use, modification, and redistribution. Running it locally in Atomic Chat costs nothing beyond your own hardware and electricity, with no API fees or per-token charges.

Yes. Once the weights are downloaded, Nex-N2-Pro runs entirely on your machine with no network connection required. In Atomic Chat every prompt and response stays on-device, so your code and documents are never sent to a server.

It is built for agentic engineering: writing and debugging code, calling tools, and running multi-step workflows on its own. It reported 80.8 on SWE-Bench Verified, a benchmark of fixing real bugs in real repositories, which puts it in useful day-to-day coding territory. Its 262,144-token context and vision support also let it reason over large codebases, long documents, and images.