What is Nex-N2-Pro?
Nex-N2-Pro is a 396.8B-parameter mixture-of-experts model from Nex-AGI and the larger of the two models in the Nex-N2 release, both of which Nex-AGI ships as open source. It is post-trained on Qwen3.5-397B-A17B, so about 17B parameters are active per token, and it targets agentic work: agentic coding, deep research, tool calling and terminal execution. Nex-AGI published the weights on Hugging Face on June 3, 2026 under Apache 2.0, alongside the smaller Nex-N2-mini. The vendor's own launch example serves the Pro across two 8x H100 nodes, so running it locally is a workstation job rather than a laptop one.
| Specification | Nex-N2-Pro |
|---|---|
| Total parameters | 396.8B |
| Active parameters | About 17B per token |
| Architecture | Mixture of experts, 512 experts, 10 routed per token, plus a shared expert |
| Base model | Qwen3.5-397B-A17B (post-trained) |
| Layers | 60 |
| Attention | Hybrid: full attention every fourth layer, gated-delta linear attention in the rest |
| Context window | 262,144 tokens |
| Modalities | Text and image input, text output |
| Multi-token prediction | One MTP layer in the checkpoint |
| Reasoning | Explicit reasoning traces, parsed with the qwen3 reasoning parser |
| Function calling | Supported, qwen3_coder tool-call parser |
| Recommended sampling | temperature 0.7, top_p 0.95, top_k 40 |
| Release date | June 3, 2026 |
| License | Apache 2.0 |
The idea Nex-AGI builds the release around is what it calls Agentic Thinking: one loop that ties requirement understanding, task planning, code implementation, environmental feedback, evaluation and debugging together. Adaptive Thinking lets the model decide when to think and how deeply, so it executes simple actions quickly and reasons thoroughly on critical decisions. Coherent Thinking carries one reasoning paradigm across general reasoning and agentic tasks, which is what Nex-AGI credits for stable capability transfer. The config matches that agent framing: one layer in four runs full attention and the rest run gated-delta linear attention, with 64 linear value heads and 16 linear key heads, over a 262,144-token position limit. A vision tower of 27 layers lets the model read images alongside text, and the checkpoint carries one multi-token prediction layer.
Nex-N2-Pro benchmarks
Nex-AGI published the launch table on its own model card, scoring both N2 variants against GPT-5.5, Opus 4.7, Kimi-K2.6, GLM-5.1, MiniMax M3 and DeepSeek-V4-Pro. These are the headline rows with the most complete columns:
| Benchmark | Nex-N2-Pro | GPT-5.5 | Opus 4.7 | Kimi-K2.6 | MiniMax M3 | DeepSeek-V4-Pro |
|---|---|---|---|---|---|---|
BrowseComp Web browsing | 83.7 | 84.4 | 79.8 | 83.2 | 83.5 | 83.4 |
GDPval Professional tasks | 1585 | 1769 | 1753 | 1481 | - | 1554 |
Toolathlon Long-horizon tools | 51.9 | 55.6 | 52.8 | 50.0 | - | 51.8 |
Terminal-Bench 2.1 Terminal agents | 75.3 | 83.4 | 69.7 | - | 66.0 | 72.0 |
SWE-Bench Verified Software engineering | 80.8 | 82.9 | 87.6 | 80.2 | 80.5 | 80.6 |
SWE-Bench Pro Harder engineering | 58.8 | 58.6 | 64.3 | 58.6 | 59.0 | 55.4 |
GPQA Diamond Expert science | 90.7 | 93.6 | 94.2 | 90.5 | - | 90.1 |
Nex-N2-Pro tops none of these rows: GPT-5.5 takes the four agent and terminal rows, Opus 4.7 takes the two SWE-Bench rows and GPQA Diamond. It is ahead of Kimi-K2.6 and DeepSeek-V4-Pro in every row here that those two columns report, and ahead of MiniMax M3 in every row except SWE-Bench Pro, where MiniMax M3 posts 59.0 against 58.8. Nex-AGI's own summary, that the Pro keeps pace with top-tier models such as GPT-5.5 and Opus 4.7, matches the numbers.
Nex-N2-Pro hardware requirements
The system requirement to check is memory, and at this size there is a lot of it to check. Nex-AGI ships only the original safetensors weights, roughly 794 GB in BF16; the ladder below uses the real file sizes from the community repo bartowski/nex-agi_Nex-N2-Pro-GGUF, where every build is split across several shards.
| Memory | Build to pick | File size |
|---|---|---|
| 96 GB | IQ1_S | 81.8 GB |
| 128 GB | IQ2_XXS | 106.3 GB |
| 192 GB | IQ3_XXS | 166.1 GB |
| 256 GB | IQ4_XS | 212.2 GB |
| 384 GB | Q5_K_M | 283.3 GB |
| 512 GB and up | Q8_0 | 421.5 GB |
When two builds both fit, take the larger one: the steps on this ladder are tens of gigabytes apart. Add the 0.9 GB mmproj file from the same repo if you want the model to read images. If the format is new to you, start with what GGUF is.
How to run Nex-N2-Pro in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Nex-N2-Pro in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
The smaller sibling, Nex-N2-mini, is post-trained on Qwen3.5-35B-A3B-Base, and Nex-AGI describes the two variants as covering different latency and quality trade-offs; both sit on the Nex models you can run locally page.
Nex-N2-Pro license
Nex-N2-Pro is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, so you can fine-tune the model, ship it inside your own products, and serve it on your own hardware without a usage fee or a per-token bill.
