What is K2 Horizon MoVA-36B-A4B?
K2 Horizon MoVA 36B-A4B combines a Mixture-of-Experts model with IFM's Mixture-of-Values attention. IFM describes 36B total parameters and roughly 4B active per token, with a 512K context window. The final checkpoint is released for text reasoning and agent workflows.
| Specification | K2 Horizon MoVA-36B-A4B |
|---|---|
| Parameters | 36B total class; 4B active per token |
| Architecture | Mixture-of-Experts with Mixture-of-Values attention |
| Layers | 48 |
| Context window | 524,288 tokens; native from midtraining |
| Modalities | Text input and output |
| Release status | Final checkpoint released |
| License | Apache-2.0, per model card |
The A4B suffix describes active computation, not a 4B download. The full inventory holds about 37.445B parameters. Its configuration has 48 layers, 100 routed experts and eight selected per token, plus a shared expert. MoVA also selects value experts inside attention; it is an additional architectural feature, not another name for MoE.
K2 Horizon MoVA-36B-A4B benchmarks
IFM publishes these percentage scores with Artificial Analysis reference scores. Muse Glimmer-30B uses high reasoning effort and the other open-model baselines use reasoning mode, according to the card. They are not benchmarks of a quantized Atomic Chat deployment. Source: official model card.
| Benchmark | K2-Horizon-MoVA-36B-A4B | Nemotron 3 Ultra | Nemotron 3 Super | G9v3-39A5B | Qwen3.6-35B-A3B | Muse Glimmer-30B | Gemma 4 31B-it |
|---|---|---|---|---|---|---|---|
tau3-Banking Banking agents | 26.8 | 14.2 | 10.3 | 22.1 | 9.3 | 23.5 | 14.8 |
Terminal-Bench 2.1 Terminal agents | 58.6 | 53.9 | 38.6 | 32.6 | 44.9 | 51.7 | 43.4 |
SciCode Scientific coding | 38.9 | 39.9 | 36.0 | 34.0 | 35.8 | 43.6 | 43.4 |
Humanity's Last Exam (without tools) Expert questions | 25.2 | 28.4 | 20.8 | 17.5 | 22.2 | 22.0 | 23.6 |
GPQA Diamond Expert science | 80.8 | 86.7 | 80.0 | 80.5 | 84.1 | 83.5 | 85.7 |
AA-LCR Long-context reasoning | 66.3 | 71.0 | 60.3 | 62.0 | 66.7 | 80.0 | 68.3 |
MoVA leads the displayed references on tau3-Banking and Terminal-Bench 2.1. It trails several of them on GPQA Diamond and long-context reasoning. The 4B active figure is useful for understanding the architecture, but it does not establish a speed advantage on your hardware without a throughput measurement.
K2 Horizon MoVA-36B-A4B hardware requirements
The original weight shards total 74.89 GB, while the official BF16 GGUF totals 74.92 GB. That GGUF is named K2-Horizon-36B-BF16.gguf; the repository identifies it as the MoVA 36B-A4B conversion. Both contain the full weight set, including experts not selected for a particular token.
| Precision | Source | Size on disk |
|---|---|---|
| BF16 Safetensors | IFM/K2-Horizon-MoVA-36B-A4B | 74.89 GB |
| BF16 GGUF (official) | IFM/K2-Horizon-MoVA-36B-A4B-GGUF K2-Horizon-36B-BF16.gguf | 74.92 GB |
The full BF16 weights exceed 64 GB. IFM validates a two-H200 SGLang recipe with a model-specific router override, so use the matching recipe when checking numerical behavior. That configuration does not prove a minimum GPU count, and the 512K cache is additional to the disk totals.
Sizes use decimal GB and cover weights only. For the format distinction, see what GGUF is. A memory tier is not listed because minimum runtime memory has not been measured for these builds.
How to run K2 Horizon MoVA-36B-A4B in Atomic Chat
Atomic Chat execution has not been verified for this checkpoint. IFM's official GGUF card requires a llama.cpp build with K2 Horizon architecture support and points to its development fork. Check architecture support before downloading these files. The usual search, download and start-chat flow is not yet a verified procedure for this model. Browse the model catalog for other checkpoints and their documented download options.
K2 Horizon MoVA-36B-A4B license
IFM labels this checkpoint Apache-2.0 in the official model card. Use the publisher's release information as the source for its license, and review any additional notices attached to the exact build you redistribute. A community conversion is not an Atomic Chat release.
Sources and file inventories checked September 15, 2026.
