K2 Horizon MoVA-36B-A4B

Updated
15.09.2026
Thinking
Tools
Reasoning
Code

K2 Horizon MoVA-36B-A4B from IFM: 512K context, sparse weights and server setup. Atomic Chat support is not yet verified.

At a glance

  • License: Apache-2.0, per model card
  • Parameters: 36B total class; 4B active per token
  • Context length: 524,288 tokens; native from midtraining
  • Modalities: Text input and output
  • Minimum hardware: Not measured; 74.89 GB BF16 files alone, plus runtime memory

What is K2 Horizon MoVA-36B-A4B?

K2 Horizon MoVA 36B-A4B combines a Mixture-of-Experts model with IFM's Mixture-of-Values attention. IFM describes 36B total parameters and roughly 4B active per token, with a 512K context window. The final checkpoint is released for text reasoning and agent workflows.

SpecificationK2 Horizon MoVA-36B-A4B
Parameters36B total class; 4B active per token
ArchitectureMixture-of-Experts with Mixture-of-Values attention
Layers48
Context window524,288 tokens; native from midtraining
ModalitiesText input and output
Release statusFinal checkpoint released
LicenseApache-2.0, per model card

The A4B suffix describes active computation, not a 4B download. The full inventory holds about 37.445B parameters. Its configuration has 48 layers, 100 routed experts and eight selected per token, plus a shared expert. MoVA also selects value experts inside attention; it is an additional architectural feature, not another name for MoE.

K2 Horizon MoVA-36B-A4B benchmarks

IFM publishes these percentage scores with Artificial Analysis reference scores. Muse Glimmer-30B uses high reasoning effort and the other open-model baselines use reasoning mode, according to the card. They are not benchmarks of a quantized Atomic Chat deployment. Source: official model card.

BenchmarkK2-Horizon-MoVA-36B-A4BNemotron 3 UltraNemotron 3 SuperG9v3-39A5BQwen3.6-35B-A3BMuse Glimmer-30BGemma 4 31B-it
tau3-Banking
Banking agents
26.814.210.322.19.323.514.8
Terminal-Bench 2.1
Terminal agents
58.653.938.632.644.951.743.4
SciCode
Scientific coding
38.939.936.034.035.843.643.4
Humanity's Last Exam (without tools)
Expert questions
25.228.420.817.522.222.023.6
GPQA Diamond
Expert science
80.886.780.080.584.183.585.7
AA-LCR
Long-context reasoning
66.371.060.362.066.780.068.3

MoVA leads the displayed references on tau3-Banking and Terminal-Bench 2.1. It trails several of them on GPQA Diamond and long-context reasoning. The 4B active figure is useful for understanding the architecture, but it does not establish a speed advantage on your hardware without a throughput measurement.

K2 Horizon MoVA-36B-A4B hardware requirements

The original weight shards total 74.89 GB, while the official BF16 GGUF totals 74.92 GB. That GGUF is named K2-Horizon-36B-BF16.gguf; the repository identifies it as the MoVA 36B-A4B conversion. Both contain the full weight set, including experts not selected for a particular token.

PrecisionSourceSize on disk
BF16 SafetensorsIFM/K2-Horizon-MoVA-36B-A4B74.89 GB
BF16 GGUF (official)IFM/K2-Horizon-MoVA-36B-A4B-GGUF
K2-Horizon-36B-BF16.gguf
74.92 GB

The full BF16 weights exceed 64 GB. IFM validates a two-H200 SGLang recipe with a model-specific router override, so use the matching recipe when checking numerical behavior. That configuration does not prove a minimum GPU count, and the 512K cache is additional to the disk totals.

Sizes use decimal GB and cover weights only. For the format distinction, see what GGUF is. A memory tier is not listed because minimum runtime memory has not been measured for these builds.

How to run K2 Horizon MoVA-36B-A4B in Atomic Chat

Atomic Chat execution has not been verified for this checkpoint. IFM's official GGUF card requires a llama.cpp build with K2 Horizon architecture support and points to its development fork. Check architecture support before downloading these files. The usual search, download and start-chat flow is not yet a verified procedure for this model. Browse the model catalog for other checkpoints and their documented download options.

K2 Horizon MoVA-36B-A4B license

IFM labels this checkpoint Apache-2.0 in the official model card. Use the publisher's release information as the source for its license, and review any additional notices attached to the exact build you redistribute. A community conversion is not an Atomic Chat release.

Sources and file inventories checked September 15, 2026.

Get the weights from Hugging Face

# Download only; check disk space first.
hf download IFM/K2-Horizon-MoVA-36B-A4B --revision de2d2efb32ed7639b7140bccbefe131a0063a982

For a compatible local server on port 8000:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"IFM/K2-Horizon-MoVA-36B-A4B","messages":[{"role":"user","content":"Write tests for a function that parses a CSV row."}],"temperature":1,"top_p":0.95,"max_tokens":32768,"chat_template_kwargs":{"reasoning_effort":"high"}}'

Install the OpenAI Python client and start a compatible local server first.

from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="local")
result = client.chat.completions.create(
    model="IFM/K2-Horizon-MoVA-36B-A4B",
    messages=[{"role": "user", "content": "Write tests for a CSV parser."}],
    temperature=1.0, top_p=0.95, max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(result.choices[0].message.content)

Requires a compatible local server and the OpenAI JavaScript client.

import OpenAI from "openai";
const client = new OpenAI({baseURL: "http://127.0.0.1:8000/v1", apiKey: "local"});
const result = await client.chat.completions.create({
  "model": "IFM/K2-Horizon-MoVA-36B-A4B",
  "messages": [
    {
      "role": "user",
      "content": "Write tests for a function that parses a CSV row."
    }
  ],
  "temperature": 1,
  "top_p": 0.95,
  "max_tokens": 32768,
  "chat_template_kwargs": {
    "reasoning_effort": "high"
  }
});
console.log(result.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

K2 Horizon MoVA 36B-A4B combines a Mixture-of-Experts model with IFM's Mixture-of-Values attention. IFM describes 36B total parameters and roughly 4B active per token, with a 512K context window. The final checkpoint is released for text reasoning and agent workflows.

Execution of this checkpoint has not been verified in Atomic Chat. The official GGUF requires a llama.cpp build with K2 Horizon support. A downloadable GGUF does not by itself confirm that the app can load it.

Its original BF16 weight files total 74.89 GB. That is disk usage, not a tested minimum RAM or VRAM figure. You also need memory for the inference engine and a cache that grows with context.

No. It means about 4B parameters participate in the computation for each token. The full stored inventory is about 37.445B parameters and its BF16 Safetensors files occupy 74.89 GB. Runtime memory and the attention cache add to the weight footprint.

IFM labels this checkpoint Apache-2.0 in the official model card. Use the publisher's release information as the source for its license, and review any additional notices attached to the exact build you redistribute. A community conversion is not an Atomic Chat release.