K2 Horizon 32B

Updated
15.09.2026
Thinking
Tools
Reasoning
Code

K2 Horizon 32B from IFM: 512K context, Stage 1 dense weights and server setup. Atomic Chat support is not yet verified.

At a glance

  • License: Apache-2.0, per model card
  • Parameters: 32B class; 34.779B stored parameters
  • Context length: 524,288 tokens; native from midtraining
  • Modalities: Text input and output
  • Minimum hardware: Not measured; 69.56 GB BF16 files alone, plus runtime memory

What is K2 Horizon 32B?

K2 Horizon 32B is IFM's large dense text model with 512K context. The available release is Stage 1, and the model card says the final checkpoint is still to come. Evaluate the weights available now, and keep their revision with your results so a later training stage does not silently replace your baseline.

SpecificationK2 Horizon 32B
Parameters32B class; 34.779B stored parameters
ArchitectureDense decoder-only
Layers64
Context window524,288 tokens; native from midtraining
ModalitiesText input and output
Release statusStage 1; final checkpoint pending
LicenseApache-2.0, per model card

Unlike the MoVA sibling, this model is dense. It uses 64 layers with hidden size 5,120, and the uploaded inventory contains about 34.779B parameters. A similar full-weight size does not imply the same computation per token as an MoE model. Release stage matters too: these scores belong to Stage 1 only.

K2 Horizon 32B benchmarks

IFM reports these Stage 1 percentage scores and takes the reference scores from Artificial Analysis. Its card specifies high reasoning effort for Muse Glimmer-30B and reasoning mode for the other open models. We have not reproduced this comparison. Source: official model card.

BenchmarkK2-Horizon-32B-Stage1Qwen3.8-27BMuse Glimmer-30BIBM Granite 4.2 30B
tau3-Banking
Banking agents
22.548.023.514.4
Terminal-Bench 2.1
Terminal agents
36.679.851.726.6
SciCode
Scientific coding
30.244.743.636.6
Humanity's Last Exam (without tools)
Expert questions
22.833.922.011.2
GPQA Diamond
Expert science
82.390.583.564.4
AA-LCR
Long-context reasoning
65.377.380.046.7

Stage 1 scores 82.3 on GPQA Diamond, but Qwen3.8-27B leads all six displayed rows. Its Terminal-Bench 2.1 score is also below Muse Glimmer-30B. The larger parameter label alone is not a reason to switch; compare with Qwen3.8-27B on the tasks you intend to run.

K2 Horizon 32B hardware requirements

The original BF16 shards total 69.56 GB, and IFM's BF16 GGUF totals 69.57 GB. The GGUF card explicitly identifies the conversion as Stage 1. Do not label it a final checkpoint or assume conversion to GGUF reduces it to a 4-bit download.

PrecisionSourceSize on disk
BF16 SafetensorsIFM/K2-Horizon-32B69.56 GB
BF16 GGUF (official)IFM/K2-Horizon-32B-GGUF
K2-Horizon-32B-BF16.gguf
69.57 GB

The full BF16 weights exceed a 64 GB memory pool before cache and runtime overhead. IFM's SGLang recipe was validated on two H200 GPUs; this is a serving example, not a measured minimum. Context length and concurrency require separate capacity tests.

Sizes use decimal GB and cover weights only. For the format distinction, see what GGUF is. A memory tier is not listed because minimum runtime memory has not been measured for these builds.

How to run K2 Horizon 32B in Atomic Chat

Atomic Chat execution has not been verified for this checkpoint. IFM's official GGUF card requires a llama.cpp build with K2 Horizon architecture support and points to its development fork. Check architecture support before downloading these files. The usual search, download and start-chat flow is not yet a verified procedure for this model. Browse the model catalog for other checkpoints and their documented download options.

K2 Horizon 32B license

IFM labels this checkpoint Apache-2.0 in the official model card. Use the publisher's release information as the source for its license, and review any additional notices attached to the exact build you redistribute. A community conversion is not an Atomic Chat release.

Sources and file inventories checked September 15, 2026.

Get the weights from Hugging Face

# Download only; check disk space first.
hf download IFM/K2-Horizon-32B --revision e0fe3043018e32da0b3af7241584f863ab0372eb

For a compatible local server on port 8000:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"IFM/K2-Horizon-32B","messages":[{"role":"user","content":"Write tests for a function that parses a CSV row."}],"temperature":1,"top_p":0.95,"max_tokens":32768,"chat_template_kwargs":{"reasoning_effort":"high"}}'

Install the OpenAI Python client and start a compatible local server first.

from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="local")
result = client.chat.completions.create(
    model="IFM/K2-Horizon-32B",
    messages=[{"role": "user", "content": "Write tests for a CSV parser."}],
    temperature=1.0, top_p=0.95, max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(result.choices[0].message.content)

Requires a compatible local server and the OpenAI JavaScript client.

import OpenAI from "openai";
const client = new OpenAI({baseURL: "http://127.0.0.1:8000/v1", apiKey: "local"});
const result = await client.chat.completions.create({
  "model": "IFM/K2-Horizon-32B",
  "messages": [
    {
      "role": "user",
      "content": "Write tests for a function that parses a CSV row."
    }
  ],
  "temperature": 1,
  "top_p": 0.95,
  "max_tokens": 32768,
  "chat_template_kwargs": {
    "reasoning_effort": "high"
  }
});
console.log(result.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

K2 Horizon 32B is IFM's large dense text model with 512K context. The available release is Stage 1, and the model card says the final checkpoint is still to come. Evaluate the weights available now, and keep their revision with your results so a later training stage does not silently replace your baseline.

Execution of this checkpoint has not been verified in Atomic Chat. The official GGUF requires a llama.cpp build with K2 Horizon support. A downloadable GGUF does not by itself confirm that the app can load it.

Its original BF16 weight files total 69.56 GB. That is disk usage, not a tested minimum RAM or VRAM figure. You also need memory for the inference engine and a cache that grows with context.

No. The checked model card and official GGUF card identify the available checkpoint as Stage 1. The final checkpoint is still pending, so the numbers on this page must not be presented as its future results.

IFM labels this checkpoint Apache-2.0 in the official model card. Use the publisher's release information as the source for its license, and review any additional notices attached to the exact build you redistribute. A community conversion is not an Atomic Chat release.