K2 Horizon 375B-A23B

Updated
15.09.2026
Thinking
Tools
Reasoning
Code

K2 Horizon 375B-A23B from IFM: 512K context, sparse weights and server setup. Atomic Chat support is not yet verified.

At a glance

  • License: Apache-2.0, per model card
  • Parameters: 375B total class; 23B active per token
  • Context length: 524,288 tokens; native from midtraining
  • Modalities: Text input and output
  • Minimum hardware: Not measured; 758.34 GB BF16 files alone, plus runtime memory

What is K2 Horizon 375B-A23B?

K2 Horizon 375B-A23B is IFM's flagship text model for multi-step agent work. The final release uses a Mixture-of-Experts architecture with 23B active parameters per token and a 512K context window. Self-hosting gives you control of the weights and serving stack, but requires server-scale storage and memory.

SpecificationK2 Horizon 375B-A23B
Parameters375B total class; 23B active per token
ArchitectureMixture-of-Experts
Layers61
Context window524,288 tokens; native from midtraining
ModalitiesText input and output
Release statusFinal checkpoint released
LicenseApache-2.0, per model card

The model selects eight of 192 routed experts per token and includes a shared expert. Its 61-layer configuration uses hidden size 6,144. The uploaded inventory contains 379.167B parameters; A23B is the active subset, not the amount of model data you need to store or make available to the runtime.

K2 Horizon 375B-A23B benchmarks

These are percentage scores from IFM's release table, not our measurements. SWE Bench Pro uses the card's strict, no-internet setting; preserve that label when comparing it with another report. A dash marks an unavailable result. Other evaluation setups must be checked before making cross-table claims. Source: official model card.

BenchmarkK2-Horizon-375B-A23BNemotron 3 UltraInkling (xhigh)MiniMax-M3GLM 5.2 (max)GPT 5.6 Luna (max)GPT 5.6 Terra (high)Claude Sonnet5 (max)
Toolathlon Verified
Long-horizon tools
65.334.345.553.759.967.564.871.6
MCPMark
MCP tools
67.745.751.248.872.466.974.065.3
Terminal-Bench 2.1
Terminal agents
70.253.955.165.277.980.975.780.5
SWE Bench Pro (strict)
Harder engineering
42.638.743.143.846.748.8----
GPQA Diamond
Expert science
87.386.787.292.989.591.189.691.1
AA-LCR
Long-context reasoning
76.071.073.380.376.778.373.377.0

K2 Horizon reaches 67.7 on MCPMark, above Claude Sonnet5 in this table but below GPT 5.6 Terra. It scores 70.2 on Terminal-Bench 2.1, below the listed GLM and GPT references. The results give you specific agent tasks to evaluate rather than a blanket claim that the model replaces a hosted frontier service.

K2 Horizon 375B-A23B hardware requirements

The original BF16 shards total 758.34 GB. A third-party Baekpica conversion provides complete 30-shard BF16 and Q8_0 GGUF sets, with a pinned source revision and converter commit. The table totals every shard; downloading the first file alone does not download the model.

PrecisionSourceSize on disk
BF16 SafetensorsIFM/K2-Horizon-375B-A23B758.34 GB
BF16 GGUF (community)Baekpica/K2-Horizon-375B-A23B-GGUF
All 30 shards
758.48 GB
Q8_0 GGUF (community)Baekpica/K2-Horizon-375B-A23B-GGUF
All 30 shards
403.08 GB

Baekpica reports structural and file-integrity checks, not a minimum-memory or Atomic Chat inference test. Even Q8_0 occupies 403.08 GB on disk before runtime allocations. IFM's official SGLang recipe uses eight H200 GPUs. That is a validated example, not proof that every compatible deployment needs exactly eight GPUs.

Sizes use decimal GB and cover weights only. For the format distinction, see what GGUF is. A memory tier is not listed because minimum runtime memory has not been measured for these builds.

How to run K2 Horizon 375B-A23B in Atomic Chat

Atomic Chat execution has not been verified for this checkpoint. The community GGUF conversion uses IFM's K2 Horizon llama.cpp branch; a conversion audit does not confirm compatibility with the app's bundled runtime. Check architecture support before downloading these files. The usual search, download and start-chat flow is not yet a verified procedure for this model. Browse the model catalog for other checkpoints and their documented download options.

K2 Horizon 375B-A23B license

IFM labels this checkpoint Apache-2.0 in the official model card. Use the publisher's release information as the source for its license, and review any additional notices attached to the exact build you redistribute. A community conversion is not an Atomic Chat release.

Sources and file inventories checked September 15, 2026.

Get the weights from Hugging Face

# Download only; check disk space first.
hf download IFM/K2-Horizon-375B-A23B --revision 1715eebf1df33acb542f9ed61b780a1d4541f5f7

For a compatible local server on port 8000:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"IFM/K2-Horizon-375B-A23B","messages":[{"role":"user","content":"Write tests for a function that parses a CSV row."}],"temperature":1,"top_p":0.95,"max_tokens":32768,"chat_template_kwargs":{"reasoning_effort":"high"}}'

Install the OpenAI Python client and start a compatible local server first.

from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="local")
result = client.chat.completions.create(
    model="IFM/K2-Horizon-375B-A23B",
    messages=[{"role": "user", "content": "Write tests for a CSV parser."}],
    temperature=1.0, top_p=0.95, max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(result.choices[0].message.content)

Requires a compatible local server and the OpenAI JavaScript client.

import OpenAI from "openai";
const client = new OpenAI({baseURL: "http://127.0.0.1:8000/v1", apiKey: "local"});
const result = await client.chat.completions.create({
  "model": "IFM/K2-Horizon-375B-A23B",
  "messages": [
    {
      "role": "user",
      "content": "Write tests for a function that parses a CSV row."
    }
  ],
  "temperature": 1,
  "top_p": 0.95,
  "max_tokens": 32768,
  "chat_template_kwargs": {
    "reasoning_effort": "high"
  }
});
console.log(result.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

K2 Horizon 375B-A23B is IFM's flagship text model for multi-step agent work. The final release uses a Mixture-of-Experts architecture with 23B active parameters per token and a 512K context window. Self-hosting gives you control of the weights and serving stack, but requires server-scale storage and memory.

Execution of this checkpoint has not been verified in Atomic Chat. The available community GGUF uses IFM's K2 Horizon development branch. A downloadable GGUF does not by itself confirm that the app can load it.

Its original BF16 weight files total 758.34 GB. That is disk usage, not a tested minimum RAM or VRAM figure. You also need memory for the inference engine and a cache that grows with context.

The Baekpica BF16 and Q8_0 builds each split the complete model into 30 files. Keep every numbered sibling for your chosen precision in the same directory and use its first shard as the entry point. The 23B active-parameter figure does not reduce the checkpoint to a 23B download.

IFM labels this checkpoint Apache-2.0 in the official model card. Use the publisher's release information as the source for its license, and review any additional notices attached to the exact build you redistribute. A community conversion is not an Atomic Chat release.