What is K2 Horizon 375B-A23B?
K2 Horizon 375B-A23B is IFM's flagship text model for multi-step agent work. The final release uses a Mixture-of-Experts architecture with 23B active parameters per token and a 512K context window. Self-hosting gives you control of the weights and serving stack, but requires server-scale storage and memory.
| Specification | K2 Horizon 375B-A23B |
|---|---|
| Parameters | 375B total class; 23B active per token |
| Architecture | Mixture-of-Experts |
| Layers | 61 |
| Context window | 524,288 tokens; native from midtraining |
| Modalities | Text input and output |
| Release status | Final checkpoint released |
| License | Apache-2.0, per model card |
The model selects eight of 192 routed experts per token and includes a shared expert. Its 61-layer configuration uses hidden size 6,144. The uploaded inventory contains 379.167B parameters; A23B is the active subset, not the amount of model data you need to store or make available to the runtime.
K2 Horizon 375B-A23B benchmarks
These are percentage scores from IFM's release table, not our measurements. SWE Bench Pro uses the card's strict, no-internet setting; preserve that label when comparing it with another report. A dash marks an unavailable result. Other evaluation setups must be checked before making cross-table claims. Source: official model card.
| Benchmark | K2-Horizon-375B-A23B | Nemotron 3 Ultra | Inkling (xhigh) | MiniMax-M3 | GLM 5.2 (max) | GPT 5.6 Luna (max) | GPT 5.6 Terra (high) | Claude Sonnet5 (max) |
|---|---|---|---|---|---|---|---|---|
Toolathlon Verified Long-horizon tools | 65.3 | 34.3 | 45.5 | 53.7 | 59.9 | 67.5 | 64.8 | 71.6 |
MCPMark MCP tools | 67.7 | 45.7 | 51.2 | 48.8 | 72.4 | 66.9 | 74.0 | 65.3 |
Terminal-Bench 2.1 Terminal agents | 70.2 | 53.9 | 55.1 | 65.2 | 77.9 | 80.9 | 75.7 | 80.5 |
SWE Bench Pro (strict) Harder engineering | 42.6 | 38.7 | 43.1 | 43.8 | 46.7 | 48.8 | -- | -- |
GPQA Diamond Expert science | 87.3 | 86.7 | 87.2 | 92.9 | 89.5 | 91.1 | 89.6 | 91.1 |
AA-LCR Long-context reasoning | 76.0 | 71.0 | 73.3 | 80.3 | 76.7 | 78.3 | 73.3 | 77.0 |
K2 Horizon reaches 67.7 on MCPMark, above Claude Sonnet5 in this table but below GPT 5.6 Terra. It scores 70.2 on Terminal-Bench 2.1, below the listed GLM and GPT references. The results give you specific agent tasks to evaluate rather than a blanket claim that the model replaces a hosted frontier service.
K2 Horizon 375B-A23B hardware requirements
The original BF16 shards total 758.34 GB. A third-party Baekpica conversion provides complete 30-shard BF16 and Q8_0 GGUF sets, with a pinned source revision and converter commit. The table totals every shard; downloading the first file alone does not download the model.
| Precision | Source | Size on disk |
|---|---|---|
| BF16 Safetensors | IFM/K2-Horizon-375B-A23B | 758.34 GB |
| BF16 GGUF (community) | Baekpica/K2-Horizon-375B-A23B-GGUF All 30 shards | 758.48 GB |
| Q8_0 GGUF (community) | Baekpica/K2-Horizon-375B-A23B-GGUF All 30 shards | 403.08 GB |
Baekpica reports structural and file-integrity checks, not a minimum-memory or Atomic Chat inference test. Even Q8_0 occupies 403.08 GB on disk before runtime allocations. IFM's official SGLang recipe uses eight H200 GPUs. That is a validated example, not proof that every compatible deployment needs exactly eight GPUs.
Sizes use decimal GB and cover weights only. For the format distinction, see what GGUF is. A memory tier is not listed because minimum runtime memory has not been measured for these builds.
How to run K2 Horizon 375B-A23B in Atomic Chat
Atomic Chat execution has not been verified for this checkpoint. The community GGUF conversion uses IFM's K2 Horizon llama.cpp branch; a conversion audit does not confirm compatibility with the app's bundled runtime. Check architecture support before downloading these files. The usual search, download and start-chat flow is not yet a verified procedure for this model. Browse the model catalog for other checkpoints and their documented download options.
K2 Horizon 375B-A23B license
IFM labels this checkpoint Apache-2.0 in the official model card. Use the publisher's release information as the source for its license, and review any additional notices attached to the exact build you redistribute. A community conversion is not an Atomic Chat release.
Sources and file inventories checked September 15, 2026.
