K2 Horizon 0.9B

Updated
15.09.2026
Thinking
Tools
Reasoning
Code

K2 Horizon 0.9B from IFM: 128K context, dense weights and server setup. Atomic Chat support is not yet verified.

At a glance

  • License: Apache-2.0 label; conflicting metadata
  • Parameters: 0.9B class; 1.078B stored parameters
  • Context length: 131,072 tokens; YaRN from 8,192
  • Modalities: Text input and output
  • Minimum hardware: Not measured; 2.16 GB BF16 files alone, plus runtime memory

What is K2 Horizon 0.9B?

K2 Horizon 0.9B is IFM's smallest dense reasoning model. It accepts text and uses multi-teacher distillation to combine math, coding and instruction-following skills in a compact checkpoint. You can evaluate it for short code generation or tightly scoped tool selection before moving to a larger model.

SpecificationK2 Horizon 0.9B
Parameters0.9B class; 1.078B stored parameters
ArchitectureDense decoder-only
Layers28
Context window131,072 tokens; YaRN from 8,192
ModalitiesText input and output
Release statusReleased checkpoint
LicenseApache-2.0 label; conflicting metadata

The 128K window comes from YaRN scaling beyond the original 8,192-token context. It is not native 128K pretraining. The 0.9B name is also a size class: Hugging Face reports 1.078B stored parameters. Use the actual file inventory when choosing a download, especially on a memory-constrained machine.

K2 Horizon 0.9B benchmarks

IFM reports these percentage scores in the model card. Qwen3.5-2B is a larger reference model. They are vendor-reported results, not tests run in Atomic Chat; the card links a technical appendix for protocol details. Source: official model card.

BenchmarkK2-Horizon-0.9BQwen3.5-0.8BOpenBMB-1BQwen3.5-2B
AIME 2025
Competition math
41.71.040.434.2
AIME 2026
Competition math
48.50.240.438.8
HMMT Feb 2026
Olympiad math
25.80.623.322.7
GPQA Diamond
Expert science
27.311.926.354.9
HumanEval+
Code generation
79.916.565.275.6
BFCL v4
Function calling
28.025.325.243.6

The 0.9B model leads these references on the selected math tests and HumanEval+. Qwen3.5-2B is ahead on GPQA Diamond and BFCL v4. If your application depends on reliable function selection, the coding scores alone are not enough to choose the smaller model.

K2 Horizon 0.9B hardware requirements

The original Safetensors download is 2.16 GB. IFM also publishes a BF16 GGUF whose filename says 1B, while its repository identifies the source as K2 Horizon 0.9B. This is a format conversion, not a low-bit quantization.

PrecisionSourceSize on disk
BF16 SafetensorsIFM/K2-Horizon-0.9B2.16 GB
BF16 GGUF (official)IFM/K2-Horizon-0.9B-GGUF
K2-Horizon-1B-BF16.gguf
2.16 GB

A 2 GB memory pool cannot hold the complete BF16 weight set plus the inference engine. The GGUF file size does not include the operating system or a growing KV cache. Start capacity testing at a short context; the advertised 128K maximum is not a measured fit on a small device.

Sizes use decimal GB and cover weights only. For the format distinction, see what GGUF is. A memory tier is not listed because minimum runtime memory has not been measured for these builds.

How to run K2 Horizon 0.9B in Atomic Chat

Atomic Chat execution has not been verified for this checkpoint. IFM's official GGUF card requires a llama.cpp build with K2 Horizon architecture support and points to its development fork. Check architecture support before downloading these files. The usual search, download and start-chat flow is not yet a verified procedure for this model. Browse the model catalog for other checkpoints and their documented download options.

K2 Horizon 0.9B license

The model card sets license: apache-2.0, but its front matter also contains license_name: internal-only and points to a LICENSE file absent from the checked inventory. Those declarations conflict. Confirm the intended license with IFM before commercial use or redistribution; we do not resolve the conflict by copying a sibling model's terms.

Sources and file inventories checked September 15, 2026.

Get the weights from Hugging Face

# Download only; check disk space first.
hf download IFM/K2-Horizon-0.9B --revision 2e1ac41e5676266596f8f2198321d0760f041af6

For a compatible local server on port 8000:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"IFM/K2-Horizon-0.9B","messages":[{"role":"user","content":"Write tests for a function that parses a CSV row."}],"temperature":0.6,"top_p":0.95,"max_tokens":32768,"chat_template_kwargs":{"reasoning_effort":"high"}}'

Install the OpenAI Python client and start a compatible local server first.

from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="local")
result = client.chat.completions.create(
    model="IFM/K2-Horizon-0.9B",
    messages=[{"role": "user", "content": "Write tests for a CSV parser."}],
    temperature=0.6, top_p=0.95, max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(result.choices[0].message.content)

Requires a compatible local server and the OpenAI JavaScript client.

import OpenAI from "openai";
const client = new OpenAI({baseURL: "http://127.0.0.1:8000/v1", apiKey: "local"});
const result = await client.chat.completions.create({
  "model": "IFM/K2-Horizon-0.9B",
  "messages": [
    {
      "role": "user",
      "content": "Write tests for a function that parses a CSV row."
    }
  ],
  "temperature": 0.6,
  "top_p": 0.95,
  "max_tokens": 32768,
  "chat_template_kwargs": {
    "reasoning_effort": "high"
  }
});
console.log(result.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

K2 Horizon 0.9B is IFM's smallest dense reasoning model. It accepts text and uses multi-teacher distillation to combine math, coding and instruction-following skills in a compact checkpoint. You can evaluate it for short code generation or tightly scoped tool selection before moving to a larger model.

Execution of this checkpoint has not been verified in Atomic Chat. The official GGUF requires a llama.cpp build with K2 Horizon support. A downloadable GGUF does not by itself confirm that the app can load it.

Its original BF16 weight files total 2.16 GB. That is disk usage, not a tested minimum RAM or VRAM figure. You also need memory for the inference engine and a cache that grows with context.

The configuration extends an 8,192-token context to 131,072 with YaRN scaling. IFM's official conversion is named K2-Horizon-1B-BF16.gguf, but its model card links back to the 0.9B checkpoint. Use that source identity rather than treating the filename as a separate release.

The model card sets license: apache-2.0, but its front matter also contains license_name: internal-only and points to a LICENSE file absent from the checked inventory. Those declarations conflict. Confirm the intended license with IFM before commercial use or redistribution; we do not resolve the conflict by copying a sibling model's terms.