K2 Horizon 3.7B

Updated
15.09.2026
Thinking
Tools
Reasoning
Code

K2 Horizon 3.7B from IFM: 512K context, dense weights and server setup. Atomic Chat support is not yet verified.

At a glance

  • License: Apache-2.0, per model card
  • Parameters: 3.7B core; 5.058B stored parameters
  • Context length: 524,288 tokens; native from midtraining
  • Modalities: Text input and output
  • Minimum hardware: Not measured; 10.12 GB BF16 files alone, plus runtime memory

What is K2 Horizon 3.7B?

K2 Horizon 3.7B is IFM's small dense model for text reasoning and agent tasks. It offers a 512K context window and publishes results for repository repair, terminal work and function calling. This size is worth testing when a compact model misses instructions but a larger download is difficult to accommodate.

SpecificationK2 Horizon 3.7B
Parameters3.7B core; 5.058B stored parameters
ArchitectureDense decoder-only
Layers36
Context window524,288 tokens; native from midtraining
ModalitiesText input and output
Release statusReleased checkpoint
LicenseApache-2.0, per model card

The name counts a 3.7B core, while the complete tensor inventory contains 5.058B parameters. The configuration has 36 layers and hidden size 2,560. Native 512K context starts in the midtraining stages; it is a supported window, not a promise that a small machine can keep that entire conversation in memory.

K2 Horizon 3.7B benchmarks

These percentage scores come from IFM's model card, which warns that baseline protocols may differ. Atomic Chat has not reproduced them. Keep the evaluation setup with each score when comparing a local run with the published table. Source: official model card.

BenchmarkK2-Horizon-3.7BQwen3.5-4BG9v3-3BGranite 4.2-3BNemotron 3 Nano-4B
HMMT Feb 2026
Olympiad math
70.561.634.157.234.7
SWE-bench Verified
Software engineering
68.641.216.432.21.8
GPQA Diamond
Expert science
65.477.143.855.951.3
HLE
Expert questions
12.99.94.56.64.9
SciCode
Scientific coding
25.916.117.724.916.4
Terminal-Bench 2.1
Terminal agents
25.125.86.013.93.7
BFCL v4
Function calling
50.955.747.950.836.8

IFM reports 68.6 on SWE-bench Verified, ahead of the listed small-model references. Qwen3.5-4B scores higher on GPQA Diamond, Terminal-Bench 2.1 and BFCL v4. The case for this checkpoint is therefore task-specific: test repository changes separately from tool selection and terminal execution.

K2 Horizon 3.7B hardware requirements

The original weight shards total 10.12 GB, not the roughly 7.4 GB you would infer from the 3.7B name alone. The official GGUF is called K2-Horizon-4B-BF16.gguf and retains BF16 precision. Its source is the 3.7B model, not a separate 4B release.

PrecisionSourceSize on disk
BF16 SafetensorsIFM/K2-Horizon-3.7B10.12 GB
BF16 GGUF (official)IFM/K2-Horizon-3.7B-GGUF
K2-Horizon-4B-BF16.gguf
10.13 GB

These files exclude the memory used by the inference runtime and attention cache. Neither an 8 GB machine nor an 8 GB GPU can hold the full BF16 model in that memory pool. Quantized files need their own measured sizes and a compatible loader; they should not inherit the vendor's unquantized benchmark scores.

Sizes use decimal GB and cover weights only. For the format distinction, see what GGUF is. A memory tier is not listed because minimum runtime memory has not been measured for these builds.

How to run K2 Horizon 3.7B in Atomic Chat

Atomic Chat execution has not been verified for this checkpoint. IFM's official GGUF card requires a llama.cpp build with K2 Horizon architecture support and points to its development fork. Check architecture support before downloading these files. The usual search, download and start-chat flow is not yet a verified procedure for this model. Browse the model catalog for other checkpoints and their documented download options.

K2 Horizon 3.7B license

IFM labels this checkpoint Apache-2.0 in the official model card. Use the publisher's release information as the source for its license, and review any additional notices attached to the exact build you redistribute. A community conversion is not an Atomic Chat release.

Sources and file inventories checked September 15, 2026.

Get the weights from Hugging Face

# Download only; check disk space first.
hf download IFM/K2-Horizon-3.7B --revision adcaf2678bb39e76690387094193caea8c5f6b2d

For a compatible local server on port 8000:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"IFM/K2-Horizon-3.7B","messages":[{"role":"user","content":"Write tests for a function that parses a CSV row."}],"temperature":1,"top_p":0.95,"max_tokens":32768,"chat_template_kwargs":{"reasoning_effort":"high"}}'

Install the OpenAI Python client and start a compatible local server first.

from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="local")
result = client.chat.completions.create(
    model="IFM/K2-Horizon-3.7B",
    messages=[{"role": "user", "content": "Write tests for a CSV parser."}],
    temperature=1.0, top_p=0.95, max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
print(result.choices[0].message.content)

Requires a compatible local server and the OpenAI JavaScript client.

import OpenAI from "openai";
const client = new OpenAI({baseURL: "http://127.0.0.1:8000/v1", apiKey: "local"});
const result = await client.chat.completions.create({
  "model": "IFM/K2-Horizon-3.7B",
  "messages": [
    {
      "role": "user",
      "content": "Write tests for a function that parses a CSV row."
    }
  ],
  "temperature": 1,
  "top_p": 0.95,
  "max_tokens": 32768,
  "chat_template_kwargs": {
    "reasoning_effort": "high"
  }
});
console.log(result.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

K2 Horizon 3.7B is IFM's small dense model for text reasoning and agent tasks. It offers a 512K context window and publishes results for repository repair, terminal work and function calling. This size is worth testing when a compact model misses instructions but a larger download is difficult to accommodate.

Execution of this checkpoint has not been verified in Atomic Chat. The official GGUF requires a llama.cpp build with K2 Horizon support. A downloadable GGUF does not by itself confirm that the app can load it.

Its original BF16 weight files total 10.12 GB. That is disk usage, not a tested minimum RAM or VRAM figure. You also need memory for the inference engine and a cache that grows with context.

IFM describes a 3.7B core, while the complete uploaded tensor inventory contains about 5.058B parameters. The Safetensors files total 10.12 GB at BF16 precision. The official GGUF filename uses 4B, but its declared source remains K2 Horizon 3.7B.

IFM labels this checkpoint Apache-2.0 in the official model card. Use the publisher's release information as the source for its license, and review any additional notices attached to the exact build you redistribute. A community conversion is not an Atomic Chat release.