Granite-4.0-H-Small

Updated
05.10.2026
Reasoning
Code
Multilingual
Tools

IBM’s 32B (9B active) hybrid Mamba-2/MoE instruct model with 128K context, strong tool-calling and multilingual support, under Apache 2.0.

At a glance

  • License: Apache 2.0
  • Parameters: 32B total, 9B active
  • Context length: 128K tokens
  • Modalities: Text
  • Minimum hardware: 16 GB of memory (Q2_K build, 11.78 GB)

What is Granite-4.0-H-Small?

Granite-4.0-H-Small is a 32B parameter long-context instruct model from IBM's Granite team, the largest of the four Granite 4.0 language models released on October 2, 2025 under Apache 2.0. It is a mixture-of-experts model with 9B parameters active per token, and 36 of its 40 layers are Mamba2 rather than attention. IBM designed it to respond to general instructions and to build AI assistants for multiple domains, including business applications, across 12 supported languages. IBM publishes its own GGUF builds, so the path to running it locally is short.

SpecificationGranite-4.0-H-Small
Total parameters32B
Active parameters9B
ArchitectureDecoder-only MoE transformer, 4 attention and 36 Mamba2 layers
Embedding size4096, shared input and output embeddings
Attention32 heads, 8 KV heads (GQA), head size 128
Mamba2128 heads, state size 128
Experts72 per MoE block, 10 active, plus a shared expert
Expert hidden size768, shared expert 1536
ActivationSwiGLU, RMSNorm
Context window128K tokens
ModalitiesText
Languages12: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, Chinese
Position embeddingNone (NoPE)
Base modelGranite-4.0-H-Small-Base
Trained onNVIDIA GB200 NVL72 cluster at CoreWeave
Release dateOctober 2, 2025
LicenseApache 2.0

Only 4 of the 40 layers are attention: 32 heads over 8 KV heads with GQA, head size 128. The other 36 are Mamba2 layers, 128 heads with a state size of 128. The feed-forward blocks are mixture-of-experts: 72 experts per block with 10 of them active, next to a shared expert of hidden size 1536. IBM lists the position embedding as NoPE, no positional encoding at all, where the attention-only 3B Micro Dense in the same family uses RoPE. The instruct model comes from Granite-4.0-H-Small-Base by supervised finetuning, reinforcement learning alignment and model merging, over permissively licensed public datasets, internal synthetic data and human-curated examples.

Granite-4.0-H-Small benchmarks

IBM's numbers, from the model card, compare the four Granite 4.0 models: the 32B H Small MoE against the 7B H Tiny MoE and the two 3B models, H Micro Dense and the attention-only Micro Dense:

BenchmarkGranite-4.0-H-SmallH TinyH MicroMicro Dense
MMLU
General knowledge
78.4468.6567.4365.98
MMLU-Pro
Academic knowledge
55.4744.9443.4844.5
GPQA
Expert science
40.6332.5932.1530.14
IFEval
Instruction following
87.5581.4484.3282.31
GSM8K
Grade-school math
87.2784.6981.3585.45
HumanEval
Python coding
88838180
BFCL v3
Function calling
64.6957.6557.5659.98
MGSM
Multilingual math
38.7245.3644.4828.56

H-Small takes every row here but the last. MGSM is the exception: on 8-shot math across five languages the 7B H Tiny scores 45.36 against 38.72. IBM notes separately that its instruction data is mostly English, so multilingual performance may not match English tasks.

IBM lists the capabilities directly: summarization, text classification, text extraction, question answering, retrieval augmented generation, code related tasks, function calling, multilingual dialog and fill-in-the-middle code completions. IBM says the Granite 4.0 instruct models feature improved instruction following and tool-calling, which makes them more effective in enterprise applications. To call tools you declare them with OpenAI's function definition schema, pass them into the chat template, and the model answers with a JSON object inside <tool_call> tags. The reference stack on the card is plain Transformers. An October 7, 2025 update added a default system prompt to the chat template, steering answers toward professional, accurate and safe responses.

Granite-4.0-H-Small hardware requirements

The system requirement to check is memory. IBM publishes its own quantized builds at ibm-granite/granite-4.0-h-small-GGUF, fourteen quantized files from an 11.78 GB Q2_K up to a 34.26 GB Q8_0, next to the 64.45 GB f16.

MemoryBuild to pickFile size
16 GBQ2_K11.78 GB
20 GBQ3_K_M15.36 GB
24 GBQ4_K_M19.48 GB
28 GBQ5_K_M22.87 GB
32 GBQ6_K26.47 GB
48 GB and upQ8_034.26 GB

The listed size is the weights alone, so every row leaves a few gigabytes on top for context and the rest of the system. Within that budget, when two builds both fit, take the larger one. If the format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else runs at that size.

How to run Granite-4.0-H-Small in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for Granite-4.0-H-Small in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

The smaller siblings run the same way; see every Granite model you can run locally, or the rest of the local model catalogue.

Granite-4.0-H-Small license

Granite-4.0-H-Small is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, and IBM says users may finetune Granite 4.0 models for languages beyond the supported 12. IBM adds one caveat: the model was aligned with safety in consideration but may still produce inaccurate, biased or unsafe responses, so run your own safety testing and tuning for your task.

Get the weights from Hugging Face

pip install -U transformers accelerate torch
huggingface-cli download ibm-granite/granite-4.0-h-small
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ibm-granite/granite-4.0-h-small",
    "messages": [{"role": "user", "content": "What is retrieval augmented generation?"}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "ibm-granite/granite-4.0-h-small"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(model_path, device_map="cuda")
chat = [{"role": "user", "content": "Summarize the Apache 2.0 license in one sentence."}]
prompt = tokenizer.apply_chat_template(chat, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(out[0]))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "not-needed" });
const res = await client.chat.completions.create({
  model: "ibm-granite/granite-4.0-h-small",
  messages: [{ role: "user", content: "Explain mixture-of-experts in two sentences." }],
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

Granite-4.0-H-Small is a Mixture-of-Experts model with 32 billion total parameters, of which roughly 9 billion are active per token during inference. It routes through 72 experts with 10 active at a time, keeping compute and memory closer to a 9B dense model than a full 32B one.

Granite-4.0-H-Small has a 128K-token context window. IBM trained the Granite 4.0 models on samples up to 512K tokens and validated performance up to 128K, so it handles long documents, multi-file code, and extended RAG inputs comfortably.

Yes. Granite-4.0-H-Small is released by IBM under the Apache 2.0 license, which permits commercial use, modification, and redistribution. The weights are publicly downloadable from Hugging Face, and the Granite 4.0 family is cryptographically signed and certified under ISO 42001.

Because only 9B of its 32B parameters are active, the hybrid Mamba-2/transformer design cuts memory use sharply. A 4-bit quantized build runs on a single 24 GB GPU; the full BF16 weights need roughly 64 GB. IBM cites over 70% lower memory and about 2x faster inference than comparable dense models in long-context use.

It officially supports 12 languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese. Most training data is English, so non-English performance can lag, but few-shot examples help. Users may also finetune it for additional languages.