What is Granite-4.0-H-Small?
Granite-4.0-H-Small is a 32B parameter long-context instruct model from IBM's Granite team, the largest of the four Granite 4.0 language models released on October 2, 2025 under Apache 2.0. It is a mixture-of-experts model with 9B parameters active per token, and 36 of its 40 layers are Mamba2 rather than attention. IBM designed it to respond to general instructions and to build AI assistants for multiple domains, including business applications, across 12 supported languages. IBM publishes its own GGUF builds, so the path to running it locally is short.
| Specification | Granite-4.0-H-Small |
|---|---|
| Total parameters | 32B |
| Active parameters | 9B |
| Architecture | Decoder-only MoE transformer, 4 attention and 36 Mamba2 layers |
| Embedding size | 4096, shared input and output embeddings |
| Attention | 32 heads, 8 KV heads (GQA), head size 128 |
| Mamba2 | 128 heads, state size 128 |
| Experts | 72 per MoE block, 10 active, plus a shared expert |
| Expert hidden size | 768, shared expert 1536 |
| Activation | SwiGLU, RMSNorm |
| Context window | 128K tokens |
| Modalities | Text |
| Languages | 12: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, Chinese |
| Position embedding | None (NoPE) |
| Base model | Granite-4.0-H-Small-Base |
| Trained on | NVIDIA GB200 NVL72 cluster at CoreWeave |
| Release date | October 2, 2025 |
| License | Apache 2.0 |
Only 4 of the 40 layers are attention: 32 heads over 8 KV heads with GQA, head size 128. The other 36 are Mamba2 layers, 128 heads with a state size of 128. The feed-forward blocks are mixture-of-experts: 72 experts per block with 10 of them active, next to a shared expert of hidden size 1536. IBM lists the position embedding as NoPE, no positional encoding at all, where the attention-only 3B Micro Dense in the same family uses RoPE. The instruct model comes from Granite-4.0-H-Small-Base by supervised finetuning, reinforcement learning alignment and model merging, over permissively licensed public datasets, internal synthetic data and human-curated examples.
Granite-4.0-H-Small benchmarks
IBM's numbers, from the model card, compare the four Granite 4.0 models: the 32B H Small MoE against the 7B H Tiny MoE and the two 3B models, H Micro Dense and the attention-only Micro Dense:
| Benchmark | Granite-4.0-H-Small | H Tiny | H Micro | Micro Dense |
|---|---|---|---|---|
MMLU General knowledge | 78.44 | 68.65 | 67.43 | 65.98 |
MMLU-Pro Academic knowledge | 55.47 | 44.94 | 43.48 | 44.5 |
GPQA Expert science | 40.63 | 32.59 | 32.15 | 30.14 |
IFEval Instruction following | 87.55 | 81.44 | 84.32 | 82.31 |
GSM8K Grade-school math | 87.27 | 84.69 | 81.35 | 85.45 |
HumanEval Python coding | 88 | 83 | 81 | 80 |
BFCL v3 Function calling | 64.69 | 57.65 | 57.56 | 59.98 |
MGSM Multilingual math | 38.72 | 45.36 | 44.48 | 28.56 |
H-Small takes every row here but the last. MGSM is the exception: on 8-shot math across five languages the 7B H Tiny scores 45.36 against 38.72. IBM notes separately that its instruction data is mostly English, so multilingual performance may not match English tasks.
IBM lists the capabilities directly: summarization, text classification, text extraction, question answering, retrieval augmented generation, code related tasks, function calling, multilingual dialog and fill-in-the-middle code completions. IBM says the Granite 4.0 instruct models feature improved instruction following and tool-calling, which makes them more effective in enterprise applications. To call tools you declare them with OpenAI's function definition schema, pass them into the chat template, and the model answers with a JSON object inside <tool_call> tags. The reference stack on the card is plain Transformers. An October 7, 2025 update added a default system prompt to the chat template, steering answers toward professional, accurate and safe responses.
Granite-4.0-H-Small hardware requirements
The system requirement to check is memory. IBM publishes its own quantized builds at ibm-granite/granite-4.0-h-small-GGUF, fourteen quantized files from an 11.78 GB Q2_K up to a 34.26 GB Q8_0, next to the 64.45 GB f16.
| Memory | Build to pick | File size |
|---|---|---|
| 16 GB | Q2_K | 11.78 GB |
| 20 GB | Q3_K_M | 15.36 GB |
| 24 GB | Q4_K_M | 19.48 GB |
| 28 GB | Q5_K_M | 22.87 GB |
| 32 GB | Q6_K | 26.47 GB |
| 48 GB and up | Q8_0 | 34.26 GB |
The listed size is the weights alone, so every row leaves a few gigabytes on top for context and the rest of the system. Within that budget, when two builds both fit, take the larger one. If the format is new to you, start with what GGUF is, and see the best local LLMs for a 16 GB Mac for what else runs at that size.
How to run Granite-4.0-H-Small in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for Granite-4.0-H-Small in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
The smaller siblings run the same way; see every Granite model you can run locally, or the rest of the local model catalogue.
Granite-4.0-H-Small license
Granite-4.0-H-Small is released under Apache 2.0. That permits commercial use, modification and redistribution with no royalties, and IBM says users may finetune Granite 4.0 models for languages beyond the supported 12. IBM adds one caveat: the model was aligned with safety in consideration but may still produce inaccurate, biased or unsafe responses, so run your own safety testing and tuning for your task.
