What is OpenELM-1_1B-Instruct?
OpenELM-1_1B-Instruct is the 1.1 billion parameter instruction-tuned member of OpenELM, Apple's family of Open Efficient Language Models, published in April 2024 in four sizes: 270M, 450M, 1.1B and 3B, each as a pretrained and an instruction-tuned checkpoint. The release is unusually complete for a corporate lab: weights, the CoreNet training library, data preparation, fine-tuning and evaluation code, and the training logs all ship together. For local use the point is size: the bf16 checkpoint is a 2.16 GB file, small enough for machines where every gigabyte of memory counts.
| Specification | OpenELM-1_1B-Instruct |
|---|---|
| Parameters | 1.1B (1,079,891,456 in the safetensors) |
| Architecture | Decoder-only transformer with layer-wise scaling, 28 layers |
| Attention | Grouped-query attention, 16 to 32 query heads and 4 to 8 KV heads growing with depth |
| Context window | 2,048 tokens |
| Modalities | Text in, text out |
| Training data | About 1.8 trillion tokens: RefinedWeb, deduplicated PILE, RedPajama and Dolma v1.6 subsets |
| Tokenizer | Llama-2 tokenizer, 32,000 vocabulary |
| Release date | April 2024 |
| License | Apple Sample Code License (apple-amlr) |
The distinctive design choice is layer-wise scaling. Instead of giving all 28 transformer layers the same shape, OpenELM widens them with depth: the first layer runs 16 query heads and an FFN multiplier of 0.5, the last runs 32 heads at a multiplier of 4.0, so parameters concentrate in the later layers where they buy the most accuracy at this size. Input and output embeddings are shared, and query and key projections are normalized before attention.
OpenELM-1_1B-Instruct benchmarks
Apple's zero-shot numbers from the model card put the instruct model against its own base checkpoint and the two neighbouring sizes in the family:
| Benchmark | OpenELM-1_1B-Instruct | OpenELM-1_1B | OpenELM-450M-Instruct | OpenELM-3B-Instruct |
|---|---|---|---|---|
ARC-c Science reasoning | 37.97 | 32.34 | 30.38 | 39.42 |
ARC-e Easier science | 52.23 | 55.43 | 50.00 | 61.74 |
BoolQ Yes/no questions | 70.00 | 63.58 | 60.37 | 68.17 |
HellaSwag Commonsense completion | 71.20 | 64.81 | 59.34 | 76.36 |
PIQA Physical commonsense | 75.03 | 75.57 | 72.63 | 79.00 |
SciQ Science exam | 89.30 | 90.60 | 88.00 | 92.50 |
WinoGrande Pronoun resolution | 62.75 | 61.72 | 58.96 | 66.85 |
Instruction tuning lifts the 1.1B from a 63.44 zero-shot average to 65.50, and it takes BoolQ from the 3B outright. Everywhere else the 3B-Instruct stays ahead, so pick the 1.1B for the memory it saves, not for peak scores.
OpenELM-1_1B-Instruct hardware requirements
The system requirement to check is memory. There is no GGUF build: llama.cpp does not support the OpenELM architecture, so the file to plan around is the bf16 safetensors checkpoint from apple/OpenELM-1_1B-Instruct, which Apple serves through Hugging Face Transformers with trust_remote_code enabled and the Llama-2 tokenizer.
| Memory | Build to pick | File size |
|---|---|---|
| 8 GB and up | model.safetensors (bf16) | 2.16 GB |
One file covers every machine, so there is no quant ladder to pick through; community MLX and CoreML conversions exist for Apple silicon. Apple's generate_openelm.py script also supports prompt-lookup and assistant-model generation, the mechanism behind speculative decoding, by passing a smaller model as the draft.
How to run OpenELM-1_1B-Instruct in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for OpenELM-1_1B-Instruct in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the rest of the family, from 270M to 3B, see every OpenELM model you can run locally, or compare it with another small local model like SmolLM3-3B.
OpenELM-1_1B-Instruct license
OpenELM-1_1B-Instruct is released under the Apple Sample Code License (apple-amlr), Apple's own terms rather than a standard permissive license, so read the LICENSE file in the repository before building a product on it. Apple ships the models without safety guarantees and expects users to run their own safety testing and filtering, and it asks you to check the license agreements of the pretraining datasets before using them.
