OpenELM-1_1B-Instruct

Updated
05.10.2026
Reasoning

Apple’s 1.1B instruction-tuned OpenELM model, built with layer-wise scaling for efficient on-device English text generation.

At a glance

  • License: Apple Sample Code License (apple-amlr)
  • Parameters: 1.1B
  • Context length: 2,048 tokens
  • Modalities: Text in, text out
  • Minimum hardware: 2.16 GB bf16 checkpoint, fits an 8 GB machine

What is OpenELM-1_1B-Instruct?

OpenELM-1_1B-Instruct is the 1.1 billion parameter instruction-tuned member of OpenELM, Apple's family of Open Efficient Language Models, published in April 2024 in four sizes: 270M, 450M, 1.1B and 3B, each as a pretrained and an instruction-tuned checkpoint. The release is unusually complete for a corporate lab: weights, the CoreNet training library, data preparation, fine-tuning and evaluation code, and the training logs all ship together. For local use the point is size: the bf16 checkpoint is a 2.16 GB file, small enough for machines where every gigabyte of memory counts.

SpecificationOpenELM-1_1B-Instruct
Parameters1.1B (1,079,891,456 in the safetensors)
ArchitectureDecoder-only transformer with layer-wise scaling, 28 layers
AttentionGrouped-query attention, 16 to 32 query heads and 4 to 8 KV heads growing with depth
Context window2,048 tokens
ModalitiesText in, text out
Training dataAbout 1.8 trillion tokens: RefinedWeb, deduplicated PILE, RedPajama and Dolma v1.6 subsets
TokenizerLlama-2 tokenizer, 32,000 vocabulary
Release dateApril 2024
LicenseApple Sample Code License (apple-amlr)

The distinctive design choice is layer-wise scaling. Instead of giving all 28 transformer layers the same shape, OpenELM widens them with depth: the first layer runs 16 query heads and an FFN multiplier of 0.5, the last runs 32 heads at a multiplier of 4.0, so parameters concentrate in the later layers where they buy the most accuracy at this size. Input and output embeddings are shared, and query and key projections are normalized before attention.

OpenELM-1_1B-Instruct benchmarks

Apple's zero-shot numbers from the model card put the instruct model against its own base checkpoint and the two neighbouring sizes in the family:

BenchmarkOpenELM-1_1B-InstructOpenELM-1_1BOpenELM-450M-InstructOpenELM-3B-Instruct
ARC-c
Science reasoning
37.9732.3430.3839.42
ARC-e
Easier science
52.2355.4350.0061.74
BoolQ
Yes/no questions
70.0063.5860.3768.17
HellaSwag
Commonsense completion
71.2064.8159.3476.36
PIQA
Physical commonsense
75.0375.5772.6379.00
SciQ
Science exam
89.3090.6088.0092.50
WinoGrande
Pronoun resolution
62.7561.7258.9666.85

Instruction tuning lifts the 1.1B from a 63.44 zero-shot average to 65.50, and it takes BoolQ from the 3B outright. Everywhere else the 3B-Instruct stays ahead, so pick the 1.1B for the memory it saves, not for peak scores.

OpenELM-1_1B-Instruct hardware requirements

The system requirement to check is memory. There is no GGUF build: llama.cpp does not support the OpenELM architecture, so the file to plan around is the bf16 safetensors checkpoint from apple/OpenELM-1_1B-Instruct, which Apple serves through Hugging Face Transformers with trust_remote_code enabled and the Llama-2 tokenizer.

MemoryBuild to pickFile size
8 GB and upmodel.safetensors (bf16)2.16 GB

One file covers every machine, so there is no quant ladder to pick through; community MLX and CoreML conversions exist for Apple silicon. Apple's generate_openelm.py script also supports prompt-lookup and assistant-model generation, the mechanism behind speculative decoding, by passing a smaller model as the draft.

How to run OpenELM-1_1B-Instruct in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for OpenELM-1_1B-Instruct in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, from 270M to 3B, see every OpenELM model you can run locally, or compare it with another small local model like SmolLM3-3B.

OpenELM-1_1B-Instruct license

OpenELM-1_1B-Instruct is released under the Apple Sample Code License (apple-amlr), Apple's own terms rather than a standard permissive license, so read the LICENSE file in the repository before building a product on it. Apple ships the models without safety guarantees and expects users to run their own safety testing and filtering, and it asks you to check the license agreements of the pretraining datasets before using them.

Get the weights from Hugging Face

pip install -U transformers sentencepiece
huggingface-cli download apple/OpenELM-1_1B-Instruct
python generate_openelm.py --model apple/OpenELM-1_1B-Instruct --prompt 'Once upon a time there was' --generate_kwargs repetition_penalty=1.2
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "apple/OpenELM-1_1B-Instruct",
    "messages": [{"role": "user", "content": "Write a short poem about the sea."}]
  }'
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("apple/OpenELM-1_1B-Instruct", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-2-7b-hf")
inputs = tokenizer("Once upon a time there was", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=128, repetition_penalty=1.2)
print(tokenizer.decode(out[0], skip_special_tokens=True))
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "local" });
const res = await client.chat.completions.create({
  model: "apple/OpenELM-1_1B-Instruct",
  messages: [{ role: "user", content: "Write a short poem about the sea." }]
});
console.log(res.choices[0].message.content);
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

OpenELM-1_1B-Instruct is a 1.08-billion-parameter instruction-tuned language model released by Apple in April 2024. It belongs to the OpenELM family, which uses a layer-wise scaling strategy to allocate parameters efficiently across transformer layers. Apple released it alongside the full training and evaluation framework built on the CoreNet library.

In BF16 precision the model's weights are about 2.2 GB, so it runs on a GPU with roughly 4 GB of VRAM and fits comfortably on most modern cards. With 4-bit quantization the memory footprint drops below 1 GB, and the model is small enough to run on CPU or Apple Silicon for testing.

The weights are openly downloadable from Hugging Face, but the model is distributed under the Apple Sample Code License (apple-amlr) rather than a standard permissive license like Apache 2.0 or MIT. That license is more restrictive than typical open-source terms, so review it before any commercial deployment. Apple did release the complete training, fine-tuning, and evaluation code to support open research.

Load it with Hugging Face Transformers using trust_remote_code=True, since OpenELM ships custom modeling code. The model uses the Llama-2 tokenizer (meta-llama/Llama-2-7b-hf) and needs add_bos_token=True. Apple also provides a generate_openelm.py helper script in the repo, and passing repetition_penalty=1.2 is recommended for cleaner generations.

OpenELM-1_1B-Instruct has a maximum context length of 2,048 tokens, set by the max_context_length value in its configuration. That window is short by current standards and suits short prompts, single-turn instructions, and lightweight chat rather than long-document tasks. The model was trained primarily on English data, so it is best used for English text.