VibeThinker-3B

Updated
05.10.2026
Thinking
Reasoning
Code

VibeThinker-3B is WeiboAI’s 3.1B reasoning model: 76.4 on IMO-AnswerBench, in range of 671B to 1T models, and 96.1% on unseen LeetCode contests.

At a glance

  • License: MIT
  • Parameters: 3.1B
  • Context length: 64K RL training window; quick start generates up to 102,400 tokens
  • Modalities: Text in, text out (English)
  • Minimum hardware: ~4 GB memory for the 1.93 GB Q4_K_M GGUF

What is VibeThinker-3B?

VibeThinker-3B is a 3.1B-parameter reasoning model from WeiboAI, fine-tuned from Qwen2.5-Coder-3B and aimed at one job: multi-step reasoning on problems whose answers can be checked, which in practice means competition math, competitive programming and STEM. WeiboAI published the weights on June 12, 2026 under the MIT license. The case for running it locally is the size: the 4-bit GGUF is a 1.93 GB file, so the model fits in about 4 GB of memory, which covers nearly any laptop.

SpecificationVibeThinker-3B
Total parameters3.1B (3,085,938,688)
Base modelQwen2.5-Coder-3B
ModalitiesText in, text out (English)
Reasoning budgetQuick start defaults to 102,400 new tokens; 60K to 100K advised for the hardest problems
Post-trainingSpectrum-to-Signal Principle: two-stage SFT, multi-domain RL, self-distillation, instruct RL
Tool callingNot trained for it; the vendor advises against agent use
Recommended samplingtemperature 1.0, top_p 0.95
Release dateJune 12, 2026
LicenseMIT

The training recipe is why a 3B model shows up next to frontier names at all. WeiboAI's Spectrum-to-Signal Principle pipeline runs a two-stage curriculum SFT that deliberately preserves multiple valid solution paths, then reinforcement learning with verifiable rewards over math, code and STEM inside a 64K context window, then distills the strongest RL trajectories back into a single student model, and finishes with an instruct RL pass. The authors call the bet behind it the Parametric Compression-Coverage Hypothesis: verifiable reasoning compresses well into few parameters, broad world knowledge does not. They are open about both halves of that trade and recommend larger general-purpose models for open-domain tasks. The card also carries a plain warning: the model was not trained on tool-calling or agent data, so point it at LeetCode-style problems, not at function calling or autonomous coding agents.

VibeThinker-3B benchmarks

The numbers are WeiboAI's own, published on the model card. The headline is IMO-AnswerBench, a benchmark of 400 IMO-level problems, where the 3B lands in the range of models hundreds of times its size:

BenchmarkVibeThinker-3BDeepSeek V3.2GLM-5Kimi K2.5
IMO-AnswerBench
Olympiad math
76.478.382.581.8
IMO-AnswerBench with CLR
Verified reasoning
80.678.382.581.8
LeetCode contests
Contest coding
96.1%---

Read the first two rows with parameter counts in mind: DeepSeek V3.2 is 671B, GLM-5 is 744B and Kimi K2.5 is 1T, and at 3B the raw 76.4 still trails all three. Claim-Level Reliability Assessment, WeiboAI's test-time scaling strategy for answer-verifiable tasks, lifts the same benchmark to 80.6, which passes DeepSeek V3.2 but stays behind GLM-5 and Kimi K2.5. The LeetCode row is not a comparison: it is 123 of 128 first-attempt passes on unseen weekly and biweekly contests from April 25 to May 31, 2026, and WeiboAI publishes no competitor numbers next to it. The card also cites results on AIME, HMMT and LiveCodeBench without publishing a numeric table.

VibeThinker-3B hardware requirements

The system requirement to check is memory. The GGUF builds and file sizes below come from the community repo prithivMLmods/VibeThinker-3B-GGUF.

MemoryBuild to pickFile size
4 GBQ4_K_M1.93 GB
6 GBQ5_K_M2.22 GB
8 GBQ6_K2.54 GB
12 GBQ8_03.29 GB
16 GB and upBF166.18 GB

Neighbouring files differ by a few hundred megabytes, so when two builds both fit, take the larger one. Leave headroom beyond the file itself: WeiboAI advises 60K to 100K token limits for the hardest problems, and a reasoning trace that long needs KV cache on top of the weights. If the format is new to you, start with what GGUF is.

How to run VibeThinker-3B in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for VibeThinker-3B in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the rest of the family, see every VibeThinker model you can run locally, or compare it with another small reasoning model, Qwen3-4B Thinking 2507.

VibeThinker-3B license

VibeThinker-3B is released under the MIT license. That permits commercial use, modification, fine-tuning and redistribution, with the single obligation of keeping the copyright and license notice, so you can build on the model and ship it in a product without a usage fee.

Get the weights from Hugging Face

huggingface-cli download WeiboAI/VibeThinker-3B
from transformers import AutoModel
model = AutoModel.from_pretrained("WeiboAI/VibeThinker-3B")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

VibeThinker-3B is a 3.1B-parameter reasoning model from WeiboAI, post-trained on Qwen2.5-Coder-3B. It is aimed at verifiable tasks like competition math, STEM problems, and code, and it ships open-weight under an MIT license. Despite its small size, WeiboAI reports it competing with much larger models on math benchmarks such as AIME and HMMT.

The Q4_K_M GGUF listed here is about 1.93 GB on disk; that is not the total memory needed to run it. The page uses roughly 4 GB as a starting estimate, with separate room needed for the operating system and other applications. Runtime settings and longer contexts increase memory use, so long reasoning sessions need additional headroom.

Yes. The weights are released under the MIT license, which allows free use including commercial projects, modification, and redistribution. Running it locally in Atomic Chat means there is no subscription, no API key, and no per-token cost.

Yes. Once the model is downloaded it runs fully on your own machine with no internet connection. In Atomic Chat every prompt and response stays on-device, so nothing is sent to an external server.

WeiboAI positions VibeThinker-3B for verifiable reasoning, including competition math, STEM and competitive programming. Its official model card says it was not trained on tool-calling or agent-based programming data and does not recommend it for function calling, API orchestration or autonomous coding agents.