What is LocateAnything-3B?
LocateAnything-3B is a 3.8B vision-language model from NVIDIA built for visual grounding: you give it an image and a plain-English request, and it returns exact bounding boxes or points for whatever you named. It pairs a MoonViT vision encoder with a Qwen2.5-3B-Instruct language model and belongs to NVIDIA's Eagle VLM family. The same grounding stack has been folded into NVIDIA's production models, including Nemotron 3 Nano Omni; for a general-purpose NVIDIA model to run alongside it, see Nemotron Nano 9B v2. NVIDIA published the weights on May 26, 2026 as a research release.
| Specification | LocateAnything-3B |
|---|---|
| Total parameters | 3.8B |
| Architecture | MoonViT vision encoder + Qwen2.5-3B-Instruct language model, MLP projector |
| Core mechanism | Parallel Box Decoding, up to 2.5x higher throughput |
| Input | RGB image up to 2.5K resolution, prompt up to 24K tokens |
| Output | Text with box and point coordinates, up to 8,192 tokens |
| Generation modes | Fast (parallel), Slow (autoregressive), Hybrid (default) |
| Training data | 12M images, 138M+ queries, 785M bounding boxes |
| Release date | May 26, 2026 |
| License | NVIDIA License, non-commercial |
The distinctive part is Parallel Box Decoding. A standard VLM writes box coordinates one token at a time; LocateAnything predicts a complete box as one fixed-length block in a single parallel step, which NVIDIA measures at up to 2.5x the throughput of prior approaches. The default hybrid mode decodes in parallel and falls back to token-by-token decoding when a box is spatially ambiguous or the output format goes irregular, the balance of speed and accuracy NVIDIA recommends. Coordinates come back as normalized integers from 0 to 1000, so a few lines of parsing turn any answer into pixel positions at the original image size.
What LocateAnything-3B is good at
NVIDIA trained the model on 12M images spanning natural scenes, robotics, driving, GUI interaction and document understanding, and the supported task list is wide for a 3.8B model: open-set and long-tail object detection, dense multi-object detection in cluttered scenes, phrase and referring-expression grounding ("people wearing red shirts"), scene text detection and OCR localization, document layout grounding, and point-based localization.
Two uses stand out. GUI element grounding gives an agent the on-screen coordinates of a button or field from a natural-language instruction, the piece a computer-use pipeline needs before it can click anything. And automated dataset labeling: the model boxes or points at everything matching a description, which NVIDIA lists as a supported use for annotation at scale. Robotics perception, autonomous driving, industrial inspection and remote sensing round out the vendor's list.
LocateAnything-3B hardware requirements
The system requirement to check is memory. NVIDIA's own instructions run the BF16 checkpoint through Transformers on Linux; for llama.cpp-based apps, the community repo yuuko-eth/LocateAnything-3B-GGUF carries the conversions, and every build also loads a 0.87 GB mmproj vision file next to the main one.
| Memory | Build to pick | File size |
|---|---|---|
| 4 GB | Q4_K_M | 2.11 GB |
| 6 GB | Q6_K | 2.80 GB |
| 8 GB | Q8_0 | 3.62 GB |
| 12 GB and up | BF16 | 6.81 GB |
When two builds both fit, take the larger one. If the format is new to you, start with what GGUF is.
How to run LocateAnything-3B in Atomic Chat
Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.
- Download Atomic Chat for your platform and open it.
- Search for LocateAnything-3B in the model browser and open Download Options.
- Pick the build that fits the memory you have, then start a chat.
For the family overview, see every LocateAnything model you can run locally.
LocateAnything-3B license
LocateAnything-3B ships under the NVIDIA License, which permits use, reproduction and modification for academic and non-profit research only; commercial use is not permitted except by NVIDIA and its affiliates, and redistribution must keep the license and attribution notices. Two components carry their own terms: the Qwen2.5-3B-Instruct language model (Qwen Research License) and the MoonViT vision encoder (MIT). Treat it as a model to prototype and research with, not one to ship inside a commercial product.
