LocateAnything-3B

Updated
05.10.2026
Thinking
Vision
Reasoning
Code

LocateAnything-3B is NVIDIA’s visual grounding model for locating objects, GUI elements and text with bounding boxes or points.

At a glance

  • License: NVIDIA License, non-commercial research only
  • Parameters: 3.8B
  • Context length: 24K-token prompts, up to 8,192 output tokens
  • Modalities: Image and text input, text output with box and point coordinates
  • Minimum hardware: 4 GB of memory (Q4_K_M GGUF, 2.11 GB plus a 0.87 GB vision file)

What is LocateAnything-3B?

LocateAnything-3B is a 3.8B vision-language model from NVIDIA built for visual grounding: you give it an image and a plain-English request, and it returns exact bounding boxes or points for whatever you named. It pairs a MoonViT vision encoder with a Qwen2.5-3B-Instruct language model and belongs to NVIDIA's Eagle VLM family. The same grounding stack has been folded into NVIDIA's production models, including Nemotron 3 Nano Omni; for a general-purpose NVIDIA model to run alongside it, see Nemotron Nano 9B v2. NVIDIA published the weights on May 26, 2026 as a research release.

SpecificationLocateAnything-3B
Total parameters3.8B
ArchitectureMoonViT vision encoder + Qwen2.5-3B-Instruct language model, MLP projector
Core mechanismParallel Box Decoding, up to 2.5x higher throughput
InputRGB image up to 2.5K resolution, prompt up to 24K tokens
OutputText with box and point coordinates, up to 8,192 tokens
Generation modesFast (parallel), Slow (autoregressive), Hybrid (default)
Training data12M images, 138M+ queries, 785M bounding boxes
Release dateMay 26, 2026
LicenseNVIDIA License, non-commercial

The distinctive part is Parallel Box Decoding. A standard VLM writes box coordinates one token at a time; LocateAnything predicts a complete box as one fixed-length block in a single parallel step, which NVIDIA measures at up to 2.5x the throughput of prior approaches. The default hybrid mode decodes in parallel and falls back to token-by-token decoding when a box is spatially ambiguous or the output format goes irregular, the balance of speed and accuracy NVIDIA recommends. Coordinates come back as normalized integers from 0 to 1000, so a few lines of parsing turn any answer into pixel positions at the original image size.

What LocateAnything-3B is good at

NVIDIA trained the model on 12M images spanning natural scenes, robotics, driving, GUI interaction and document understanding, and the supported task list is wide for a 3.8B model: open-set and long-tail object detection, dense multi-object detection in cluttered scenes, phrase and referring-expression grounding ("people wearing red shirts"), scene text detection and OCR localization, document layout grounding, and point-based localization.

Two uses stand out. GUI element grounding gives an agent the on-screen coordinates of a button or field from a natural-language instruction, the piece a computer-use pipeline needs before it can click anything. And automated dataset labeling: the model boxes or points at everything matching a description, which NVIDIA lists as a supported use for annotation at scale. Robotics perception, autonomous driving, industrial inspection and remote sensing round out the vendor's list.

LocateAnything-3B hardware requirements

The system requirement to check is memory. NVIDIA's own instructions run the BF16 checkpoint through Transformers on Linux; for llama.cpp-based apps, the community repo yuuko-eth/LocateAnything-3B-GGUF carries the conversions, and every build also loads a 0.87 GB mmproj vision file next to the main one.

MemoryBuild to pickFile size
4 GBQ4_K_M2.11 GB
6 GBQ6_K2.80 GB
8 GBQ8_03.62 GB
12 GB and upBF166.81 GB

When two builds both fit, take the larger one. If the format is new to you, start with what GGUF is.

How to run LocateAnything-3B in Atomic Chat

Atomic Chat is a free local app for macOS, Windows and Linux. It includes a Hugging Face model browser and a built-in chat, with no manual llama.cpp build required.

  1. Download Atomic Chat for your platform and open it.
  2. Search for LocateAnything-3B in the model browser and open Download Options.
  3. Pick the build that fits the memory you have, then start a chat.

For the family overview, see every LocateAnything model you can run locally.

LocateAnything-3B license

LocateAnything-3B ships under the NVIDIA License, which permits use, reproduction and modification for academic and non-profit research only; commercial use is not permitted except by NVIDIA and its affiliates, and redistribution must keep the license and attribution notices. Two components carry their own terms: the Qwen2.5-3B-Instruct language model (Qwen Research License) and the MoonViT vision encoder (MIT). Treat it as a model to prototype and research with, not one to ship inside a commercial product.

Get the weights from Hugging Face

huggingface-cli download nvidia/LocateAnything-3B
from transformers import AutoModel
model = AutoModel.from_pretrained("nvidia/LocateAnything-3B")
Desktop
macOS
(Intel and Apple Silicon)
Download
Windows
(x64)
Download
Linux
(x86_64)
Download

Frequently asked questions

LocateAnything-3B is a 3.8B-parameter vision-language model from NVIDIA, built on its Eagle line, for visual grounding and open-vocabulary object detection. You give it an image and a text prompt, and it returns the exact locations of the objects you named as bounding boxes or points. It handles open-ended categories, referring expressions, GUI elements, and OCR text rather than a fixed list of classes.

In FP16 it needs roughly 8.4 GB of VRAM for inference, covering weights, activations, and KV cache. Quantized to INT4 that falls to about 2.1 GB, which lets an 8 GB card like an RTX 4060 run it. The model uses BF16, so you need an Ampere-or-newer NVIDIA GPU (RTX 30/40/50-series, A100, or H100).

The weights are free to download from Hugging Face. It is released under the NVIDIA License (non-commercial), which allows use, modification, and reproduction for academic and non-profit research only. Commercial use is not permitted, so review the license before using it in a paid product.

Yes. Once you download the weights with huggingface-cli, the model runs entirely on your own hardware with no internet connection needed. In Atomic Chat it runs on-device, so the images you analyze and the prompts you write never leave your machine. There is also a C++ ggml port (locate-anything.cpp) for CPU-only offline inference without a Python runtime.

Unlike a fixed-class detector such as YOLO, it does open-vocabulary detection: you can name any category in plain text and it finds it, no retraining required. It performs strongly in dense scenes and on long-tail objects, and it can resolve referring expressions like "the person on the left holding a phone." That makes it a fit for GUI agents, robotics, and document-understanding pipelines that need spatial coordinates from language.