GroundAnythingVLM

Visual primitives for physical intelligence.

Media Type
Image
Task Type
ModelGroundAnything-VLMMax output tokens1024Read-only runtime configuration. Only the current image and task prompt enter inference.
Checking inference service…
Open-Vocabulary Grounding

Turn category names into visual primitives: what is present, and where it is.

Separate categories with commas. Select a result to inspect its original pixels.
Image ready
Image canvas
100%

Enter to run · Shift + Enter for a new line

Explore a visual targetChoose an example, run grounding, then select a category to inspect.

Explore Examples

Curated visual scenarios
TRAINING SOURCEdemo_groundingIMAGE SHA256c6abef94a28290caCurated research examples for qualitative exploration, not a held-out benchmark. Inference receives only the image and task prompt.

Inference Metrics

Status
Ready
Task
Grounding
Model
GroundAnything-VLM
Detections
0
Latency
—
Output tokens
—
Transparent inputs: only this image and compiled prompt are sent to the model. No annotations, training answers, or expected counts are included.
GroundAnything-VLM · Towards Physical IntelligenceResearch demonstration Qualitative examples, not benchmark results.