GroundAnythingVLM
Visual primitives for physical intelligence.
Media Type
Image
Task Type
ModelGroundAnything-VLMMax output tokens1024Read-only runtime configuration. Only the current image and task prompt enter inference.
Checking inference service…
Open-Vocabulary Grounding
Turn category names into visual primitives: what is present, and where it is.
Separate categories with commas. Select a result to inspect its original pixels.Image ready
Enter to run · Shift + Enter for a new line
Explore a visual targetChoose an example, run grounding, then select a category to inspect.
Explore Examples
Curated visual scenariosTRAINING SOURCEdemo_groundingIMAGE SHA256
c6abef94a28290caCurated research examples for qualitative exploration, not a held-out benchmark. Inference receives only the image and task prompt.Inference Metrics
- Status
- Ready
- Task
- Grounding
- Model
- GroundAnything-VLM
- Detections
- 0
- Latency
- —
- Output tokens
- —
Target Detail
Select a category below the canvas to magnify its original pixels.Detail ViewOriginal pixels · magnified
Transparent inputs: only this image and compiled prompt are sent to the model. No annotations, training answers, or expected counts are included.