JumpScore · repeated action localization
- 74.9Glint
- 30.1Qwen
- 39.6Keye
- 11.0Intern
- 13.1PLM
- 2.1LLaVA
An 8B vision-language model for long-video understanding, temporal localization, and spatial reasoning
Glint-VL-Video-8B is an 8B vision-language model that handles images and videos in one architecture. It supports video QA, temporal localization, spatial reasoning, and video object segmentation and tracking—explaining what happened, when it happened, and where the target appeared.
Its codec-stream input reads I/P frame structure, motion vectors, and residual signals from the compressed stream, spending a limited visual-token budget on frames that contain meaningful events and motion.
Allocates visual tokens by event density so important actions retain detail.
Localizes the correct occurrence among dense, repeated actions.
Video QA, temporal localization, spatial reasoning, and tracking share one parameter set.
codec-stream affects only visual encoder input and fits existing video-understanding pipelines.
The comparison set contains open-source 8B-class models. See how Glint-VL-Video-8B compares with five peer models on each benchmark.
Benchmark scores range from 0 to 100; higher is better.
Understand meetings, surveillance clips, or lessons and answer content and event questions.
Return second-level timestamps for a natural-language event description.
Distinguish a specific repetition in sports, exercise, or production lines.
Answer direction, spatial relation, count, and trajectory questions.
Segment and track a text-referred target throughout a video.
Handle common image, chart, document, and counting tasks.