Cookies on this website

We use cookies to ensure that we give you the best experience on our website. If you click 'Accept all cookies' we'll assume that you are happy to receive all cookies and you won't see this message again. If you click 'Reject all non-essential cookies' only necessary cookies providing core functionality such as security, network management, and accessibility will be enabled. Click 'Find out more' for information on how to change your cookie settings.

Medical Video Large Language Models (Medical Video LLMs) typically process only a limited number of video frames because of computational constraints, making temporal frame selection a critical compo- nent of efficient medical video understanding. Existing approaches pri- marily rely on fixed sampling strategies or generic frame importance estimation, while overlooking the rich supervisory signals already avail- able in medical video annotations. In this work, we investigate adaptive frame selection through Task-Derived Evidence Supervision (TDES), a framework that converts existing task annotations into frame-level rele- vance supervision for training a lightweight Task-Aware Adaptive Video Sampler (TAVS). The learned sampler operates independently of the downstream Medical Video LLM and can therefore be integrated into existing inference pipelines without architectural modification. We con- duct a controlled comparison of evidence-guided and conventional frame selection strategies on the MedVidBench benchmark using two repre- sentative Medical Video LLM backbones, Qwen2.5-VL-7B-Instruct and UAI-NEXUS-MedVLM-1.0a-7B-RL, and compare them with uniform, heuristic, contiguous-window, and learned sampling baselines, as well as prompt-based temporal metadata augmentation. Our experiments show that evidence-guided frame selection yields the strongest Surgical Assess- ment performance on Qwen while remaining competitive across other tasks. In contrast, MedVLM exhibits remarkable robustness to differ- ent frame selection strategies, and explicit temporal metadata provides only marginal benefit for either backbone. Rather than identifying a universally superior frame selection strategy, our findings reveal that the effectiveness of adaptive temporal sampling depends strongly on the underlying Medical Video LLM backbone. This study suggests that fu- ture progress in temporal sampling should consider both the supervision strategy and the characteristics of the downstream backbone.

More information

Type

Conference paper

Publication Date

2026-08-19T00:00:00+00:00