dataset link on Hugging dataset
https://github.com/OmkarThawakar/BSE-CoVR
Arxiv link
https://arxiv.org/abs/2508.14039
Description of the dataset
This task evaluates dense, fine-grained compositional retrieval using video inputs. Given a source reference video and a text query detailing a specific modification or edit to the scene/action, the model must accurately retrieve the target video that matches the combined video-and-text context.
dataset link on Hugging dataset
https://github.com/OmkarThawakar/BSE-CoVR
Arxiv link
https://arxiv.org/abs/2508.14039
Description of the dataset
This task evaluates dense, fine-grained compositional retrieval using video inputs. Given a source reference video and a text query detailing a specific modification or edit to the scene/action, the model must accurately retrieve the target video that matches the combined video-and-text context.