dataset link on Hugging dataset
https://imagine.enpc.fr/~ventural/covr/
Arxiv link
https://arxiv.org/abs/2308.14746
Description of the dataset
This task focuses on retrieving videos based on a multimodal query. The query consists of a static reference image (specifically, the middle frame of a video) paired with a text prompt that specifies an action, temporal change, or modification. The model must synthesize these two inputs to retrieve the correct corresponding video.
dataset link on Hugging dataset
https://imagine.enpc.fr/~ventural/covr/
Arxiv link
https://arxiv.org/abs/2308.14746
Description of the dataset
This task focuses on retrieving videos based on a multimodal query. The query consists of a static reference image (specifically, the middle frame of a video) paired with a text prompt that specifies an action, temporal change, or modification. The model must synthesize these two inputs to retrieve the correct corresponding video.