TRECVID 2023 · Ad-hoc Video Search
Finding the right moment
Overview
Combining vision-language models, diffusion-generated image queries, and relevance feedback for ad-hoc video search.
Approach
The system combines CLIP-family models with generated image queries and ranking fusion. Relevance feedback adjusts the retrieval process through interactive re-ranking.
My contribution
Led the AVS project and contributed to the retrieval pipeline, experiments, technical report, and presentation of the approach.
Outcome
WHU-NERCMS ranked first in the automatic and interactive AVS tracks at TRECVID 2023.