Abstract: Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results