Document type
Bachelor thesisPublication date
Publication license
Please use this identifier to cite or link to this item: https://hdl.handle.net/2445/203127
Efficient transformers applied to video classification
Journal Title
Authors
Director/Tutor
Journal ISSN
Volume Title
Related resource
Abstract
[en] Transformers, with the self-attention mechanism on its core, have shown great performance on several Machine Learning areas such as NLP or Computer Vision since its appearance at 2017 [1]. However, its quadratic time and memory complexity on the input length makes its application prohibitive when dealing with large input sequences. This motivated the appearance of several self-attention reformulations in order to lower its complexity and make its development less costly.
We focus on three of these self-attention mechanisms applied to video classification: Cosformer [2], Nyströmformer [3] and Linformer [4]. Concretely, our goal in this project is to suggest which of them is best suited for this task. To evaluate each model performance, we design a personalizable Transformer with interchangeable self attention mechanisms and train it using a simplified dataset derived from EpicKitchens-100 [5]. We carefully describe the Transformer architecture, explaining the purpose of each of its modules, and provide and overall description of how internally works. Preliminary results indicate that Nyströmformer is the best option, being the model which converged faster and achieved the best trade off between computational cost and classification metrics. Linformer obtained similar results and Cosformer apparently failed to perform the classification.
The theoretical formalization of the aforementioned self-attention mechanisms is essential for their results interpretation. Hence, we also provide an in-depth mathematical description of both the original self-attention mechanism presented by Vaswani [1] and the three efficient mechanisms. We realize a complexity analy-
sis of all mechanisms and expose its main properties, linking the theoretical basis with the results.
Description
Treballs Finals de Grau de Matemàtiques, Facultat de Matemàtiques, Universitat de Barcelona, Any: 2023, Director: Sergio Escalera Guerrero, Albert Clapés i Sintes i David Pujol
Citation
Collections
Citation
MARTÍNEZ PÉREZ, Oriol. Efficient transformers applied to video classification. [consulted: 8 of August of 2026]. Available at: https://hdl.handle.net/2445/203127