Amb motiu del tancament d'estiu, la validació de documents es reprendrà a partir del 28 d'agost de 2026. Disculpeu les molèsties.
Con motivo del cierre de verano, la validación de documentos se reanudará a partir del 28 de agosto de 2026. Disculpad las molestias
Due to the summer closure, document validation will resume starting August 28, 2026. We apologize for any inconvenience.

Document type

Bachelor thesis

Publication date

Publication license

cc-by-nc-nd (c) Oriol Martínez Pérez, 2023
Please use this identifier to cite or link to this item: https://hdl.handle.net/2445/203127

Efficient transformers applied to video classification

Journal Title

Journal ISSN

Volume Title

Related resource

Abstract

[en] Transformers, with the self-attention mechanism on its core, have shown great performance on several Machine Learning areas such as NLP or Computer Vision since its appearance at 2017 [1]. However, its quadratic time and memory complexity on the input length makes its application prohibitive when dealing with large input sequences. This motivated the appearance of several self-attention reformulations in order to lower its complexity and make its development less costly. We focus on three of these self-attention mechanisms applied to video classification: Cosformer [2], Nyströmformer [3] and Linformer [4]. Concretely, our goal in this project is to suggest which of them is best suited for this task. To evaluate each model performance, we design a personalizable Transformer with interchangeable self attention mechanisms and train it using a simplified dataset derived from EpicKitchens-100 [5]. We carefully describe the Transformer architecture, explaining the purpose of each of its modules, and provide and overall description of how internally works. Preliminary results indicate that Nyströmformer is the best option, being the model which converged faster and achieved the best trade off between computational cost and classification metrics. Linformer obtained similar results and Cosformer apparently failed to perform the classification. The theoretical formalization of the aforementioned self-attention mechanisms is essential for their results interpretation. Hence, we also provide an in-depth mathematical description of both the original self-attention mechanism presented by Vaswani [1] and the three efficient mechanisms. We realize a complexity analy- sis of all mechanisms and expose its main properties, linking the theoretical basis with the results.

Description

Treballs Finals de Grau de Matemàtiques, Facultat de Matemàtiques, Universitat de Barcelona, Any: 2023, Director: Sergio Escalera Guerrero, Albert Clapés i Sintes i David Pujol

Citation

Citation

MARTÍNEZ PÉREZ, Oriol. Efficient transformers applied to video classification. [consulted: 8 of August of 2026]. Available at: https://hdl.handle.net/2445/203127

Export metadata

JSON - METS

Share record