Amb motiu del tancament d'estiu, la validació de documents es reprendrà a partir del 28 d'agost de 2026. Disculpeu les molèsties.
Con motivo del cierre de verano, la validación de documentos se reanudará a partir del 28 de agosto de 2026. Disculpad las molestias
Due to the summer closure, document validation will resume starting August 28, 2026. We apologize for any inconvenience.

Tipus de document

Article

Versió

Versió acceptada

Data de publicació

Llicència de publicació

cc-by-nc-nd (c) Elsevier B.V., 2020
Si us plau utilitzeu sempre aquest identificador per citar o enllaçar aquest document: https://hdl.handle.net/2445/182997

Integrating lexical and prosodic features for automatic paragraph segmentation

Títol de la revista

Director/Tutor

ISSN de la revista

Títol del volum

Resum

Spoken documents, such as podcasts or lectures, are a growing presence in everyday life. Being able to automatically identify their discourse structure is an important step to understanding what a spoken document is about. Moreover, finer-grained units, such as paragraphs, are highly desirable for presenting and analyzing spoken content. However, little work has been done on discourse based speech segmentation below the level of broad topics. In order to examine how discourse transitions are cued in speech, we investigate automatic paragraph segmentation of TED talks using lexical and prosodic features. Experiments using Support Vector Machines, AdaBoost, and Neural Networks show that models using supra-sentential prosodic features and induced cue words perform better than those based on the type of lexical cohesion measures often used in broad topic segmentation. Moreover, combining a wide range of individually weak lexical and prosodic predictors improves performance, and modelling contextual information using recurrent neural networks outperforms other approaches by a large margin. Our best results come from using late fusion methods that integrate representations generated by separate lexical and prosodic models while allowing interactions between these features streams rather than treating them as independent information sources. Application to ASR outputs shows that adding prosodic features, particularly using late fusion, can significantly ameliorate decreases in performance due to transcription errors.

Citació

Citació

LAI, Catherine, FARRÚS, Mireia and MOORE, Johanna D. Integrating lexical and prosodic features for automatic paragraph segmentation. Speech Communication. 2020. Vol. 121, num. 44-57. ISSN 0167-6393. [consulted: 23 of August of 2026]. Available at: https://hdl.handle.net/2445/182997

Exportar metadades

JSON - METS

Compartir registre