{"id":800,"date":"2022-10-28T12:08:33","date_gmt":"2022-10-28T12:08:33","guid":{"rendered":"http:\/\/iberspeech2022.ugr.es\/?page_id=800"},"modified":"2022-11-11T13:23:32","modified_gmt":"2022-11-11T13:23:32","slug":"technical-program-day-2","status":"publish","type":"page","link":"https:\/\/iberspeech2022.ugr.es\/?page_id=800","title":{"rendered":"Technical Program, Day 2"},"content":{"rendered":"<h3 style=\"background-color: #f3f3f3;\"><span style=\"color: #2e75b7;\"><strong>Tuesday, 15 November<br \/>\n<\/strong><\/span><\/h3>\n<p><strong>Oral 3: Speech and Audio Processing<\/strong><br \/>\n<strong><span style=\"color: #d62013;\">Tuesday, 15 November 2022 (9:00-10:40)<\/span><\/strong><br \/>\n<strong>Chair: Alfonso Ortega<\/strong><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"150\"><strong>O3.1<\/strong><br \/>\n09:00\u00a0 &#8211; 09:20<\/td>\n<td>On the potential of jointly-optimised solutions to spoofing attack detection and automatic speaker verification (<span class=\"collapseomatic \" id=\"id6a6ebb61ae7e8\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61ae7e8\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">The spoofing-aware speaker verification (SASV) challenge was designed to pro<\/span><span dir=\"ltr\" role=\"presentation\">mote the study of jointly-optimised solutions to accomplish the traditionally separately-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">optimised tasks of spoofing detection and speaker verification.<\/span> <span dir=\"ltr\" role=\"presentation\">Jointly-optimised<\/span> <span dir=\"ltr\" role=\"presentation\">systems have the potential to operate in synergy as a better performing solution<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">to the single task of reliable speaker verification. However, none of the 23 submis<\/span><span dir=\"ltr\" role=\"presentation\">sions to SASV 2022 are jointly optimised. We have hence sought to determine why<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">separately-optimised sub-systems perform best or why joint optimisation was not <\/span><span dir=\"ltr\" role=\"presentation\">successful. Experiments reported in this paper show that joint optimisation is suc-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">cessful in improving robustness to spoofing but that it degrades speaker verification <\/span><span dir=\"ltr\" role=\"presentation\">performance. The findings suggest that spoofing detection and speaker verification<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">sub-systems should be optimised jointly in a manner which reflects the differences<\/span> <span dir=\"ltr\" role=\"presentation\">in how information provided by each sub-system is complementary to that provided<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">by the other.<\/span> <span dir=\"ltr\" role=\"presentation\">Progress will also likely depend upon the collection of data from a<\/span> <span dir=\"ltr\" role=\"presentation\">larger number of speakers.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Wanying Ge, Hemlata Tak, Massimiliano Todisco and Nicholas Evans<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"150\"><strong>O3.2<\/strong><br \/>\n09:20\u00a0 &#8211; 09:40<\/td>\n<td>A Study on the Use of wav2vec Representations for Multiclass Audio Segmentation (<span class=\"collapseomatic \" id=\"id6a6ebb61ae868\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61ae868\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">This paper presents a study on the use of new unsupervised representations <\/span><span dir=\"ltr\" role=\"presentation\">through wav2vec models seeking to jointly model speech and music fragments of <\/span><span dir=\"ltr\" role=\"presentation\">audio signals in a multiclass audio segmentation task. Previous studies have already<\/span> <span dir=\"ltr\" role=\"presentation\">described the capabilities of deep neural networks in binary and multiclass audio <\/span><span dir=\"ltr\" role=\"presentation\">segmentation tasks. Particularly, the separation of speech, music and noise signals<\/span> <span dir=\"ltr\" role=\"presentation\">through audio segmentation shows competitive results using a combination of per<\/span><span dir=\"ltr\" role=\"presentation\">ceptual and musical features as input to a neural network. Wav2vec representations<\/span> <span dir=\"ltr\" role=\"presentation\">have been successfully applied to several speech processing applications.<\/span> <span dir=\"ltr\" role=\"presentation\">In this <\/span><span dir=\"ltr\" role=\"presentation\">study, they are considered for the multiclass audio segmentation task presented in<\/span> <span dir=\"ltr\" role=\"presentation\">the Albayz<\/span> <span dir=\"ltr\" role=\"presentation\"> \u0301<\/span><span dir=\"ltr\" role=\"presentation\">\u0131n 2010 evaluation.<\/span> <span dir=\"ltr\" role=\"presentation\">We compare the use of different representations <\/span><span dir=\"ltr\" role=\"presentation\">obtained through unsupervised learning with our previous results in this database<\/span> <span dir=\"ltr\" role=\"presentation\">using a traditional set of features under different conditions. Experimental results <\/span><span dir=\"ltr\" role=\"presentation\">show that wav2vec representations can improve the performance of audio segmen<\/span><span dir=\"ltr\" role=\"presentation\">tation systems for classes containing speech, while showing a degradation in the<\/span> <span dir=\"ltr\" role=\"presentation\">segmentation of isolated music. This trend is consistent among all experiments de<\/span><span dir=\"ltr\" role=\"presentation\">veloped.<\/span> <span dir=\"ltr\" role=\"presentation\">On average, the use of unsupervised representation learning leads to a <\/span><span dir=\"ltr\" role=\"presentation\">relative improvement close to 6.8% on the segmentation task.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Pablo Gimeno, Alfonso Ortega, Antonio Miguel and Eduardo Lleida<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"150\"><strong>O3.3<\/strong><br \/>\n09:40\u00a0 &#8211; 10:00<\/td>\n<td>Respiratory Sound Classification Using an Attention LSTM Model with Mixup Data Augmentation (<span class=\"collapseomatic \" id=\"id6a6ebb61ae8b9\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61ae8b9\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Auscultation is the most common method for the diagnosis of respiratory dis<\/span><span dir=\"ltr\" role=\"presentation\">eases, although it depends largely on the physician\u2019s ability. In order to alleviate this <\/span><span dir=\"ltr\" role=\"presentation\">drawback, in this paper, we present an automatic system capable of distinguishing<\/span> <span dir=\"ltr\" role=\"presentation\">between different types of lung sounds (neutral, wheeze, crackle) in patient\u2019s res<\/span><span dir=\"ltr\" role=\"presentation\">piratory recordings.<\/span> <span dir=\"ltr\" role=\"presentation\">In particular, the proposed system is based on Long Short<\/span> <span dir=\"ltr\" role=\"presentation\">Term-Memory (LSTM) networks fed with log-mel spectrograms, on which several <\/span><span dir=\"ltr\" role=\"presentation\">improvements have been developed. Firstly, the frequency bands that contain more<\/span> <span dir=\"ltr\" role=\"presentation\">useful information have been experimentally determined in order to enhance the <\/span><span dir=\"ltr\" role=\"presentation\">input acoustic features. Secondly, an Attention Mechanism has been incorporated<\/span> <span dir=\"ltr\" role=\"presentation\">into the LSTM model in order to emphasize the more relevant audio frames to the <\/span><span dir=\"ltr\" role=\"presentation\">task under consideration. Finally, a Mixup data augmentation technique has been<\/span> <span dir=\"ltr\" role=\"presentation\">adopted in order to mitigate the problem of data imbalance and improve the sensi<\/span><span dir=\"ltr\" role=\"presentation\">tivity of the system. The proposed methods have been evaluated over the publicly <\/span><span dir=\"ltr\" role=\"presentation\">available ICBHI 2017 dataset, achieving good results in comparison to the baseline.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Noelia Salor-Burdalo and Ascension Gallardo-Antolin<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O3.4<\/strong><br \/>\n10:00\u00a0 &#8211; 10:20<\/td>\n<td>The Vicomtech Spoofing-Aware Biometric System for the SASV Challenge (<span class=\"collapseomatic \" id=\"id6a6ebb61ae908\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61ae908\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">This paper describes our proposed system for the spoofingaware speaker verifica<\/span><span dir=\"ltr\" role=\"presentation\">tion challenge (SASV Challenge 2022). The system follows an integrated approach <\/span><span dir=\"ltr\" role=\"presentation\">that uses speaker verification and antispoofing embeddings extracted from special<\/span><span dir=\"ltr\" role=\"presentation\">ized neural networks. Firstly, a shallow neural network, fed with the test utterance\u2019s <\/span><span dir=\"ltr\" role=\"presentation\">verification and spoofing embeddings, is used to compute a spoof-based score. The<\/span> <span dir=\"ltr\" role=\"presentation\">final scoring decision is then obtained by combining this score with the cosine simi<\/span><span dir=\"ltr\" role=\"presentation\">larity between speaker verification embeddings. The integration network was trained<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">using a one-class loss to discriminate between target and unauthorized trials. Our<\/span> <span dir=\"ltr\" role=\"presentation\">proposed system is evaluated over the ASVspoof19 database and shows competitive<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">performance compared to other integration approaches.<\/span> <span dir=\"ltr\" role=\"presentation\">In addition, we compare<\/span> <span dir=\"ltr\" role=\"presentation\">our approach with further state-of-theart speaker verification and antispoofing sys<\/span><span dir=\"ltr\" role=\"presentation\">tems based on selfsupervised learning, yielding high-performance speech biometric<\/span> <span dir=\"ltr\" role=\"presentation\">systems comparable with the best challenge submissions.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Juan Manuel Mart\u00edn-Do\u00f1as, Iv\u00e1n Gonz\u00e1lez Torre, Aitor \u00c1lvarez and Joaquin Arellano<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O3.5<\/strong><br \/>\n10:20\u00a0 &#8211; 10:40<\/td>\n<td>VoxCeleb-PT &#8211; a dataset for a speech processing course (<span class=\"collapseomatic \" id=\"id6a6ebb61ae955\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61ae955\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">This paper introduces VoxCeleb-PT, a small dataset of voices of Portuguese<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">celebrities that can be used as a language-specific extension of the widely used Vox-<\/span><span dir=\"ltr\" role=\"presentation\">Celeb corpus. Besides introducing the corpus, we also describe three lab assignments<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">where it was used in a one-semester speech processing course: age regression, speaker <\/span><span dir=\"ltr\" role=\"presentation\">verification and speech recognition, hoping to highlight the relevance of this dataset<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">as a pedagogical tool.<\/span> <span dir=\"ltr\" role=\"presentation\">Additionally, this paper confirms the overall limitations of<\/span> <span dir=\"ltr\" role=\"presentation\">current systems when evaluated in different languages and acoustic conditions: we <\/span><span dir=\"ltr\" role=\"presentation\">found an overall degradation of performance on all of the proposed tasks.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>John Mendonca and Isabel Trancoso<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Keynote 2<\/strong><br \/>\n<span style=\"color: #d62013;\"><strong>Tuesday, 15 November 2022 (11:00-12:00)<\/strong><\/span><\/p>\n<table>\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>KN2<\/strong><br \/>\n11:00\u00a0 &#8211; 12:00<\/td>\n<td>Disease biomarkers in speech (<span class=\"collapseomatic \" id=\"id6a6ebb61ae9a4\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61ae9a4\" class=\"collapseomatic_content \">Speech encodes information about a plethora of diseases, which go beyond the so-called speech and language disorders, and include neurodegenerative diseases, such as Parkinson\u2019s, Alzheimer\u2019s, and Huntington\u2019s disease, mood and anxiety-related diseases, such as Depression and Bipolar Disease, and diseases that concern respiratory organs such as the common Cold, or Obstructive Sleep Apnea. This talk addresses the potential of speech as a health biomarker which allows a non-invasive route to early diagnosis and monitoring of a range of conditions related to human physiology and cognition. The talk will also address the many challenges that lie ahead, namely in the context of an ageing population with frequent multimorbidity, and the need to build robust models that provide explanations compatible with clinical reasoning. That would be a major step towards a future where collecting speech samples for health screening may become as common as a blood test nowadays. Speech can indeed encode health information au par with many other characteristics that make it viewed as Personal Identifiable Information. The last part of this talk will briefly discuss the privacy issues that this enormous potential may entail. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Isabel Trancoso<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Posters 2: Special Sessions <\/strong><br \/>\n<span style=\"color: #d62013;\"><strong>Tuesday, 15 November 2022 (12:00-13:30)<\/strong><\/span><br \/>\n<strong>Chair:\u00a0 Inma Hern\u00e1ez<\/strong><\/p>\n<p><strong>Ph.D. Thesis<\/strong><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.1<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Representation and Metric Learning Advances for Deep Neural Network Face and Speaker Biometric Systems (<span class=\"collapseomatic \" id=\"id6a6ebb61ae9ef\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61ae9ef\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">\u00a0<\/span> Nowadays, the use of technological devices and face and speaker biometric recognition systems are becoming increasingly common in people daily lives. This fact has motivated a great deal of research interest in the development of effective and robust systems. However, although face and voice recognition systems are mature technologies, there are still some challenges<br \/>\nwhich need further improvement and continued research when Deep Neural Networks (DNNs) are employed in these systems. In this manuscript, we present an overview of the main findings of Victoria Mingote\u2019s Thesis where different approaches to address these issues are proposed. The advances<br \/>\npresented are focused on two streams of research. First, in the representation learning part, we propose several approaches to obtain robust representations of the signals for text-dependent speaker verification systems. While in the metric learning part, we focus on introducing new loss functions to train DNNs directly to optimize the goal task for text-dependent speaker, language and face verification and also multimodal diarization.<\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em> Victoria Mingote and Antonio Miguel<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.2<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Voice Biometric Systems based on Deep Neural Networks: A Ph.D. Thesis Overview (<span class=\"collapseomatic \" id=\"id6a6ebb61aea49\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aea49\" class=\"collapseomatic_content \">Voice biometric systems based on automatic speaker verification (ASV) are exposed to spoofing attacks which may compromise their security. To increase the robustness against such attacks, anti-spoofing systems have been proposed for the detection of replay, synthesis and voice conversion based attacks. This paper summarizes the work carried out for the first author\u2019s PhD Thesis, which focused on the development of robust biometric systems which are able to detect zero-effort, spoofing and adversarial attacks. First, we propose a gated recurrent convolutional neural network (GRCNN) for detecting both logical and physical access spoofing attacks. Second, we propose a new loss function for training neural networks classifiers based on a probabilistic framework known as kernel density estimation (KDE). Third, we propose a top-performing integration of ASV and anti-spoofing systems with a new loss function which tries to optimize the whole voice biometric system on an expected<br \/>\nrange of operating points. Finally, we propose a generative adversarial network (GAN) for generating adversarial spoofing attacks in order to use them as a defense for building higher robust voice biometric systems. Experimental results show that the proposed techniques outperform many other state-of-the-art systems trained and evaluated in the same conditions with standard public datasets. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em> Alejandro Gomez-Alanis, Jose Andres Gonzalez-Lopez and Antonio Miguel Peinado Herreros<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.3<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Online Multichannel Speech Enhancement combining Statistical Signal Processing and Deep Neural Networks: A Ph.D. Thesis Overview (<span class=\"collapseomatic \" id=\"id6a6ebb61aea97\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aea97\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\"> Speech-related applications on mobile devices require highperformance speech enhancement algorithms to tackle challenging, noisy real-world environments. In addition, current mobile devices often embed several microphones, allowing them to exploit spatial information. The main goal of this Thesis is the development of online multichannel speech enhancement algorithms for speech services in mobile devices. The proposed techniques use multichannel signal processing to increase the noise reduction performance without degrading the quality of\u00a0 the speech signal. Moreover, deep neural networks are applied in specific parts of the algorithm where modeling by classical methods would be, otherwise, unfeasible or very limiting. Our<br \/>\ncontributions focus on different noisy environments where these mobile speech technologies can be applied. These include dualmicrophone smartphones in noisy and reverberant environments<br \/>\nand general multi-microphone devices for speech enhancement and target source separation. Moreover, we study the training of deep learning methods for speech processing using perceptual considerations. Our contributions successfully integrate signal processing and deep learning methods to exploit spectral, spatial, and temporal speech features jointly. As a result, the proposed techniques provide us with a manifold framework for robust speech processing under very challenging acoustic environments, thus allowing us to improve perceptual quality and intelligibility measures.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em> Juan Manuel Mart\u00edn-Do\u00f1as, Antonio M. Peinado and Angel M. Gomez<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Research and Development Projects<\/strong><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.4<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>ReSSInt project: voice restoration using Silent Speech Interfaces\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aeae4\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aeae4\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">ReSSInt is a project funded by the Spanish Ministry of Science <\/span><span dir=\"ltr\" role=\"presentation\">and Innovation aiming at investigating the use of Silent\u00a0 speech interfaces (SSIs) for restoring communication to individuals who have been deprived of the ability to speak. These interfaces capture non-acoustic biosignals generated during the <\/span><span dir=\"ltr\" role=\"presentation\">speech production process and use them to predict the intended message. In the project two different biosignals are being investigated:\u00a0 electromyography (EMG) signals representing electrical activity driving the facial muscles and intracraneal electroencephalography (iEEG) neural signals captured by means of invasive electrodes implanted on the brain. From the whole spectrum of speech disorders which may affect a person\u2019s voice, ReSSInt will address two particular conditions: (i) voice loss after total laryngectomy and (ii) neurodegenerative diseases and other traumatic injuries which may leave an individual paralyzed and, eventually, unable to speak. In this paper we describe the current status of the project as well as the problems and difficulties encountered in its development.<\/span> <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Inma Hernaez, Jose Andres Gonzalez Lopez, Eva Navas, Jose Luis P\u00e9rez C\u00f3rdoba, Ibon Saratxaga, Gonzalo Olivares, Jon Sanchez de la Fuente, Alberto Gald\u00f3n, Victor Garcia, Jes\u00fas del Castillo, Inge Salomons and Eder del Blanco Sierra<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.5<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>ELE Project: an overview of the desk research (<span class=\"collapseomatic \" id=\"id6a6ebb61aeb2f\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aeb2f\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">This paper provides an overview of the European Language Equality (ELE) project. The main objective of ELE is to prepare<br \/>\nthe European Language Equality program in the form of<br \/>\na strategic research and innovation agenda that may be utilized<br \/>\nas a road map for achieving full digital language equality in<br \/>\nEurope by 2030. The desk research phase of ELE concentrated<br \/>\non the systematic collection and analysis of the existing international,<br \/>\nnational, and regional strategic research agendas, studies,<br \/>\nreports, and initiatives related to language technology and LTrelated<br \/>\nartificial intelligence. A brief survey of the findings is<br \/>\npresented here, with a special focus on the Spanish ecosystem.<\/span> <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Itziar Aldabe, Aritz Farwell, Eva Navas, Inma Hernaez, German Rigau<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.6<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Snorble: An Interactive Children Companion\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aeb7a\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aeb7a\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">This paper presents an interactive companion called Snorble, created to engage with children and promote the development of healthy habits under the Snorble project.<br \/>\nSnorble is a smart companion capable of having a conversation with children, playing games, and helping them to go to sleep, all made possible thanks to speech recognition.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Mike Rizkalla, Thomas Chan, Emilio Granell, Chara Tsoukala, Aitor Carricondo, Carlos Bailon, Mar\u00eda Teresa Gonz\u00e1lez and Vicent Alabau<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.7<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Fusion of Classical Digital Signal Processing and Deep Learning methods (FTCAPPS)\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aebc9\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aebc9\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">The use of deep learning approaches in Signal Processing is<br \/>\nfinally showing a trend towards a rational use. After an effervescent<br \/>\nperiod where research activity seemed to focus on<br \/>\nseeking old problems to apply solutions entirely based on neural<br \/>\nnetworks, we have reached a more mature stage where integrative<br \/>\napproaches are on the rise. These approaches gather<br \/>\nthe best from each paradigm: on the one hand, the knowledge<br \/>\nand elegance of classical signal processing and, on the other,<br \/>\nthe great ability to model and learn from data which is inherent<br \/>\nto deep learning methods. In this project we aim towards a new<br \/>\nsignal processing paradigm where classical and deep learning<br \/>\ntechniques not only collaborate, but fuse themselves. In particular,<br \/>\nwe focus on two objectives: 1) the development of deep<br \/>\nlearning architectures based on or inspired by signal processing<br \/>\nschemes, and 2) the improvement of current deep learning training<br \/>\nmethods by means of classical techniques and algorithms,<br \/>\nparticularly, by exploiting the knowledge legacy they treasure.<br \/>\nThese innovations will be applied to two socially and scientifically<br \/>\nrelevant topics in which our research group has been working<br \/>\nfor years. The first one is the enhancement of speech signal<br \/>\nacquired under acoustic adverse conditions (e.g., noise, reverberation,<br \/>\nother speakers, &#8230;). The second one is the development<br \/>\nof anti-fraud measures for biometric voice authentication,<br \/>\nin which banking corporations and other large companies are<br \/>\nstrongly interested.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Angel M. G\u00f3mez, Victoria E. Sanchez, Antonio M. Peinado, Juan M. Mart\u00edn-Do\u00f1as, Alejandro G\u00f3mez-Alanis, Amelia Villegas-Morcillo, Eros Rosello, Manuel Chica, Celia Garc\u00eda and Ivan L\u00f3pez-Espejo<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.8<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Spanish Lipreading in Realistic Scenarios: the LLEER project\u00a0\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aec16\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aec16\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Automatic speech recognition has been usually performed by<br \/>\nusing only the audio data, but speech communication is affected<br \/>\nas well by other non-audio sources, mainly visual cues. Visual<br \/>\ninformation includes body expression, face expression, and lip<br \/>\nmovements, among other. Lip reading, also known as Visual<br \/>\nSpeech Recognition, aims at decoding speech by only using<br \/>\nthe image of the lip movements. Current approaches for automatic<br \/>\nlip reading follow the same lines than for speech processing:<br \/>\nuse of massive data for training deep learning models<br \/>\nthat allow to perform speech recognition. However, most of the<br \/>\ndatasets and models are devoted to languages such as English<br \/>\nor Chinese, while other languages, particularly Spanish, are underrepresented.<br \/>\nThe LLEER (Lectura de Labios en Espa\u02dcnol en<br \/>\nEscenarios Realistas) project aims at the acquisition of largescale<br \/>\nvisual corpora for Spanish lip reading, the development<br \/>\nof visual processing techniques that allow to extract important<br \/>\ninformation for the task, the implementation of models for automatic<br \/>\nlip reading, and the integration with speech recognition<br \/>\nmodels for audiovisual speech recognition.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Carlos David Martinez Hinarejos, David Gimeno-Gomez, Francisco Casacuberta, Emilio Granell, Roberto Paredes, Mois\u00e9s Pastor and Enrique Vidal<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.9<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Clinical Applications of Neuroscience: Locating Language Areas in Epileptic Patients and Restoring Speech in Paralyzed People\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aec63\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aec63\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">The goal of this project is to study the neurological bases of<br \/>\nlanguage using intracranial electroencephalography (iEEG) signals<br \/>\nrecorded from drug-resistant epilepsy patients. In particular,<br \/>\nwe aim to address two current clinical challenges. Firstly,<br \/>\nwe intend to individually identify the brain regions involved<br \/>\nin the production and understanding of language, in order to<br \/>\npreserve these regions during brain surgery for epilepsy treatment.<br \/>\nSecondly, this project also aims to develop novel pattern<br \/>\nrecognition algorithms that can decode speech from iEEG signals<br \/>\nobtained from participants performing language production<br \/>\ntasks. The ultimate goal is to evaluate the feasibility of a neuroprosthetic<br \/>\ndevice that could restore oral communication in persons<br \/>\nthat cannot speak following a neurodegenerative disease<br \/>\nor brain damage. For both goals, a series of experimental tasks<br \/>\nwill be developed in order to thoroughly evaluate language production<br \/>\nand comprehension. Furthermore, data derived from<br \/>\nthese tasks will be analyzed using state-of-the-art multivariate<br \/>\nstatistical methods and machine learning techniques (e.g., deep<br \/>\nlearning). In addition to having a social impact, the results of<br \/>\nthis project will also help in advancing the knowledge about the<br \/>\nneural substrates that underpin language production and comprehension.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Jose Andres Gonzalez Lopez, Alberto Gald\u00f3n, Gonzalo Olivares, Sneha Raman, David Mu\u00f1oz, Daniela Paolieri, Pedro Macizo, Jos\u00e9 L. P\u00e9rez-C\u00f3rdoba, Antonio M. Peinado, Angel Gomez, Victoria E. Sanchez and Ana B. Chica<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.10<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>ORKESTA Comprehensive Solution for the Orchestration of Services and Soci-Sanitary Care at Home\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aecae\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aecae\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">In this paper we present the main goals of the ORKESTA<br \/>\nproject. This is an industrial project carried out by a consortium<br \/>\nof companies aimed at providing products and services<br \/>\ncontributing to improve the wellbeing of the old adults and enlarge<br \/>\nthe years of independent life. To this end the consortium<br \/>\ncollaborates with the Vicomtech Tecnological Center and the<br \/>\nSpeech Interactive research Group at the UPV\/EHU. Both provide<br \/>\nspeech and language Technologies to the project.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Juan Alos, Julien Boulli\u00e9, M. In\u00e9s Torres, Eneko Ruiz, Andoni Beristain, Jacobo L\u00f3pez Fern\u00e1ndez, I\u00f1aki Teller\u00eda, Janeth Carolina Carre\u00f1o, Iker Garay, Arkaitz Carbajo, Amaia Santamar\u00eda, Urtzi Zubiate, Jon Ander Arzallus, Francisco Mart\u00ednez and Adriana Mart\u00ednez<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.11<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>The CITA GO-ON trial: A person-centered, digital, intergenerational, and cost- effective dementia prevention multi-modal intervention model to guide strategic policies facing the demographic challenges of progressive aging\u00a0\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aecf9\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aecf9\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">This paper presents a general overview of the CITA GOON<br \/>\nstudy, a controlled and randomized trial aimed to<br \/>\ndemonstrate the efficacy and cost-effectiveness of a 2 years<br \/>\nmulti-modal intervention to control risk factors and change<br \/>\nlifestyles in cognitively frail people at increased risk of<br \/>\ndementia. In this framework, the applicability of a virtual<br \/>\nagent to increase adherence and effectiveness (the \u201cGo-ON<br \/>\ndigital coach\u201d) will be explored.<br \/>\nThe multidisciplinary nature of the study brings together 7<br \/>\npartners including non-profit organizations, universities,<br \/>\ntechnological centers and companies.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Mikel Tainta, Javier Mikel Olaso, M. In\u00e9s Torres, Mirian Ecay-Torres, Nekane Balluerka, Naia Ros, Mikel Izquierdo, Mikel Sa\u00e9z de Asteasu, Usune Etxebarria, Luc\u00eda Gayoso, Maider Mateo, Oliver Ibarrondo, Elena Alberdi, Est\u00edbaliz Capetillo-Z\u00e1rate, Jesus Angel Bravo and Pablo Mart\u00ednez-Lage<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.12<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>The BioVoz Project: Secure Speech Biometrics by Deep Processing Techniques\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aed45\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aed45\" class=\"collapseomatic_content \">Currently, voice biometrics systems are attracting a growing<br \/>\ninterest driven by the need for new authentication modalities.<br \/>\nThe BioVoz project focuses on the reliability of these systems,<br \/>\nthreatened by various types of attacks, from a simple playback<br \/>\nof prerecorded speech to more sophisticated variants such as impersonation<br \/>\nbased on voice conversion or synthesis. One problem<br \/>\nin detecting spoofed speech is the lack of suitable models<br \/>\nbased on classical signal processing techniques. Therefore, the<br \/>\ncurrent trend is based on the use of deep neural networks, either<br \/>\nfor direct attack detection, or for obtaining deep feature vectors<br \/>\nto represent the audio signals. However, these solutions raise<br \/>\nmany questions that are still unanswered and are the subject<br \/>\nof the research proposed here. These include what spectral or<br \/>\ntemporal information should be used to feed the network, how<br \/>\nto compensate for the effect of acoustic noise, what network architecture<br \/>\nis appropriate, or what methodology should be used<br \/>\nfor training in order to provide the network with discriminative<br \/>\ngeneralization capabilities. The present project focuses on<br \/>\nthe search for solutions to the aforementioned problems without<br \/>\nforgetting a fundamental issue, little studied so far, such as the<br \/>\nintegration of fraud detection in the whole biometrics system.<\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Antonio M. Peinado, Alejandro Gomez-Alanis, Jose Andres Gonzalez-Lopez, Angel M. Gomez, Eros Rosello, Manuel Chica-Villar, Jose C. Sanchez-Valera, Jose L. Perez-Cordoba and Victoria Sanchez<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.13<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Automatic evaluation of the pronunciation of people with Down syndrome in an educational video game (EvaProDown)\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aed92\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aed92\" class=\"collapseomatic_content \">The deficiencies in oral communication of people with Down<br \/>\nsyndrome (DS) represent an important barrier towards their social<br \/>\nintegration. Interventions based on performing exercises<br \/>\nof speech and language therapy have proven to be effective in<br \/>\nimproving their communication skills. Our research group has<br \/>\nbeen involved in the development of a serious video game for<br \/>\nthe practice of oral communication of people with Down syndrome.<br \/>\nThe video game has proven its usefulness by being able<br \/>\nto motivate users to carry out practical exercises designed to improve<br \/>\ntheir communication skills related to prosody, an important<br \/>\naspect of spoken communication. The video game has also<br \/>\nfacilitated the compilation of a speech corpus called Prautocal<br \/>\nwith a large number of utterances of people with DS. The objective<br \/>\nof this project is to extend the functionality of the video<br \/>\ngame to include exercises focused on pronunciation and on improving<br \/>\narticulation and speech intelligibility. To do this, an<br \/>\nautomatic pronunciation assessment module will be developed<br \/>\nand incorporated into the existing video game in order to complement<br \/>\nits functionality. In this way, using the video game,<br \/>\nusers will be able to perform exercises autonomously to work<br \/>\non aspects of speech related to both pronunciation and prosody.<\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>\u00a0C\u00e9sar Gonz\u00e1lez-Ferreras, Valent\u00edn Carde\u00f1oso-Payo, David Escudero, Carlos Vivaracho-Pascual, Lourdes Aguilar, Valle Flores-Lucas and Mario Corrales Astorgano<br \/>\n<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><strong>\u00a0<\/strong><\/p>\n<p><strong>Demos<\/strong><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.14<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>SONOC Platform for Audio and Speech Analytics in Call Centers\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aedde\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aedde\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">This paper presents a platform for processing audio data of call centers to obtain statistical information on the telephone call. The system computes several metrics of the audio and speech to define a representation of the call flow, the audio quality, and the paralinguistic performance of the call. This way, it can model the behavior and feelings of the agent and customer involved in the conversation. This solution applies to many industries such as call centers, social communities, metaverse, customer identifications, online and offline meetings, etc. In summary, the platform leverages an already trained artificial intelligence business network to get non-verbal communication information from audio. This information translates into valuable business insights for further decision-making.<\/span> <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>\u00a0Dayana Ribas, Antonio Miguel, Luis Guillen, Jose Javier Castejon, Juan Antonio Navarro, Alfonso Ortega and Luis Benavente<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><strong>\u00a0<\/strong><\/p>\n<p><strong>Entrepreneurship<\/strong><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.15<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>\u00a0 ELSA Speak\u00a0 (<span class=\"collapseomatic \" id=\"id6a6ebb61aee34\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aee34\" class=\"collapseomatic_content \">In 2015 Xavier Anguera and Vu Van co-founded ELSA (English Language Speech Assistant) an app (and AI technology) to help learners of English to improve their pronunciation skills. Fast forward to 2022, the company has grown to more than 100 employees and with offices in US, Portugal, India and Vietnam. Our application (ELSA Speak) has been downloaded over 20M times and we are serving users from over 100 countries, who speak to the app and get feedback in real time.<br \/>\nMoving from research to a startup environment and to a product requires a mindset change in some areas (e.g. you need to always be razor-focused on what you spend time on) and is very similar in others (e.g. long hours of work, you need to be very resilient when things look bad). In The session we will share some of the learnings we acquired on our particular journey.<\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em> Xavier Anguera<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P2.16<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Monoceros Labs: From Voice Applications To Voice Synthesis In The Spanish Market (<span class=\"collapseomatic \" id=\"id6a6ebb61aee81\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aee81\" class=\"collapseomatic_content \">Creating a company in speech technologies in Spain and applying learnings from the research on dialogue systems at the University of Granada from 2009 to 2013 was not intended at first. In 2018, voice assistants landed in Spain, allowing us to extend them by creating voice and multimodal applications. We worked with users (from kids to older adults) and companies from different sectors (insurances, media) to expand their content and services to Amazon Alexa. Our motivation is breaking down barriers between technology and people using advances in the speech technology area. Voice is natural, efficient and accessible in many contexts. Our focus on people led us to learn the nuances of their needs, test in real scenarios, and launch to the market as soon as possible. After a few years, currently available synthetic voices in Spanish were not creating the best experiences we aimed for users in some use cases. We started working on Spanish neural TTS to close the gap between SOTA and the market. We are currently building our TTS platform; meanwhile working with companies and content creators to validate and learn from the possible uses of TTS, impact and benefits, which goes from content accessibility to scalability. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em> Nieves Abalos and Carlos Mu\u00f1oz-Romero<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Oral 4: Affective Computing and Applications<\/strong><br \/>\n<span style=\"color: #d62013;\"><strong>Tuesday, 15 November 2022 (15:00-17:00)<\/strong><\/span><br \/>\n<strong>Chair: Carmen Pel\u00e1ez Moreno<\/strong><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O4.1<\/strong><br \/>\n15:00\u00a0 &#8211; 15:20<\/td>\n<td>Cross-Corpus Speech Emotion Recognition with HuBERT Self-Supervised Representation (<span class=\"collapseomatic \" id=\"id6a6ebb61aeece\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aeece\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Speech Emotion Recognition (SER) is a task related to many applications in <\/span><span dir=\"ltr\" role=\"presentation\">the framework of human-machine interaction. However, the lack of suitable speech <\/span><span dir=\"ltr\" role=\"presentation\">emotional datasets compromises the performance of the SER systems.<\/span> <span dir=\"ltr\" role=\"presentation\">A lot of<\/span> <span dir=\"ltr\" role=\"presentation\">labeled data are required to accomplish successful training, especially for current <\/span><span dir=\"ltr\" role=\"presentation\">Deep Neural Network (DNN)-based solutions.<\/span> <span dir=\"ltr\" role=\"presentation\">Previous works have explored dif<\/span><span dir=\"ltr\" role=\"presentation\">ferent <\/span><span dir=\"ltr\" role=\"presentation\">strategies for extending the training set using some emotion speech corpora<\/span> <span dir=\"ltr\" role=\"presentation\">available.<\/span> <span dir=\"ltr\" role=\"presentation\">In this paper, we evaluate the impact on the performance of crosscor<\/span><span dir=\"ltr\" role=\"presentation\">pus as a data augmentation strategy for spectral representations and the recent <\/span><span dir=\"ltr\" role=\"presentation\">Self-Supervised (SS) representation of Hu- BERT in an SER system.<\/span> <span dir=\"ltr\" role=\"presentation\">Experimen<\/span><span dir=\"ltr\" role=\"presentation\">tal results show improvements in the accuracy of SER in the IEMOCAP dataset<\/span> <span dir=\"ltr\" role=\"presentation\">when extending the training set with two other datasets, EmoDB in German and<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">RAVDESS in English.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Miguel Pastor, Dayana Ribas, Alfonso Ortega, Antonio Miguel and Eduardo Lleida<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O4.2<\/strong><br \/>\n15:20\u00a0 &#8211; 15:40<\/td>\n<td>Analysis of Trustworthiness Recognition models from an aural and emotional perspective (<span class=\"collapseomatic \" id=\"id6a6ebb61aef1b\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aef1b\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Trustworthiness and deception recognition attracts the research community at<\/span><span dir=\"ltr\" role=\"presentation\">tention due to their relevant role in social negotiations and other relevant areas.<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Despite the increasing interest in the field, there are still many questions about <\/span><span dir=\"ltr\" role=\"presentation\">how to perform automatic deception detection or which features explain better how<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">people perceive trustworthiness. Previous studies have demonstrated that emotions <\/span><span dir=\"ltr\" role=\"presentation\">and sentiments correlate with deception. However, not many articles employed deep-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">learning models pre-trained on emotion recognition tasks to predict trustworthiness. <\/span><span dir=\"ltr\" role=\"presentation\">For this reason, this paper will compare traditional statistical functional feature sets<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">proposed for performing emotion recognition, such as eGeMAPS, with features ex<\/span><span dir=\"ltr\" role=\"presentation\">tracted from deep-learning models, like AlexNet, CNN-14 or xlsr-Wav2Vec2.0 pre<\/span><span dir=\"ltr\" role=\"presentation\">trained on emotion recognition tasks. After obtaining each set of features, we will<\/span> <span dir=\"ltr\" role=\"presentation\">train a Support Vector Machine (SVM) model on deception detection.<\/span> <span dir=\"ltr\" role=\"presentation\">These ex<\/span><span dir=\"ltr\" role=\"presentation\">periments provide a baseline to understand how methodologies exploited in emotion<\/span> <span dir=\"ltr\" role=\"presentation\">recognition tasks could be applied to speech trustworthiness recognition. Utilizing<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the eGeMAPs feature set on deception detection achieved an accuracy of 65.98% <\/span><span dir=\"ltr\" role=\"presentation\">at turn level, and employing transfer-learning on the embeddings extracted from<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">a pre-trained xlsr-Wav2Vec2.0 let improve this rate until a 68.11%, surpassing the <\/span><span dir=\"ltr\" role=\"presentation\">baseline on audio modality from previous works by an 8.5%.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Cristina Luna Jim\u00e9nez, Ricardo Kleinlein, Syaheerah Lebai Lutfi, Juan M. Montero and Fernando Fern\u00e1ndez-Mart\u00ednez<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O4.3<\/strong><br \/>\n15:40\u00a0 &#8211; 16:00<\/td>\n<td>Speech and Text Processing for Major Depressive Disorder Detection (<span class=\"collapseomatic \" id=\"id6a6ebb61aef69\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aef69\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Major Depressive Disorder (MDD) is a common mental health issue these days.<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Its early diagnostic is vital to avoid bigger consequences and provide an appropriate <\/span><span dir=\"ltr\" role=\"presentation\">treatment. Speech and utterance\u2019s transcription of patients\u2019 interviews contain use<\/span><span dir=\"ltr\" role=\"presentation\">ful information sources for the automatic screening of MDD. In this sense, speech <\/span><span dir=\"ltr\" role=\"presentation\">and text-based systems are proposed in this paper, using the DAIC-WOZ dataset<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">as experimental framework. The speech-based one is a Sequence-to-Sequence (S2S) <\/span><span dir=\"ltr\" role=\"presentation\">model with a local attention mechanism.<\/span> <span dir=\"ltr\" role=\"presentation\">The text-based one is based on GloVe<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">features and a Convolutional Neural Network as classifier. A description of some of<\/span> <span dir=\"ltr\" role=\"presentation\">the more relevant results achieved by other research publications on DAIC-WOZ are<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">described as well. The goal is to provide a better understanding of the context of<\/span> <span dir=\"ltr\" role=\"presentation\">our systems results. In general, the S2S architecture provides mostly better results<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">than previous speechbased systems.<\/span> <span dir=\"ltr\" role=\"presentation\">The GloVe-CNN system shows even a better <\/span><span dir=\"ltr\" role=\"presentation\">performance, leading to the idea that text is a more suitable information source for<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the detection of MDD when it is manually developed.<\/span> <span dir=\"ltr\" role=\"presentation\">However, to automatically<\/span> <span dir=\"ltr\" role=\"presentation\">obtain high quality transcriptions is not a straightforward task, which makes nec<\/span><span dir=\"ltr\" role=\"presentation\">essary the development of effective speech-based systems as the presented in this<\/span> <span dir=\"ltr\" role=\"presentation\">research work.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Edward L. Campbell, Laura Doc\u00edo Fern\u00e1ndez, Nicholas Cummins and Carmen Garc\u00eda Mateo<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O4.4<\/strong><br \/>\n16:00\u00a0 &#8211; 16:20<\/td>\n<td>Bridging the Semantic Gap with Affective Acoustic Scene Analysis: an Information Retrieval-based Approach (<span class=\"collapseomatic \" id=\"id6a6ebb61aefe6\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61aefe6\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Human emotions induce physiological and physical changes in the body and <\/span><span dir=\"ltr\" role=\"presentation\">can ultimately influence our actions.<\/span> <span dir=\"ltr\" role=\"presentation\">Their study belongs to the field of Affective<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Computing, to improve human-computer interaction tasks.<\/span> <span dir=\"ltr\" role=\"presentation\">Defining an \u2019affective <\/span><span dir=\"ltr\" role=\"presentation\">acoustic scene\u2019 as an acoustic environment that can induce specific emotions, in this<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">work we aim to characterize acoustic scenes that elicit affective states regarding the <\/span><span dir=\"ltr\" role=\"presentation\">acoustic events occurring and the available acoustic information. This is achieved by<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">generating emotion embeddings to define the \u2019affective acoustic fingerprint\u2019 of such<\/span> <span dir=\"ltr\" role=\"presentation\">affective acoustic scenes. We use YAMNet, an acoustic events\u2019 classifier trained in<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Audioset to classify acoustic events in the WEMAC Audiovisual stimuli dataset.<\/span> <span dir=\"ltr\" role=\"presentation\">Each video in this dataset is labelled by crowd-sourcing with the categorical emo<\/span><span dir=\"ltr\" role=\"presentation\">tion it induces.<\/span> <span dir=\"ltr\" role=\"presentation\">Thus we determine the relevance of the detected acoustic events<\/span> <span dir=\"ltr\" role=\"presentation\">that induce each emotion by performing an affective acoustic mapping, creating<\/span> <span dir=\"ltr\" role=\"presentation\">interpretable acoustic fingerprints of such emotions, by means of the well-known<\/span> <span dir=\"ltr\" role=\"presentation\">information-retrieval-based TF-IDF algorithm. This paper intends to shed light on <\/span><span dir=\"ltr\" role=\"presentation\">the path to the definition of emotional acoustic embeddings.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Clara Luis-Mingueza, Esther Rituerto-Gonz\u00e1lez and Carmen Pel\u00e1ez-Moreno<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O4.5<\/strong><br \/>\n16:20\u00a0 &#8211; 16:40<\/td>\n<td>Detecting Gender-based Violence aftereffects from Emotional Speech Paralinguistic Features (<span class=\"collapseomatic \" id=\"id6a6ebb61af03e\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61af03e\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Speech is known to provide information regarding the person speaking, such as <\/span><span dir=\"ltr\" role=\"presentation\">their gender, identity, emotions, and even disorders or trauma. In this paper we aim <\/span><span dir=\"ltr\" role=\"presentation\">to answer the following question, can women who have suffered from gender-based<\/span> <span dir=\"ltr\" role=\"presentation\">violence (GBV) be distinguished from those who have not, just by using speech <\/span><span dir=\"ltr\" role=\"presentation\">paralinguistic cues?<\/span> <span dir=\"ltr\" role=\"presentation\">In this work, we intend to demonstrate whether there exist<\/span> <span dir=\"ltr\" role=\"presentation\">measurable differences between the emotional expression in the voice of GBV vic<\/span><span dir=\"ltr\" role=\"presentation\">tims (GBVV) and non-victims (Non-GBVV). The present study was carried out<\/span> <span dir=\"ltr\" role=\"presentation\">in the framework of the project EMPATIA-CM, whose aim is to understand the <\/span><span dir=\"ltr\" role=\"presentation\">reaction of GBVV to dangerous situations and develop automatic mechanisms to<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">protect them.<\/span> <span dir=\"ltr\" role=\"presentation\">For this purpose, we use data collected and partly published from<\/span> <span dir=\"ltr\" role=\"presentation\">the WEMAC Database, a multimodal database containing physiological and speech<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">data from women who have and have not suffered from GBV while visualizing dif<\/span><span dir=\"ltr\" role=\"presentation\">ferent emotion-eliciting video clips. With the performed analysis, it is proven that<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">such differences exist indeed and, therefore, that suffering from GBV alters the way <\/span><span dir=\"ltr\" role=\"presentation\">women react to the same emotion eliciting stimulus in terms of physical variables,<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">specifically certain voice features<\/span>.<\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Emma Reyner Fuentes, Esther Rituerto Gonz\u00e1lez, Clara Luis Mingueza, Carmen Pel\u00e1ez Moreno and Celia L\u00f3pez Ongil<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O4.6<\/strong><br \/>\n16:40\u00a0 &#8211; 17:00<\/td>\n<td>Extraction of structural and semantic features for the identification of Psychosis in European Portuguese (<span class=\"collapseomatic \" id=\"id6a6ebb61af094\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61af094\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Psychosis is a brain condition that affects the subject and the way it perceives <\/span><span dir=\"ltr\" role=\"presentation\">the world around, impairing its cognitive and speech capabilities, and creating a <\/span><span dir=\"ltr\" role=\"presentation\">disconnection from reality in which the subject is inserted. Psychosis lacks formal<\/span> <span dir=\"ltr\" role=\"presentation\">and precise diagnostic tools, relying on self-reports from patients, their families, <\/span><span dir=\"ltr\" role=\"presentation\">and specialized clinicians. Previous studies have focused on the identification and<\/span> <span dir=\"ltr\" role=\"presentation\">prediction of psychosis through surface-level analysis of diagnosed patients target<\/span><span dir=\"ltr\" role=\"presentation\">ing audio, time, and paucity features to predict or identify psychosis. More recent<\/span> <span dir=\"ltr\" role=\"presentation\">studies have started focusing on high-level and complex language analysis such as se<\/span><span dir=\"ltr\" role=\"presentation\">mantics, structure, and pragmatics. Only a reduced number of studies have targeted<\/span> <span dir=\"ltr\" role=\"presentation\">the Portuguese language. Currently, no study has targeted structural or semantic <\/span><span dir=\"ltr\" role=\"presentation\">features in European Portuguese, thus this is our objective.<\/span> <span dir=\"ltr\" role=\"presentation\">The results obtained<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">through our work suggest that the use of structural and semantic features, particu<\/span><span dir=\"ltr\" role=\"presentation\">larly for European Portuguese, holds some power in classifying subjects as diagnosed<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">with psychosis or not. However, further research is required to identify possible im<\/span><span dir=\"ltr\" role=\"presentation\">provements to the techniques employed and to concretely identify which particular<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">features hold the most power during the classification tasks.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Rodrigo Sousa, Helena Sofia Pinto, Alberto Abad, Daniel Neto and Joaquim Gago<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Albayzin Evaluations<\/strong><br \/>\n<span style=\"color: #d62013;\"><strong>Tuesday, 15 November 2022 (17:20 \u2013 19:20)<\/strong><\/span><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"150\"><strong>A.1<\/strong><br \/>\n17:20\u00a0 &#8211; 19:20<\/td>\n<td>The Vicomtech-UPM Speech Transcription Systems for the Albayz\u00edn-RTVE 2022 Speech to Text Transcription Challenge (<span class=\"collapseomatic \" id=\"id6a6ebb61af0e4\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61af0e4\" class=\"collapseomatic_content \">This paper describes the Vicomtech-UPM submission to the<br \/>\nAlbayz\u00b4\u0131n-RTVE 2022 Speech to Text Transcription Challenge,<br \/>\nwhich calls for automatic speech transcription systems to be<br \/>\nevaluated in realistic TV shows. A total of 4 systems were built<br \/>\nand presented to the evaluation challenge, considering the primary<br \/>\nsystem alongside three contrastive systems. Each system<br \/>\nwas built on top of one different architecture, with the aim of<br \/>\ntesting several state-of-the-art modelling approaches focused on<br \/>\ndifferent learning techniques and typologies of neural networks.<br \/>\nThe primary system used the self-supervised Wav2vec2.0<br \/>\nmodel as the pre-trained model of the transcription engine. This<br \/>\nmodel was fine-tuned with in-domain labelled data and the initial<br \/>\nhypothesis re-scored with a pruned 4-gram based language<br \/>\nmodel. The first contrastive system corresponds to a pruned<br \/>\nRNN-Transducer model, composed of a Conformer encoder<br \/>\nand a stateless prediction network using BPE word-pieces as<br \/>\noutput symbols. As the second contrastive system, we built<br \/>\na Multistream-CNN acoustic model based system with a nonpruned<br \/>\n3-gram model for decoding, and a RNN based language<br \/>\nmodel for rescoring the initial lattices. Finally, results obtained<br \/>\nwith the publicly available Large model of the recently published<br \/>\nWhisper engine were also presented within the third contrastive<br \/>\nsystem, with the aim of serving as a reference benchmark<br \/>\nfor other engines. Along with the description of the systems,<br \/>\nthe results obtained on the Albayzin-RTVE 2020 and<br \/>\n2022 test sets by each engine are presented as well. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Haritz Arzelus, Iv\u00e1n G. Torres, Juan Manuel Mart\u00edn-Do\u00f1as, Ander Gonz\u00e1lez-Docasal and Aitor Alvarez<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"150\"><strong>A.2<\/strong><br \/>\n17:20\u00a0 &#8211; 19:20<\/td>\n<td>TID Spanish ASR system for the Albayzin 2022 Speech-to-Text Transcription Challenge (<span class=\"collapseomatic \" id=\"id6a6ebb61af136\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61af136\" class=\"collapseomatic_content \">This paper describes Telef\u00b4onica I+D\u2019s participation in the<br \/>\nIberSPEECH-RTVE 2022 Speech-to-Text Transcription Challenge.<br \/>\nWe built an acoustic end-to-end Automatic Speech<br \/>\nRecognition (ASR) based on the large XLS-R architecture. We<br \/>\nfirst trained it with already aligned data from CommonVoice.<br \/>\nAfter we adapted it to the TV broadcasting domain with a<br \/>\nself-supervised method. For that purpose, we used an iterative<br \/>\npseudo-forced alignment algorithm fed with frame-wise character<br \/>\nposteriors produced by our ASR. This allowed us to recover<br \/>\nup to 166 hours from RTVE2018 and RTVE2022 databases. We<br \/>\nadditionally explored using a transformer-based seq2seq translator<br \/>\nsystem as a Language Model (LM) to correct the transcripts<br \/>\nof the acoustic ASR. Our best system achieved 24.27%<br \/>\nWER in the test split of RTVE2020. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Fernando L\u00f3pez and Jordi Luque<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"150\"><strong>A.3<\/strong><br \/>\n17:20\u00a0 &#8211; 19:20<\/td>\n<td>BCN2BRNO: ASR System Fusion for Albayzin 2022 Speech to Text Challenge (<span class=\"collapseomatic \" id=\"id6a6ebb61af184\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61af184\" class=\"collapseomatic_content \">This paper describes the joint effort of BUT and Telef\u00f3nica Research<br \/>\non the development of Automatic Speech Recognition<br \/>\nsystems for the Albayzin 2022 Challenge. We train and evaluate<br \/>\nboth hybrid systems and those based on end-to-end models.<br \/>\nWe also investigate the use of self-supervised learning speech<br \/>\nrepresentations from pre-trained models and their impact on<br \/>\nASR performance (as opposed to training models directly from<br \/>\nscratch). Additionally, we also apply the Whisper model in a<br \/>\nzero-shot fashion, postprocessing its output to fit the required<br \/>\ntranscription format. On top of tuning the model architectures<br \/>\nand overall training schemes, we improve the robustness of our<br \/>\nmodels by augmenting the training data with noises extracted<br \/>\nfrom the target domain. Moreover, we apply rescoring with<br \/>\nan external LM on top of N-best hypotheses to adjust each<br \/>\nsentence score and pick the single best hypothesis. All these<br \/>\nefforts lead to a significant WER reduction. Our single best<br \/>\nsystem and the fusion of selected systems achieved 16.3% and<br \/>\n13.7% WER respectively on RTVE2020 test partition, i.e. the<br \/>\nofficial evaluation partition from the previous Albayzin challenge. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Martin Kocour, Jahnavi Umesh, Martin Karafiat, J\u00e1n \u0160vec, Fernando L\u00f3pez, Jordi Luque, Karel Bene\u0161, Mireia Diez, Igor Szoke, Karel Vesel\u00fd, Luk\u00e1\u0161 Burget and Jan \u010cernock\u00fd<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"150\"><del><strong>A.4<\/strong><\/del><br \/>\n17:20\u00a0 &#8211; 19:20<\/td>\n<td><del>BUT System for Albayzin 2022 Text and Speech Alignment Challenge<\/del> \u00a0 \u00a0 <strong>Withdrawn<\/strong><\/td>\n<\/tr>\n<tr>\n<td><del><em>Martin Kocour, Jahnavi Umesh, Martin Karafiat, Igor Szoke, Karel Bene\u0161 and Jan \u010cernock\u00fd<\/em><\/del><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"150\"><strong>A.5<\/strong><br \/>\n17:20\u00a0 &#8211; 19:20<\/td>\n<td>Intelligent Voice Speaker Recognition and Diarization System for IberSpeech 2022 Albayzin Evaluations Speaker Diarization and Identity Assignment Challenge (<span class=\"collapseomatic \" id=\"id6a6ebb61af1d0\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61af1d0\" class=\"collapseomatic_content \">This paper describes the system developed by Intelligent Voice<br \/>\nfor IberSpeech 2022 Albayzin Evaluations Speaker Diarization<br \/>\nand Identity Assignment Challenge (SDIAC). The presented<br \/>\nVariational Bayes x-vector Voice Print Extraction (VBxVPE)<br \/>\nsystem is capable of capturing the vocal variations using multiple<br \/>\nx-vector representations with two-stage clustering and outlier<br \/>\ndetection refinement and implements Deep-Encoder Convolutional<br \/>\nAutoencoder Denoiser (DE-CADE) network for denoising<br \/>\nsegments with noise and music for robust speaker<br \/>\nrecognition and diarization. When evaluated against the Radiotelevision<br \/>\nEspanola (RTVE) 2022 evaluation dataset, the<br \/>\nsystem was able to obtain a Diarization Error Rate (DER) of<br \/>\n..% and Error Rate of ..% . <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Roman Shrestha, Cornelius Glackin, Julie Wall and Nigel Cannings<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"150\"><strong>A.6<\/strong><br \/>\n17:20\u00a0 &#8211; 19:20<\/td>\n<td>ViVoLAB System Description for the S2TC IberSPEECH-RTVE 2022 challenge (<span class=\"collapseomatic \" id=\"id6a6ebb61af230\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61af230\" class=\"collapseomatic_content \">In this paper we describe the ViVoLAB system for the<br \/>\nIberSPEECH-RTVE 2022 Speech to Text Transcription Challenge.<br \/>\nThe system is a combination of several subsystems designed<br \/>\nto perform a full subtitle edition process from the raw audio<br \/>\nto the creation of aligned subtitle transcribed partitions. The<br \/>\nsubsystems include a phonetic recognizer, a phonetic subword<br \/>\nrecognizer, a speaker-aware subtitle partitioner, a sequence-tosequence<br \/>\ntranslation model working with orthographic tokens<br \/>\nto produce the desired transcription, and an optional diarization<br \/>\nstep with the previously estimated segments. Additionally, we<br \/>\nuse recurrent network based language models to improve results<br \/>\nfor steps that involve search algorithms like the subword<br \/>\ndecoder and the sequence-to-sequence model. The technologies<br \/>\ninvolved include unsupervised models like Wavlm to deal with<br \/>\nthe raw waveform, convolutional, recurrent, and transformer<br \/>\nlayers. As a general design pattern, we allow all the systems to<br \/>\naccess previous outputs or inner information, but the choice of<br \/>\nsuccessful communication mechanisms has been a difficult process<br \/>\ndue to the size of the datasets and long training times. The<br \/>\nbest solution found will be described and evaluated for some<br \/>\nreference tests of 2018 and 2020 IberSPEECH-RTVE S2TC<br \/>\nevaluations. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Antonio Miguel, Alfonso Ortega and Eduardo Lleida<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"150\"><strong>A.7<\/strong><br \/>\n17:20\u00a0 &#8211; 19:20<\/td>\n<td>GTTS Systems for the Albayzin 2022 Speech and Text Alignment Challenge (<span class=\"collapseomatic \" id=\"id6a6ebb61af27e\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6ebb61af27e\" class=\"collapseomatic_content \">This paper describes the most relevant features of the alignment<br \/>\napproach used by our research group (GTTS) for the Albayzin<br \/>\n2022 Text and Speech Alignment Challenge: Alignment of respoken<br \/>\nsubtitles (TaSAC-ST). It also presents and analyzes the<br \/>\nresults obtained by our primary and contrastive systems, focusing<br \/>\non the variability observed in the RTVE broadcasts used for<br \/>\nthis evaluation. The task is to provide some hypothesized start<br \/>\nand end times for each subtitle to be aligned. To that end, our<br \/>\nsystems decode the audio at the phonetic level using acoustic<br \/>\nmodels trained on external (non-RTVE) data, then align the recognized<br \/>\nsequence of phones with the phonetic transcription of<br \/>\nthe corresponding text and transfer the timestamps of the recognized<br \/>\nphones to the aligned text. The alignment error for each<br \/>\nsubtitle is computed as the sum of the absolute values of the<br \/>\nstart and end alignment errors (with regard to a manually supervised<br \/>\nground truth). The median of the alignment errors (MAE)<br \/>\nfor each broadcast is reported to compare system performance.<br \/>\nOur primary system yielded MAEs between 0.20 and 0.36 seconds<br \/>\non the development set, and between 0.22 and 1.30 seconds<br \/>\non the test set, with average MAEs of 0.295 and 0.395,<br \/>\nrespectively. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Germ\u00e1n Bordel, Luis Javier Rodriguez-Fuentes, Mikel Pe\u00f1agarikano and Amparo Varona<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<hr \/>\n","protected":false},"excerpt":{"rendered":"<p>Tuesday, 15 November Oral 3: Speech and Audio Processing Tuesday, 15 November 2022 (9:00-10:40) Chair: Alfonso Ortega O3.1 09:00\u00a0 &#8211; 09:20 On the potential of jointly-optimised solutions to spoofing attack detection and automatic speaker verification () Wanying Ge, Hemlata Tak, Massimiliano Todisco and Nicholas Evans O3.2 09:20\u00a0 &#8211; 09:40 A Study on the Use of&hellip;&nbsp;<a href=\"https:\/\/iberspeech2022.ugr.es\/?page_id=800\" rel=\"bookmark\">Leer m\u00e1s &raquo;<span class=\"screen-reader-text\">Technical Program, Day 2<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-800","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/pages\/800","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=800"}],"version-history":[{"count":26,"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/pages\/800\/revisions"}],"predecessor-version":[{"id":969,"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/pages\/800\/revisions\/969"}],"wp:attachment":[{"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=800"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}