{"id":795,"date":"2022-10-28T12:05:30","date_gmt":"2022-10-28T12:05:30","guid":{"rendered":"http:\/\/iberspeech2022.ugr.es\/?page_id=795"},"modified":"2022-11-09T08:57:30","modified_gmt":"2022-11-09T08:57:30","slug":"technical-program-monday-november-14th","status":"publish","type":"page","link":"https:\/\/iberspeech2022.ugr.es\/?page_id=795","title":{"rendered":"Technical Program, Day 1"},"content":{"rendered":"<h3 style=\"background-color: #f3f3f3;\"><span style=\"color: #2e75b7;\"><strong>Monday, November 14<\/strong><\/span><\/h3>\n<p><strong>Oral 1: Speech Synthesis<\/strong><br \/>\n<strong><span style=\"color: #d62013;\">Monday, 14 November 2022 (9:20-10:40)<\/span><\/strong><br \/>\n<strong>Chair: Antonio Bonafonte<\/strong><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O1.1<\/strong><br \/>\n9:20\u00a0 &#8211; 09:40<\/td>\n<td>Discrete Acoustic Space for an Efficient Sampling in Neural Text-To-Speech (<span class=\"collapseomatic \" id=\"id6a6eaa832d2c2\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d2c2\" class=\"collapseomatic_content \">We present a Split Vector Quantized Variational Autoencoder (SVQ-VAE) architecture using a split vector quantizer for NTTS, as an enhancement to the well-known Variational Autoencoder (VAE) and Vector Quantized Variational Autoencoder (VQ-VAE) architectures. Compared to these previous architectures, our proposed model retains the benefits of using an utterance-level bottleneck, while keeping significant representation power and a discretized latent space small enough for efficient prediction from text. We train the model on recordings in the expressive task-oriented dialogues domain and show that SVQ-VAE achieves a statistically significant improvement in naturalness over the VAE and VQ-VAE models. Furthermore, we demonstrate that the SVQ-VAE latent acoustic space is predictable from text, reducing the gap between the standard constant vector synthesis and vocoded recordings by 32%.<\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Marek Strelec, Jonas Rohnke, Antonio Bonafonte, Mateusz Lajszczak, Trevor Wood<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O1.2<\/strong><br \/>\n09:40 &#8211; 10:00<\/td>\n<td>An animated realistic head with vocal tract for the finite element simulation of vowel \/a\/ (<span class=\"collapseomatic \" id=\"id6a6eaa832d33d\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d33d\" class=\"collapseomatic_content \">Three-dimensional (3D) acoustic models can accurately simulate the voice production mechanism. These models require detailed 3D vocal tract geometries through which sound waves propagate. A few open source databases typically based on magnetic resonance imaging (MRI) are already available in literature. However, the 3D geometries they contain are mainly focused on the vocal tract and remove the head, which limits the computational domain of the simulations. This work develops a unified model consisting of an MRI-based vocal tract geometry set in a realistic head. The head is generated from scratch based on anatomical data of another subject, and contains different layers that add an organic appearance to the character. It is then not only designed to allow accurate finite element simulations of vowels, but more importantly, it can also be animated to add a realistic visual layer to the generated sound. This is expected to help in the dissemination of results and also to open potential applications in the audiovisual and animation sector. This paper<br \/>\nshows the first results of the model focusing on the vowel \/a\/. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Marc Arnela, Leonardo Pereira-Vivas, Jorge Egea<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O1.3<\/strong><br \/>\n10:00\u00a0 &#8211; 10:20<\/td>\n<td>Exploring the limits of neural voice cloning: A case study on two well-known personalities (<span class=\"collapseomatic \" id=\"id6a6eaa832d38e\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d38e\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">This work describes one successful and one failed Voice Cloning processes of two <\/span><span dir=\"ltr\" role=\"presentation\">famous personalities in order to be broadcast in a high-impact podcast and in a <\/span><span dir=\"ltr\" role=\"presentation\">Spanish public television program.<\/span> <span dir=\"ltr\" role=\"presentation\">Whilst a good quality synthesised voice could<\/span> <span dir=\"ltr\" role=\"presentation\">be generated for the first public figure, the second one was not adequate enough for <\/span><span dir=\"ltr\" role=\"presentation\">its broadcast on television given its low speech quality. In this study, we explore the<\/span> <span dir=\"ltr\" role=\"presentation\">limits of the neural voice cloning considering the different conditions of the training <\/span><span dir=\"ltr\" role=\"presentation\">material employed in each case and, based on several objective measures (amount<\/span> <span dir=\"ltr\" role=\"presentation\">of training data, phoneme coverage, SNR, MCD and PESQ), we analysed the main <\/span><span dir=\"ltr\" role=\"presentation\">features to be considered for a high-quality synthetic voice generation. In addition,<\/span> <span dir=\"ltr\" role=\"presentation\">a webpage is provided in which samples of the resulting audios are available for each <\/span><span dir=\"ltr\" role=\"presentation\">cloning model.<\/span> <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Ander Gonz\u00e1lez-Docasal, Aitor \u00c1lvarez, Haritz Arzelus<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O1.4<\/strong><br \/>\n10:20\u00a0 &#8211; 10:40<\/td>\n<td>Analysis of iterative adaptive and quasi closed phase inverse filtering techniques on OPENGLOT synthetic vowels (<span class=\"collapseomatic \" id=\"id6a6eaa832d3e7\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d3e7\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Three-dimensional source-filter models allow for the articulatory-based genera<\/span><span dir=\"ltr\" role=\"presentation\">tion of voice but with limited expressiveness yet.<\/span> <span dir=\"ltr\" role=\"presentation\">From the analysis of expressive <\/span><span dir=\"ltr\" role=\"presentation\">speech corpora through glottal inverse filtering techniques, it has been observed <\/span><span dir=\"ltr\" role=\"presentation\">that both the vocal tract and the glottal source play a key role in the generation <\/span><span dir=\"ltr\" role=\"presentation\">of different phonation types. However, the accuracy of the source-filter decomposi<\/span><span dir=\"ltr\" role=\"presentation\">tion depends on the considered technique. Current Quasi Closed Phase (QCP) and<\/span> <span dir=\"ltr\" role=\"presentation\">Iterative Adaptive Inverse Filtering (IAIF) based approaches present pretty good<\/span> <span dir=\"ltr\" role=\"presentation\">results, despite difficult to compare as they are obtained from different experiments. <\/span><span dir=\"ltr\" role=\"presentation\">This work aims at evaluating the performance of these stateof- the-art methods<\/span> <span dir=\"ltr\" role=\"presentation\">on the reference OPENGLOT database, using its repository with synthetic vow<\/span><span dir=\"ltr\" role=\"presentation\">els generated with different phonation types and fundamental frequencies.<\/span> <span dir=\"ltr\" role=\"presentation\">After <\/span><span dir=\"ltr\" role=\"presentation\">optimizing the parameters of each inverse filtering approach, their performance is<\/span> <span dir=\"ltr\" role=\"presentation\">compared considering typical glottal flow error measures.<\/span> <span dir=\"ltr\" role=\"presentation\">The results show that <\/span><span dir=\"ltr\" role=\"presentation\">QCP-based techniques attain statistically significant lower values in most measures.<\/span> <span dir=\"ltr\" role=\"presentation\">IAIF variants achieve a significant improvement on the spectral tilt error measure <\/span><span dir=\"ltr\" role=\"presentation\">with respect to the original IAIF, but they are surpassed by QCP when spectral tilt <\/span><span dir=\"ltr\" role=\"presentation\">compensation is applied.<\/span> <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Marc Freixes, Joan Claudi Socor\u00f3 and Francesc Al\u00edas<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Keynote 1<\/strong><br \/>\n<span style=\"color: #d62013;\"><strong>Monday, 14 November 2022 (11:00-12:00)<\/strong><\/span><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>KN1<\/strong><br \/>\n11:00\u00a0 &#8211; 12:00<\/td>\n<td>Secure and explainable voice biometrics (<span class=\"collapseomatic \" id=\"id6a6eaa832d436\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d436\" class=\"collapseomatic_content \">Anti-spoofing for voice biometrics is now an established area of research, thanks to the four competitive ASVspoof challenges (the fifth is currently underway) that have taken place over the past decade. Growing research effort has invested, firstly, in the development of front-end representations that capture more reliably the tell-tale artefacts that are indicative of utterances generated with text-to-speech and voice conversion algorithms and, secondly, in the development of deep and end-to-end solutions. Despite enormous efforts and positive achievements, little is still known about the artefacts these recognisers use to identify spoofing utterances or distinguish between bona fide and spoofed. Although many unanswered questions remain, this talk aims to provide insights and inspirations, through examples, into the behaviour of voice anti-spoofing systems. Particular attention will be given to data augmentation and boosting methods that have been shown instrumental to reliability. The ultimate goal is to better understand these artefacts from a physical and perceptual point of view and how they are actually seen by automatic processes, which puts us in a better position to design more reliable countermeasures. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Massimiliano Todisco<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Posters 1: Topics on Speech and Language Technologies <\/strong><br \/>\n<span style=\"color: #d62013;\"><strong>Monday, 14 November 2022 (12:00-13:30)<\/strong><\/span><br \/>\n<strong>Chair:\u00a0 Antonio Teixeira<\/strong><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.1<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>An Experimental Study on Light Speech Features for Small-Footprint Keyword Spotting (<span class=\"collapseomatic \" id=\"id6a6eaa832d484\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d484\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Keyword spotting (KWS) is, in many instances, intended to run on smart elec<\/span><span dir=\"ltr\" role=\"presentation\">tronic devices characterized by limited computational resources. To meet computa<\/span><span dir=\"ltr\" role=\"presentation\">tional constraints, a series of techniques \u2014ranging from feature and acoustic model<\/span> <span dir=\"ltr\" role=\"presentation\">parameter quantization to the reduction of the number of model parameters and <\/span><span dir=\"ltr\" role=\"presentation\">required multiplications\u2014 has been explored in the literature. With this same aim, <\/span><span dir=\"ltr\" role=\"presentation\">in this paper, we study a straightforward alternative consisting of the reduction of<\/span> <span dir=\"ltr\" role=\"presentation\">the spectro\/cepstro-temporal resolution of log-Mel and Melfrequency cepstral coeffi<\/span><span dir=\"ltr\" role=\"presentation\">cient feature matrices commonly employed in KWS. We show that the feature matrix <\/span><span dir=\"ltr\" role=\"presentation\">size has a strong impact on the number of multiplications\/energy consumption of a<\/span> <span dir=\"ltr\" role=\"presentation\">state-of-the-art KWS acoustic model based on convolutional neural network. Exper<\/span><span dir=\"ltr\" role=\"presentation\">imental results demonstrate that the number of elements in commonly used speech <\/span><span dir=\"ltr\" role=\"presentation\">feature matrices can be reduced by a factor of 8 while essentially maintaining KWS<\/span> <span dir=\"ltr\" role=\"presentation\">performance. Even more interestingly, this size reduction leads to a 9.6<\/span><span dir=\"ltr\" role=\"presentation\">\u00d7<\/span> <span dir=\"ltr\" role=\"presentation\">number <\/span><span dir=\"ltr\" role=\"presentation\">of multiplications\/energy consumption, 4.0<\/span><span dir=\"ltr\" role=\"presentation\">\u00d7<\/span> <span dir=\"ltr\" role=\"presentation\">training time and 3.7<\/span><span dir=\"ltr\" role=\"presentation\">\u00d7<\/span> <span dir=\"ltr\" role=\"presentation\">inference time<\/span> <span dir=\"ltr\" role=\"presentation\">reduction.<\/span> <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Iv\u00e1n L\u00f3pez-Espejo, Zheng-Hua Tan and Jesper Jensen<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.2<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>S3prl-Disorder: Open-Source Voice Disorder Detection System based in the Framework of S3PRL-toolkit (<span class=\"collapseomatic \" id=\"id6a6eaa832d4d1\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d4d1\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">This paper introduces S3prl-Disorder, an open-source toolkit for Automatic Voice <\/span><span dir=\"ltr\" role=\"presentation\">Disorder Detection (AVDD) developed in the framework of the S3prl toolkit.<\/span> <span dir=\"ltr\" role=\"presentation\">It <\/span><span dir=\"ltr\" role=\"presentation\">focuses on a binary classification task between healthy and pathological speech in<\/span> <span dir=\"ltr\" role=\"presentation\">the Saarbruecken Voice Database (SVD). However, the framework left room for <\/span><span dir=\"ltr\" role=\"presentation\">following extensions to multi-class classification to differentiate among pathologies<\/span> <span dir=\"ltr\" role=\"presentation\">and to incorporate more datasets. This work aims to contribute on the development <\/span><span dir=\"ltr\" role=\"presentation\">of automatic systems for diagnosis, treatment, and monitoring of voice pathologies in<\/span> <span dir=\"ltr\" role=\"presentation\">a common framework, that allows reproducibility and comparability among systems <\/span><span dir=\"ltr\" role=\"presentation\">and results<\/span>. <\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Dayana Ribas, Miguel Angel Pastor Yoldi, Antonio Miguel, David Mart\u00ednez, Alfonso Ortega and Eduardo Lleida<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.3<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Active Learning Improves the Teacher&#8217;s Experience: A Case Study in a Language Grounding Scenario (<span class=\"collapseomatic \" id=\"id6a6eaa832d51e\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d51e\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Active Learning, that is, assigning the responsibility of learning to the students, is <\/span><span dir=\"ltr\" role=\"presentation\">an important tool in education as it makes the students become engaged in and think <\/span><span dir=\"ltr\" role=\"presentation\">about the things they do. A similar concept was adopted in the context of Machine<\/span> <span dir=\"ltr\" role=\"presentation\">Learning as a means to reduce the annotation effort by selecting the examples that <\/span><span dir=\"ltr\" role=\"presentation\">are most relevant or provide more information at a given time.<\/span> <span dir=\"ltr\" role=\"presentation\">Most studies on<\/span> <span dir=\"ltr\" role=\"presentation\">this subject focus on the learner\u2019s performance. However, in interactive scenarios, <\/span><span dir=\"ltr\" role=\"presentation\">the teacher\u2019s experience is also a relevant aspect, as it affects their willingness to<\/span> <span dir=\"ltr\" role=\"presentation\">interact with artificial learners. In this paper, we address that aspect by performing <\/span><span dir=\"ltr\" role=\"presentation\">a case study in a language grounding scenario, in which humans have to engage in<\/span> <span dir=\"ltr\" role=\"presentation\">dialog with a learning agent and teach it how to recognize observations of certain <\/span><span dir=\"ltr\" role=\"presentation\">objects. Overall, the results of our experiments show that humans prefer to interact<\/span> <span dir=\"ltr\" role=\"presentation\">with an active learner, as it seems more intelligent, gives them a better perception <\/span><span dir=\"ltr\" role=\"presentation\">of its knowledge, and makes the dialog more natural and enjoyable.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Filipe Reynaud, Eug\u00e9nio Ribeiro and David Martins de Matos<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.4<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>The role of window length and shift in complex-domain DNN-based speech enhancement (<span class=\"collapseomatic \" id=\"id6a6eaa832d56f\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d56f\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Deep learning techniques have widely been applied to speech enhancement as <\/span><span dir=\"ltr\" role=\"presentation\">they show outstanding modeling capabilities that are needed for proper speech-<\/span><span dir=\"ltr\" role=\"presentation\">noise separation. In contrast to other end-to-end approaches, masking-based meth<\/span><span dir=\"ltr\" role=\"presentation\">ods consider speech spectra as input to the deep neural network, providing spectral <\/span><span dir=\"ltr\" role=\"presentation\">masks for noise removal or attenuation. In these approaches, the Short-Time Fourier<\/span> <span dir=\"ltr\" role=\"presentation\">Transform (STFT) and, particularly, the parameters used for the analysis\/synthesis <\/span><span dir=\"ltr\" role=\"presentation\">window, plays an important role which is often neglected. In this paper, we analyze<\/span> <span dir=\"ltr\" role=\"presentation\">the effects of window length and shift on a complex-domain convolutional-recurrent <\/span><span dir=\"ltr\" role=\"presentation\">neural network (DCCRN) which is able to provide, separately, magnitude and phase<\/span> <span dir=\"ltr\" role=\"presentation\">corrections.<\/span> <span dir=\"ltr\" role=\"presentation\">Different perceptual quality and intelligibility objective metrics are <\/span><span dir=\"ltr\" role=\"presentation\">used to assess its performance.<\/span> <span dir=\"ltr\" role=\"presentation\">As a result, we have observed that phase correc-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">tions have an increased impact with shorter window sizes.<\/span> <span dir=\"ltr\" role=\"presentation\">Similarly, as window<\/span> <span dir=\"ltr\" role=\"presentation\">overlap increases, phase takes more relevance than magnitude spectrum in speech <\/span><span dir=\"ltr\" role=\"presentation\">enhancement.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Celia Garc\u00eda-Ruiz, Angel M. Gomez and Juan M. Mart\u00edn-Do\u00f1as<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.5<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Neural Detection of Cross-lingual Syntactic Knowledge (<span class=\"collapseomatic \" id=\"id6a6eaa832d5bd\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d5bd\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">In recent years, there has been prominent development in pretrained multilingual<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">language models, such as mBERT, XLMR, etc., which are able to capture and<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">learn linguistic knowledge from input across a variety of languages simultaneously.<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">However, little is known about where multilingual models localise what they have<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">learnt across languages. In this paper, we specifically evaluate cross-lingual syntactic<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">information embedded in CINO, a more recent multilingual pre-trained language<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">model.<\/span> <span dir=\"ltr\" role=\"presentation\">We probe CINO on Universal Dependencies treebank datasets of English<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">and Chinese Mandarin for two syntax-related layerwise evaluation tasks: Part-of-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Speech Tagging at token level and Syntax Tree-depth Prediction at sentence level.<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">The results of our layer-wise probing experiments show that token-level syntax is<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">localisable in higher layers and consistency is shown across the typologically different<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">languages, whereas sentencelevel syntax is distributed across the layers in typology-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">specific and universal manners.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Yongjian Chen and Mireia Farr\u00fas<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.6<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Efficient Transformers for End-to-End Neural Speaker Diarization (<span class=\"collapseomatic \" id=\"id6a6eaa832d60c\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d60c\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">The recently proposed End-to-End Neural speaker Diarization framework (EEND)<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">handles speech overlap and speech activity detection natively. While extensions of<\/span> <span dir=\"ltr\" role=\"presentation\">this work have reported remarkable results in both two-speaker and multispeaker<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">diarization scenarios, these come at the cost of a long training process that re<\/span><span dir=\"ltr\" role=\"presentation\">quires considerable memory and computational power.<\/span> <span dir=\"ltr\" role=\"presentation\">In this work, we explore<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the integration of efficient transformer variants into the Self-Attentive EEND with<\/span> <span dir=\"ltr\" role=\"presentation\">Encoder-Decoder based Attractors (SA-EEND EDA) architecture. Since it is based<\/span> <span dir=\"ltr\" role=\"presentation\">on Transformers, the cost of training SA-EEND EDA is driven by the quadratic time<\/span> <span dir=\"ltr\" role=\"presentation\">and memory complexity of their self-attention mechanism. We verify that the use of <\/span><span dir=\"ltr\" role=\"presentation\">a linear attention mechanism in SA-EEND EDA decreases GPU memory usage by<\/span> <span dir=\"ltr\" role=\"presentation\">22%. We conduct experiments to measure how the increased efficiency of the train<\/span><span dir=\"ltr\" role=\"presentation\">ing process translates into the two-speaker diarization error rate on CALLHOME, <\/span><span dir=\"ltr\" role=\"presentation\">quantifying the impact of increasing the size of the batch, the model or the sequence<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">length on training time and diarization performance.<\/span> <span dir=\"ltr\" role=\"presentation\">In addition, we propose an <\/span><span dir=\"ltr\" role=\"presentation\">architecture combining linear and softmax attention that achieves an acceleration<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">of 12% with a small relative DER degradation of 2%, while using the same GPU<\/span> <span dir=\"ltr\" role=\"presentation\">memory as the softmax attention baseline.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Sergio Izquierdo del Alamo, Beltr\u00e1n Labrador, Alicia Lozano-Diez and Doroteo T. Toledano<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.7<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>CORAA NURC-SP Minimal Corpus: a manually annotated corpus of Brazilian Portuguese spontaneous speech (<span class=\"collapseomatic \" id=\"id6a6eaa832d65a\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d65a\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">With the advent of technology, the availability of linguistic data in digital format <\/span><span dir=\"ltr\" role=\"presentation\">has been increasingly encouraged to facilitate its use not only in different areas of<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Linguistics but also in related areas, such as natural language processing. Inspired <\/span><span dir=\"ltr\" role=\"presentation\">by a protocol for digitizing the NURC (\u2018Cultured Linguistic Urban Norm\u2019) project<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">collection \u2014 one of the most influential in Brazilian Linguistics \u2014, this paper aims to <\/span><span dir=\"ltr\" role=\"presentation\">present the text-to-speech alignment process of the NURC-Sao Paulo Minimal \u0303 Cor<\/span><span dir=\"ltr\" role=\"presentation\">pus. <\/span><span dir=\"ltr\" role=\"presentation\">This subcorpus comprises 21 audio files and audioaligned multilevel transcripts<\/span> <span dir=\"ltr\" role=\"presentation\">according to linguistically motivated intonation units (<\/span><span dir=\"ltr\" role=\"presentation\">\u2243<\/span><span dir=\"ltr\" role=\"presentation\">18 hours,<\/span> <span dir=\"ltr\" role=\"presentation\">\u2243<\/span><span dir=\"ltr\" role=\"presentation\">155 k words),<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">covering three text genres. The dataset \u2014 currently used to evaluate methods for <\/span><span dir=\"ltr\" role=\"presentation\">processing the entire NURC-SP corpus \u2014 is publicly available on the Portulan Clarin<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">repository [CC BY-NC-ND 4.0] (https:\/\/hdl.handle.net\/21.11129\/0000-000F-73CA-<\/span><span dir=\"ltr\" role=\"presentation\">C).<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Vin\u00edcius G. Santos, Caroline Adriane Alves, Bruno Baldissera Carlotto, Bruno Angelo Papa Dias, Lucas Rafael Stefanel Gris, Renan de Lima Izaias, Maria Luiza Azevedo de Morais, Paula Marin de Oliveira, Rafael Sicoli, Flaviane Romani Fernandes Svartman, Marli Quadros Leite and Sandra Maria Alu\u00edsio<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.8<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Speaker Characterization by means of Attention Pooling (<span class=\"collapseomatic \" id=\"id6a6eaa832d6b5\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d6b5\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">State-of-the-art Deep Learning systems for speaker verification are commonly<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">based on speaker embedding extractors. These architectures are usually composed of<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">a feature extractor front-end together with a pooling layer to encode variable length<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">utterances into fixed-length speaker vectors. The authors have recently proposed the<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">use of a Double Multi-Head SelfAttention pooling for speaker recognition, placed<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">between a CNN-based front-end and a set of fully connected layers. This has shown<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">to be an excellent approach to efficiently select the most relevant features captured<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">by the front-end from the speech signal. In this paper we show excellent experimental<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">results by adapting this architecture to other different speaker characterization tasks,<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">such as emotion recognition, sex classification and COVID-19 detection.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Federico Costa, Miquel India and Javier Hernando<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.9<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Enhancing the Design of a Conversational Agent for an Ethical Interaction with Children (<span class=\"collapseomatic \" id=\"id6a6eaa832d703\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d703\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Conversational agents (CAs) have become one of the most popular applications <\/span><span dir=\"ltr\" role=\"presentation\">of speech and language technologies in the last decade. Those agents employ speech <\/span><span dir=\"ltr\" role=\"presentation\">interaction to perform several tasks, from information retrieval to purchase goods<\/span> <span dir=\"ltr\" role=\"presentation\">from on-line stores. However, these agents are defined to address a general sector <\/span><span dir=\"ltr\" role=\"presentation\">of population, mainly adults without speech production problems, and then they<\/span> <span dir=\"ltr\" role=\"presentation\">fail to obtain a similar performance with specific groups, such as elderly or children. <\/span><span dir=\"ltr\" role=\"presentation\">The case of children is particularly interesting because they naturally engage in<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">interaction with those CAs and they have special needs in terms of technical and<\/span> <span dir=\"ltr\" role=\"presentation\">ethical considerations. Therefore, CAs must fulfil some conditions that could affect <\/span><span dir=\"ltr\" role=\"presentation\">their general design in order to provide a trustworthy interaction with children.<\/span> <span dir=\"ltr\" role=\"presentation\">In this article we present how to improve a general CA design to fulfil the specific <\/span><span dir=\"ltr\" role=\"presentation\">ethical needs of children interaction. We address the development of a CA devoted to<\/span> <span dir=\"ltr\" role=\"presentation\">complete a wish list of games using user preferences, and its improvements towards <\/span><span dir=\"ltr\" role=\"presentation\">children.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Marina Escobar-Planas, Emilia G\u00f3mez, Carlos-D. Mart\u00ednez-Hinarejos <\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.10<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Sentiment Analysis in Portuguese Dialogues (<span class=\"collapseomatic \" id=\"id6a6eaa832d752\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d752\" class=\"collapseomatic_content \">S<span dir=\"ltr\" role=\"presentation\">entiment analysis in dialogue aims at detecting the sentiment expressed in the<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">utterances of a conversation, which may improve human-computer interaction in nat<\/span><span dir=\"ltr\" role=\"presentation\">ural language. In this paper, we explore different approaches for sentiment analysis<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">in written Portuguese dialogues, mainly related to customer support in Telecommu<\/span><span dir=\"ltr\" role=\"presentation\">nications. If integrated into a conversational agent, this will enable the automatic<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">identification and a quick reaction upon clients manifesting negative sentiments,<\/span> <span dir=\"ltr\" role=\"presentation\">possibly with human intervention, hopefully minimising the damage. Experiments<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">were performed in two manually annotated real datasets: one with dialogues from <\/span><span dir=\"ltr\" role=\"presentation\">the call-center of a Telecommunications company (TeleComSA); another of Twitter<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">conversations primarily involving accounts of Telecommunications companies.<\/span> <span dir=\"ltr\" role=\"presentation\">We <\/span><span dir=\"ltr\" role=\"presentation\">compare the performance of different machine learning approaches, from traditional<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">to more recent, with and without considering previous utterances. The Finetuned<\/span> <span dir=\"ltr\" role=\"presentation\">BERT achieved the highest F1 Scores in both datasets, 0.87 in the Twitter dataset,<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">without context, and 0.93 in the TeleComSA, considering context.<\/span> <span dir=\"ltr\" role=\"presentation\">These are in<\/span><span dir=\"ltr\" role=\"presentation\">teresting results and suggest that automated customer-support may benefit from<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">sentiment detection. Another interesting finding was that most models did not ben<\/span><span dir=\"ltr\" role=\"presentation\">efit from using previous utterances, suggesting that, in this scenario, context does<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">not contribute much, and classifying the current utterance can be enough.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Isabel Carvalho, Hugo Gon\u00e7alo Oliveira and Catarina Silva<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.11<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>On the application of conformers to logical access voice spoofing attack detection (<span class=\"collapseomatic \" id=\"id6a6eaa832d7a2\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d7a2\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Biometric systems are exposed to spoofing attacks which may compromise their <\/span><span dir=\"ltr\" role=\"presentation\">security, and automatic speaker verification (ASV) is no exception.<\/span> <span dir=\"ltr\" role=\"presentation\">To increase<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the robustness against such attacks, anti-spoofing systems have been proposed for<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the detection of spoofed audio attacks.<\/span> <span dir=\"ltr\" role=\"presentation\">However, most of these systems can not<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">capture long-term feature dependencies and can only extract local features. While<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">transformers are an excellent solution for the exploitation of these long-distance<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">correlations, they may degrade local details. On the contrary, convolutional neural<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">networks (CNNs) are a powerful tool for extracting local features but not so much<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">for capturing global representations. The conformer is a model that combines the<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">best of both techniques, CNNs and transformers, to model both local and global<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">dependencies and has been used for speech recognition achieving state-of-the-art<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">performance. While conformers have been mainly applied to sequence-to-sequence<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">problems, in this work we make a preliminary study of their adaptation to a bi-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">nary classification task such as anti-spoofing, with focus on synthesis and voice-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">conversion-based attacks. To evaluate our proposals, experiments were carried out<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">on the ASVspoof 2019 logical access database. The experimental results show that<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the proposed system can obtain encouraging results, although more research will be<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">required in order to outperform other state-of-the-art systems.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Eros Rosello, Alejandro Gomez-Alanis, Manuel Chica, Angel M. Gomez, Jose A. Gonzalez and Antonio M. Peinado<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.12<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Speech emotion recognition in Spanish TV Debates (<span class=\"collapseomatic \" id=\"id6a6eaa832d7fd\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d7fd\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Emotion recognition from speech is an active field of study that can help build<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">more natural human\u2013machine interaction systems. Even though the advancement<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">of deep learning technology has brought improvements in this task, it is still a very<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">challenging field.<\/span> <span dir=\"ltr\" role=\"presentation\">For instance, when considering real life scenarios, things such<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">as tendency toward neutrality or the ambiguous definition of emotion can make<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">labeling a difficult task causing the data-set to be severally imbalanced and not<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">very representative.<\/span> <span dir=\"ltr\" role=\"presentation\">In this work we considered a real life scenario to carry out a<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">series of emotion classification experiments. Specifically, we worked with a labeled<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">corpus consisting of a set of audios from Spanish TV debates and their respective<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">transcriptions.<\/span> <span dir=\"ltr\" role=\"presentation\">First, an analysis of the emotional information within the corpus<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">was conducted.<\/span> <span dir=\"ltr\" role=\"presentation\">Then different data representations were analyzed as to choose<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the best one for our task; Spectrograms and UniSpeech-SAT were used for audio<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">representation and DistilBERT for text representation. As a final step, Multimodal<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Machine Learning was used with the aim of improving the obtained classification<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">results by combining acoustic and textual information.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Irune Zubiaga, Raquel Justo, M. In\u00e9s Torres and Mikel De Velasco<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.13<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Assessing Transfer Learning and automatically annotated data in the development of Named Entity Recognizers for new domains (<span class=\"collapseomatic \" id=\"id6a6eaa832d84b\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d84b\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">With recent advances Deep Learning, pretrained models and Transfer Learn<\/span><span dir=\"ltr\" role=\"presentation\">ing, the lack of labeled data has become the biggest bottleneck preventing use of <\/span><span dir=\"ltr\" role=\"presentation\">Named Entity Recognition (NER) in more domains and languages. To relieve the<\/span> <span dir=\"ltr\" role=\"presentation\">pressure of costs and time in the creation of annotated data for new domains, we <\/span><span dir=\"ltr\" role=\"presentation\">proposed recently automatic annotation by an ensemble of NERs to get data to<\/span> <span dir=\"ltr\" role=\"presentation\">train a Bidirectional Encoder Representations from Transformers (BERT) based <\/span><span dir=\"ltr\" role=\"presentation\">NER for Portuguese and made a first evaluation. Results demonstrated the method<\/span> <span dir=\"ltr\" role=\"presentation\">has potential but were limited to one domain. Having as main objective a more in<\/span><span dir=\"ltr\" role=\"presentation\">depth assessment of the method capabilities, this paper presents: (1) evaluation of<\/span> <span dir=\"ltr\" role=\"presentation\">the method in other domains; (2) assessment of the generalization capabilities of the <\/span><span dir=\"ltr\" role=\"presentation\">trained models, by applying them to new domains without retraining; (3) assessment<\/span> <span dir=\"ltr\" role=\"presentation\">of additional training with in-domain data, also automatically annotated.<\/span> <span dir=\"ltr\" role=\"presentation\">Evalu<\/span><span dir=\"ltr\" role=\"presentation\">ation, performed using the test part of MiniHAREM, Paramopama and LeNER<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Portuguese datasets, confirmed the potential of the approach and demonstrated the<\/span> <span dir=\"ltr\" role=\"presentation\">capability of models previously trained for tourism domain to recognize entities in<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">new domains, with better performance for entities of types PERSON, LOCAL and<\/span> <span dir=\"ltr\" role=\"presentation\">ORGANIZATION.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Emanuel Matos, M\u00e1rio Rodrigues and Ant\u00f3nio Teixeira<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.14<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>On the detection of acoustic events for public security: the challenges of the counter-terrorism domain (<span class=\"collapseomatic \" id=\"id6a6eaa832d899\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d899\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Massive amounts of audio-visual contents are shared in public platforms every<\/span><span dir=\"ltr\" role=\"presentation\">day. These contents are created with many purposes, from entertaining or teaching,<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">to extremist propaganda. Civil security actors need to monitor these platforms to<\/span> <span dir=\"ltr\" role=\"presentation\">detect and neutralize security threats. Generating actionable knowledge from multi<\/span><span dir=\"ltr\" role=\"presentation\">media contents requires the extraction of multiple information, from linguistic data<\/span> <span dir=\"ltr\" role=\"presentation\">to sounds and background noises. Information extraction demands audio-visual an-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">notations, a costly, time-consuming task when performed manually, which hinders <\/span><span dir=\"ltr\" role=\"presentation\">the analysis of such an overwhelming amount of data. This work, performed in the<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">context of the EU Horizon 2020 Project AIDA, addresses the challenge of building<\/span> <span dir=\"ltr\" role=\"presentation\">a robust sound detector focused on events relevant to the counterterrorism domain.<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Our classification framework combines PLP features with a convolutional architec<\/span><span dir=\"ltr\" role=\"presentation\">ture to train a scalable model on a large number of events that is later fine-tuned on<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the subset of interest. The fusion of different corpora was also investigated, showing <\/span><span dir=\"ltr\" role=\"presentation\">the difficulties posed by this task.<\/span> <span dir=\"ltr\" role=\"presentation\">With our framework, results attained an aver<\/span><span dir=\"ltr\" role=\"presentation\">age F1-score of 0.53% on the target set of events. Of relevance, during the fine-tune<\/span> <span dir=\"ltr\" role=\"presentation\">phase a general-purpose class was introduced, which allowed the model to generalize<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">on \u2019unseen\u2019 events, highlighting the importance of a robust fine-tune.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Anna Pompili, Tiago Lu\u00eds, Nuno Monteiro, Jo\u00e3o Miranda, Carlo Mendes and S\u00e9rgio Paulo<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.15<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Database dependence comparison in detection of physical access voice spoofing attacks (<span class=\"collapseomatic \" id=\"id6a6eaa832d8e8\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d8e8\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">The antispoofing challenges are designed to work on a single database, on which <\/span><span dir=\"ltr\" role=\"presentation\">we can test our model.<\/span> <span dir=\"ltr\" role=\"presentation\">The automatic speaker verification spoofing and counter<\/span><span dir=\"ltr\" role=\"presentation\">measures (ASVspoof) [1] challenge series is a community-led initiative that aims<\/span> <span dir=\"ltr\" role=\"presentation\">to promote the consideration of spoofing and the development of countermeasures. <\/span><span dir=\"ltr\" role=\"presentation\">In general, the idea of analyzing the databases individually has been the dominant<\/span> <span dir=\"ltr\" role=\"presentation\">approach but this could be rather misleading. This paper provides a study of the<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">generalization capability of antispoofing systems based on neural networks by com<\/span><span dir=\"ltr\" role=\"presentation\">bining different databases for training and testing.<\/span> <span dir=\"ltr\" role=\"presentation\">We will try to give a broader<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">vision of the advantages of grouping different datasets. We will delve into the \u201dre<\/span><span dir=\"ltr\" role=\"presentation\">play attacks\u201d on physical data. This type of attack is one of the most difficult to<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">detect since only a few minutes of audio samples are needed to impersonate the<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">voice of a genuine speaker and gain access to the ASV system.<\/span> <span dir=\"ltr\" role=\"presentation\">To carry out this<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">task, the ASV databases from ASVspoof-challenge [2], [3],[4] have been chosen and<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">will be used to have a more concrete and accurate vision of them. We report results<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">on these databases using different neural network architectures and set-ups<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Manuel Chica, Alejandro Gomez-Alanis, Eros Rosello, Angel M. Gomez, Jose A. Gonzalez and Antonio M. Peinado<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>P1.16<\/strong><br \/>\n12:00 \u2013 13:30<\/td>\n<td>Measuring trust at zero-acquaintance using acted-emotional videos (<span class=\"collapseomatic \" id=\"id6a6eaa832d947\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d947\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">rustworthiness recognition attracts the attention of the research community due<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">to its main role in social communications. However, few datasets are available and<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">there are still many dimensions of trust to investigate. This paper presents a study<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">of an annotation tool for creating of a trustworthiness corpus. Specifically, we asked<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the participants to rate short emotional videos extracted from RAVDESS at zero<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">acquaintance and studied the relationship between their trustworthiness score and<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">other characteristics of the subjects of each video. Eloquence (<\/span><span dir=\"ltr\" role=\"presentation\">\u03c1<\/span> <span dir=\"ltr\" role=\"presentation\">= 0<\/span><span dir=\"ltr\" role=\"presentation\">.<\/span><span dir=\"ltr\" role=\"presentation\">41), kindness<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">(<\/span><span dir=\"ltr\" role=\"presentation\">\u03c1<\/span> <span dir=\"ltr\" role=\"presentation\">= 0<\/span><span dir=\"ltr\" role=\"presentation\">.<\/span><span dir=\"ltr\" role=\"presentation\">32), attractiveness (<\/span><span dir=\"ltr\" role=\"presentation\">\u03c1<\/span> <span dir=\"ltr\" role=\"presentation\">= 0<\/span><span dir=\"ltr\" role=\"presentation\">.<\/span><span dir=\"ltr\" role=\"presentation\">34), and authenticity of emotion transmitted<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">(<\/span><span dir=\"ltr\" role=\"presentation\">\u03c1<\/span> <span dir=\"ltr\" role=\"presentation\">= 0<\/span><span dir=\"ltr\" role=\"presentation\">.<\/span><span dir=\"ltr\" role=\"presentation\">6) are shown to be important determinants of perceived trustworthiness. In<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">addition, we have measured a strong association between some of the variables under<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">study. For example, physical beauty and voice pleasantness obtain a<\/span> <span dir=\"ltr\" role=\"presentation\">\u03c1<\/span> <span dir=\"ltr\" role=\"presentation\">= 0<\/span><span dir=\"ltr\" role=\"presentation\">.<\/span><span dir=\"ltr\" role=\"presentation\">71, or<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">eloquence and expressiveness (<\/span><span dir=\"ltr\" role=\"presentation\">\u03c1<\/span> <span dir=\"ltr\" role=\"presentation\">= 0<\/span><span dir=\"ltr\" role=\"presentation\">.<\/span><span dir=\"ltr\" role=\"presentation\">65), which opens a future line of investigation to<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">study how people understand attractiveness and eloquence from these perspectives.<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Finally, an attribute selection strategy identified that frequency and spectral-related<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">attributes could be accurate aural indicators of perceived trustworthiness.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Cristina Luna Jim\u00e9nez, Syaheerah Lebai Lutfi, Manuel Gil-Mart\u00edn, Ricardo Kleinlein, Juan M. Montero and Fernando Fern\u00e1ndez-Mart\u00ednez<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>Oral 2: Automatic Speech Recognition\u00a0 <\/strong><br \/>\n<span style=\"color: #d62013;\"><strong>Monday, 14 November 2022 (15:00-17:00)<\/strong><\/span><br \/>\n<strong>Chair: Hermann Ney<\/strong><\/p>\n<table width=\"0\">\n<tbody>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O2.1<\/strong><br \/>\n15:00\u00a0 &#8211; 15:20<\/td>\n<td>Galician\u2019s Language Technologies in the Digital Age\u00a0 (<span class=\"collapseomatic \" id=\"id6a6eaa832d997\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d997\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">This study was carried out under the initial state of the European Language<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">Equality project to report technology support for Europe\u2019s languages. In this pa<\/span><span dir=\"ltr\" role=\"presentation\">per, we show an overview of the current state of automatic speech recognition tech<\/span><span dir=\"ltr\" role=\"presentation\">nologies for Galician.<\/span> <span dir=\"ltr\" role=\"presentation\">In addition, we compare, over a small set of Galician TV<\/span> <span dir=\"ltr\" role=\"presentation\">shows, the performance of two of the most reported automatic recognition system<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">with support for Galician: the one developed by the University of Vigo and the one <\/span><span dir=\"ltr\" role=\"presentation\">offered by Google.<\/span> <span dir=\"ltr\" role=\"presentation\">Our research shows impressive growth in the amount of data<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">and resources created for Galician in the last four years. However, the scope of the<\/span> <span dir=\"ltr\" role=\"presentation\">resources and the range of tools are still limited, especially in the actual context of<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">services and technologies based on artificial intelligence and big data. The current<\/span> <span dir=\"ltr\" role=\"presentation\">state of support, resources, and tools for Galician makes it one of the European<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">languages in danger of being left behind in the future.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Jos\u00e9 Manuel Ram\u00edrez S\u00e1nchez, Laura Docio-Fernandez and Carmen Garcia Mateo<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O2.2<\/strong><br \/>\n15:20\u00a0 &#8211; 15:40<\/td>\n<td>Contextual-Utterance Training for Automatic Speech Recognition (<span class=\"collapseomatic \" id=\"id6a6eaa832d9e5\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832d9e5\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Recent studies of streaming automatic speech recognition (ASR) recurrent neural<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">network transducer (RNN-T)-based systems have fed the encoder with past contex<\/span><span dir=\"ltr\" role=\"presentation\">tual information in order to improve its word error rate (WER) performance. In this<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">paper, we first propose a contextual-utterance training technique which makes use<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">of the previous and future contextual utterances in order to do an implicit adapta-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">tion to the speaker, topic and acoustic environment. Also, we propose a dual-mode<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">contextual-utterance training technique for streaming ASR systems. This proposed<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">approach allows to make a better use of the available acoustic context in streaming<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">models by distilling \u201cin-place\u201d the knowledge of a teacher (non-streaming mode),<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">which is able to see both past and future contextual utterances, to the student<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">(streaming mode) which can only see the current and past contextual utterances.<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">The experimental results show that a state-of-the-art conformer-transducer system<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">trained with the proposed techniques outperforms the same system trained with the<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">classical RNN-T loss.<\/span> <span dir=\"ltr\" role=\"presentation\">Specifically, the proposed technique is able to reduce both<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the WER and the average last token emission latency by more than 6% and 40 ms<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">relative, respectively.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Alejandro Gomez-Alanis, Lukas Drude, Andreas Schwarz, Rupak Vignesh Swaminathan and Simon Wiesler<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O2.3<\/strong><br \/>\n15:40\u00a0 &#8211; 16:00<\/td>\n<td>Phone classification using electromyographic signals\u00a0(<span class=\"collapseomatic \" id=\"id6a6eaa832da33\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832da33\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Silent speech interfaces aim at generating speech from biosignals obtained from<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">the human speech production system. In order to provide resources for the devel<\/span><span dir=\"ltr\" role=\"presentation\">opment of these interfaces, language-specific databases are required. Several silent<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">speech electromyography (EMG) databases for English exist. However, a database <\/span><span dir=\"ltr\" role=\"presentation\">for the Spanish language had yet to be developed.<\/span> <span dir=\"ltr\" role=\"presentation\">The aim of this research is to<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">validate the experimental design of the first silent speech EMG database for Span<\/span><span dir=\"ltr\" role=\"presentation\">ish, namely the new ReSSInt-EMG database.<\/span> <span dir=\"ltr\" role=\"presentation\">The EMG signals in this database<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">are obtained using eight surface EMG bipolar electrode pairs located in the face<\/span> <span dir=\"ltr\" role=\"presentation\">and neck and are recorded in parallel with either audible or silent speech.<\/span> <span dir=\"ltr\" role=\"presentation\">Phone<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">classification experiments are performed, using a set of time-domain features typi<\/span><span dir=\"ltr\" role=\"presentation\">cally used in related works. As a validation reference, the EMG-UKA Trial Corpus<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">is used, which is the most commonly used silent speech EMG database for English. <\/span><span dir=\"ltr\" role=\"presentation\">The results show an average test accuracy of 40.85% for ReSSInt- EMG, suggesting<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">that the data acquisition procedure for the new database is valid.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Eder Del Blanco, Inge Salomons, Eva Navas and Inma Hern\u00e1ez<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O2.4<\/strong><br \/>\n16:00\u00a0 &#8211; 16:20<\/td>\n<td>Semisupervised training of a fully bilingual ASR system for Basque and Spanish (<span class=\"collapseomatic \" id=\"id6a6eaa832da81\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832da81\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Automatic speech recognition (ASR) of speech signals with code-switching (an <\/span><span dir=\"ltr\" role=\"presentation\">abrupt language change common in bilingual communities) typically requires spo<\/span><span dir=\"ltr\" role=\"presentation\">ken language recognition to get single-language segments. In this paper, we present<\/span> <span dir=\"ltr\" role=\"presentation\">a fully bilingual ASR system for Basque and Spanish which does not require such <\/span><span dir=\"ltr\" role=\"presentation\">segmentation but naturally deals with both languages using a single set of acoustic<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">units and a single (aggregated) language model. We also present the Basque Par<\/span><span dir=\"ltr\" role=\"presentation\">liament Database (BPDB) used for the experiments in this work. A semisupervised<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">method is applied, which starts by training baseline acoustic models on small acous<\/span><span dir=\"ltr\" role=\"presentation\">tic datasets in Basque and Spanish. These models are then used to perform phone<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">recognition on the BPDB training set, for which only approximate transcriptions<\/span> <span dir=\"ltr\" role=\"presentation\">are available.<\/span> <span dir=\"ltr\" role=\"presentation\">A similarity score derived from the alignment of the nominal and<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">recognized phonetic sequences is used to rank a set of training segments. Acoustic <\/span><span dir=\"ltr\" role=\"presentation\">models are updated with those BPDB training segments for which the similarity<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">score exceeds a heuristically fixed threshold. Using the updated models, Word Er<\/span><span dir=\"ltr\" role=\"presentation\">ror Rate (WER) reduced from 16.46 to 6.99 on the validation set, and from 15.06<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">to 5.16 on the test set, meaning 57.5% and 65.74% relative WER reductions over<\/span> <span dir=\"ltr\" role=\"presentation\">baseline models, respectively.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Mikel Penagarikano, Amparo Varona, German Bordel and Luis J. Rodriguez-Fuentes<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O2.5<\/strong><br \/>\n16:20\u00a0 &#8211; 16:40<\/td>\n<td>Speaker-Adapted End-to-End Visual Speech Recognition for Continuous Spanish (<span class=\"collapseomatic \" id=\"id6a6eaa832dacf\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832dacf\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">Different studies have shown the importance of visual cues throughout the speech <\/span><span dir=\"ltr\" role=\"presentation\">perception process. In fact, the development of audiovisual approaches has led to <\/span><span dir=\"ltr\" role=\"presentation\">advances in the field of speech technologies.<\/span> <span dir=\"ltr\" role=\"presentation\">However, although noticeable results <\/span><span dir=\"ltr\" role=\"presentation\">have recently been achieved, visual speech recognition remains an open research<\/span> <span dir=\"ltr\" role=\"presentation\">problem.<\/span> <span dir=\"ltr\" role=\"presentation\">It is a task in which, by dispensing with the auditory sense, challenges<\/span> <span dir=\"ltr\" role=\"presentation\">such as visual ambiguities and the complexity of modeling silence must be faced. <\/span><span dir=\"ltr\" role=\"presentation\">Nonetheless, some of these challenges can be alleviated when the problem is ap<\/span><span dir=\"ltr\" role=\"presentation\">proached from a speaker-dependent perspective. Thus, this paper studies, using the<\/span> <span dir=\"ltr\" role=\"presentation\">Spanish LIPRTVE database, how the estimation of specialized end-to-end systems <\/span><span dir=\"ltr\" role=\"presentation\">for a specific person could affect the quality of speech recognition. First, different<\/span> <span dir=\"ltr\" role=\"presentation\">adaptation strategies based on the fine-tuning technique were proposed.<\/span> <span dir=\"ltr\" role=\"presentation\">Then, a<\/span> <span dir=\"ltr\" role=\"presentation\">pre-trained CTC\/Attention architecture was used as a baseline throughout our ex<\/span><span dir=\"ltr\" role=\"presentation\">periments. Our findings showed that a two-step finetuning process, where the VSR<\/span> <span dir=\"ltr\" role=\"presentation\">system is first adapted to the task domain, provided significant improvements when<\/span> <span dir=\"ltr\" role=\"presentation\">the speaker adaptation was addressed. Furthermore, results comparable to the cur<\/span><span dir=\"ltr\" role=\"presentation\">rent state of the art were reached even when only a limited amount of data was <\/span><span dir=\"ltr\" role=\"presentation\">available.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>David Gimeno-Gomez and Carlos David Martinez Hinarejos<\/em><\/td>\n<\/tr>\n<tr>\n<td rowspan=\"2\" width=\"140\"><strong>O2.6<\/strong><br \/>\n16:40\u00a0 &#8211; 17:00<\/td>\n<td>Iterative pseudo-forced alignment by acoustic CTC loss for self-supervised ASR domain adaptation (<span class=\"collapseomatic \" id=\"id6a6eaa832db1d\"  tabindex=\"0\" title=\"abs\"    >abs<\/span><div id=\"target-id6a6eaa832db1d\" class=\"collapseomatic_content \"><span dir=\"ltr\" role=\"presentation\">High-quality data labeling from specific domains is costly and human time-<\/span><span dir=\"ltr\" role=\"presentation\">consuming. In this work, we propose a selfsupervised domain adaptation method,<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">based upon an iterative pseudo-forced alignment algorithm.<\/span> <span dir=\"ltr\" role=\"presentation\">The produced align<\/span><span dir=\"ltr\" role=\"presentation\">ments are employed to customize an end-to-end Automatic Speech Recognition<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">(ASR) and iteratively refined. The algorithm is fed with frame-wise character pos<\/span><span dir=\"ltr\" role=\"presentation\">teriors produced by a seed ASR, trained with out-of-domain data, and optimized<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">throughout a Connectionist Temporal Classification (CTC) loss.<\/span> <span dir=\"ltr\" role=\"presentation\">The alignments<\/span> <span dir=\"ltr\" role=\"presentation\">are computed iteratively upon a corpus of broadcast TV. The process is repeated by<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">reducing the quantity of text to be aligned or expanding the alignment window until <\/span><span dir=\"ltr\" role=\"presentation\">finding the best possible audio-text alignment. The starting timestamps, or tempo-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">ral anchors, are produced uniquely based on the confidence score of the last aligned<\/span> <span dir=\"ltr\" role=\"presentation\">utterance.<\/span> <span dir=\"ltr\" role=\"presentation\">This score is computed with the paths of the CTC-alignment matrix.<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">With this methodology, no human-revised text references are required. Alignments <\/span><span dir=\"ltr\" role=\"presentation\">from long audio files with low-quality transcriptions, like TV captions, are filtered<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">out by confidence score and ready for further ASR adaptation. The obtained results, <\/span><span dir=\"ltr\" role=\"presentation\">on both the Spanish RTVE2022 and CommonVoice databases, underpin the feasibil-<\/span><br role=\"presentation\" \/><span dir=\"ltr\" role=\"presentation\">ity of using CTC-based systems to perform: highly accurate audio-text alignments,<\/span> <span dir=\"ltr\" role=\"presentation\">domain adaptation and semi-supervised training of end-to-end ASR.<\/span><\/div>)<\/td>\n<\/tr>\n<tr>\n<td><em>Fernando L\u00f3pez and Jordi Luque<\/em><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><strong>RTTH Assembly <\/strong><br \/>\n<span style=\"color: #d62013;\"><strong>Monday, 14 November 2022 (17:20-18:50)<\/strong><\/span><\/p>\n<p>&nbsp;<\/p>\n<hr \/>\n","protected":false},"excerpt":{"rendered":"<p>Monday, November 14 Oral 1: Speech Synthesis Monday, 14 November 2022 (9:20-10:40) Chair: Antonio Bonafonte O1.1 9:20\u00a0 &#8211; 09:40 Discrete Acoustic Space for an Efficient Sampling in Neural Text-To-Speech () Marek Strelec, Jonas Rohnke, Antonio Bonafonte, Mateusz Lajszczak, Trevor Wood O1.2 09:40 &#8211; 10:00 An animated realistic head with vocal tract for the finite element&hellip;&nbsp;<a href=\"https:\/\/iberspeech2022.ugr.es\/?page_id=795\" rel=\"bookmark\">Leer m\u00e1s &raquo;<span class=\"screen-reader-text\">Technical Program, Day 1<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"neve_meta_sidebar":"","neve_meta_container":"","neve_meta_enable_content_width":"off","neve_meta_content_width":100,"neve_meta_title_alignment":"","neve_meta_author_avatar":"","neve_post_elements_order":"","neve_meta_disable_header":"","neve_meta_disable_footer":"","neve_meta_disable_title":"","footnotes":""},"class_list":["post-795","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/pages\/795","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=795"}],"version-history":[{"count":9,"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/pages\/795\/revisions"}],"predecessor-version":[{"id":825,"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=\/wp\/v2\/pages\/795\/revisions\/825"}],"wp:attachment":[{"href":"https:\/\/iberspeech2022.ugr.es\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=795"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}