Regionalized models for Spanish language variations based on Twitter

Eric S Tellez; Daniela Moctezuma; Sabino Miranda; Mario Graff; Guillermo Ruiz

doi:10.1007/s10579-023-09640-9

Regionalized models for Spanish language variations based on Twitter

Lang Resour Eval. 2023 Mar 2:1-31. doi: 10.1007/s10579-023-09640-9. Online ahead of print.

Authors

Eric S Tellez^{1

2

3}, Daniela Moctezuma^#⁴, Sabino Miranda^#^{1

2

5}, Mario Graff^#^{1

2}, Guillermo Ruiz^#^{1

4}

Affiliations

¹ Conacyt, Consejo Nacional de Ciencia y Tecnología., Av. Insurgentes Sur 1582, Col. Crédito Constructor., 03940 CDMX, Mexico.
² INFOTEC, Centro de Investigación e Innovación en Tecnologías de la Información y Comunicación, Circuito Tecnopolo Norte, No.112 Col. Tecnopolo Pocitos II, 20326 Aguascalientes, Aguascalientes Mexico.
³ CICESE, Centro de Investigación Científica y de Educación Superior de Ensenada, Carr. Tijuana-Ensenada, No.3918, Zona Playitas, 22860 Ensenada, Baja California Mexico.
⁴ CentroGEO, Centro de Investigación en Ciencias de Información Geoespacial., Circuito Tecnopolo Norte, No.107 Col. Tecnopolo Pocitos II, 20313 Aguascalientes, Aguascalientes Mexico.
⁵ UPIITA-IPN, Instituto Politécnico Nacional, Av. Instituto Politécnico Nacional 2580 Col. Barrio la Laguna Ticomàn, Gustavo A. Madero, 07360 Mexico City, Mexico.

^# Contributed equally.

Abstract

Spanish is one of the most spoken languages in the world. Its proliferation comes with variations in written and spoken communication among different regions. Understanding language variations can help improve model performances on regional tasks, such as those involving figurative language and local context information. This manuscript presents and describes a set of regionalized resources for the Spanish language built on 4-year Twitter public messages geotagged in 26 Spanish-speaking countries. We introduce word embeddings based on FastText, language models based on BERT, and per-region sample corpora. We also provide a broad comparison among regions covering lexical and semantical similarities and examples of using regional resources on message classification tasks.

Keywords: Linguistic resources; Semantic space; Spanish Twitter.

© The Author(s), under exclusive licence to Springer Nature B.V. 2023, Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.