Publicación:
Diseño e implementación de un módulo de reconocimiento de voz para el control de sistemas robóticos

dc.contributor.advisorRivera, Luis Alberto
dc.contributor.authorLópez Godínez, Oscar Fernando
dc.contributor.juryRivera, Luis Alberto
dc.contributor.juryEsquit Hernández, Carlos Alberto
dc.contributor.juryZea Arenales, Miguel Enrique
dc.date.accessioned2026-08-05T16:50:56Z
dc.date.issued2026
dc.descriptionFormato PDF digital — 60 páginas — incluye gráficos, tablas y referencias bibliográficas.
dc.description.abstractEste trabajo presenta el diseño, la implementación y la validación de un módulo de reconocimiento de voz en español para el control de sistemas robóticos. El objetivo principal es reducir la carga operativa relacionada al uso de las manos. El sistema se estructura de cuatro componentes. El primero transcribe en tiempo real la voz del usuario a texto legible mediante un servicio speech-to-text en modalidad de streaming . El segundo, un ánalizador sintáctico-semántico, toma las expresiones transcritas y las traduce a intenciones entendibles para el sistema robótico. El tercero es un controlador capaz de ejecutar órdenes sobre el robot que considere los límites articulares y los comandos con instrucciones preestablecidas. El cuarto es una interfaz gráfica que permite visualizar las transcripciones, las acciones interpretadas, el estado y los ángulos del robot. Para la validación experimental de este proyecto, se utilizó un myCobot280 M5 y un micrófono JBL Quantum, con 8 sujetos de prueba. El sistema alcanzó un desempeño de comandos interpretados y ejecutados de 96.25 % en modo absoluto y 95 % en modo relativo, lo que demuestra su viabilidad operativa.spa
dc.description.abstractThis work presents the design, implementation, and validation of a Spanish voice recognition module for the control of robotic systems. The main objective is to reduce the operational burden associated with the use of hands when manipulating robotic systems. The system is structured into four components. The first component transcribes the user’s speech into readable text in real time through a streaming speech-to-text service. The second component is a syntacticsemantic analyzer that processes the transcribed expressions and translates them into system-understandable robotic commands. The third component is a controller capable of executing robot commands while enforcing joint limits and predefined instruction constraints. The fourth component is a graphical user interface that displays live transcriptions, interpreted actions, system status, and robot joint angles. For experimental validation, a myCobot280 M5 manipulator and a JBL Quantum microphone were used, involving eight test subjects. The system achieved a correctly interpreted and executed command rate of 96.25% in absolute mode and 95% in relative mode, demonstrating operational feasibility.
dc.description.degreelevelPregrado
dc.description.degreenameLicenciado en Ingeniería Mecatrónica
dc.format.extent60 p.
dc.format.mimetypeapplication/pdf
dc.identifier.urihttps://repositorio.uvg.edu.gt/handle/123456789/6783
dc.language.isospa
dc.publisherUniversidad del Valle de Guatemala
dc.publisher.branchCampus Central
dc.publisher.facultyFacultad de Ingeniería
dc.publisher.placeGuatemala
dc.publisher.programLicenciatura en Ingeniería Mecatrónica
dc.relation.referencesS. Boch, «Optimización de la herramienta de procesamiento de imágenes para el sistema Brainlab de Humana, Fase IV,» Trabajo de graduación de licenciatura, Universidad del Valle de Guatemala, 2024.
dc.relation.referencesR. Hernández, «Facultad de ciencias de la computación,» Trabajo de graduación de licenciatura, Benémerita Universidad Autónoma de Puebla, ago. de 2018. Dirección: https://hdl.handle.net/20.500.12371/7914.
dc.relation.referencesJ. Piza, «Parser semántico basado en redes neuronales de tipo Deep Q-Network,» ago. de 2024. doi : https://doi.org/10.48713/10336_43268.
dc.relation.referencesE. Verbit, «Reconocimiento automático de voz (ASR),» sep. de 2023. Dirección: https: //verbit-ai.translate.goog/transcription/automatic-speech-recognition-asr/?_x_tr_ sl=en&_x_tr_tl=es&_x_tr_hl=es&_x_tr_p.
dc.relation.referencesJ. Banesty, «Springer Handbook of Speech,» 2008. doi : https://doi.org/10.1007/978- 3-540-49127-9.
dc.relation.referencesR. Maher, A Tutorial on Acoustical Transducers: Microphones and Loudspeakers . Dirección: https://www.montana.edu/rmaher/ee417/transducer_tutorial.pdf.
dc.relation.referencesII.5. Micrófonos: tipos y manejo , es, II.5, Circular técnica, Serie CIRCULARES. Documento descargable sin metadatos de fecha ni autor individual., Taller de Alabanza, Área Técnica, n.d.
dc.relation.referencesM. Brandstein y D. Ward, eds., Microphone Arrays: Signal Processing Techniques and Applications . Springer, 2001. doi : 10.1007/978-3-662-04619-7.
dc.relation.referencesJ. Sohn, N. S. Kim y W. Sung, «A Statistical Model-Based Voice Activity Detection,» IEEE Signal Processing Letters , vol. 6, n. o 1, págs. 1-3, 1999. doi : 10.1109/97.736233. Dirección: https : / / www . researchgate . net / publication / 3342424 _ A _ statistical _ model_based_voice_activity_detector.
dc.relation.referencesS. B. Davis y P. Mermelstein, «Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences,» IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. ASSP-28, n. o 4, págs. 357-366, 1980. doi : 10.1109/TASSP.1980.1163420. Dirección: https://doi.org/10.1109/TASSP.1980. 1163420.
dc.relation.referencesS. Singh, «Impact of Dataset on Acoustic Models for Automatic Speech Recognition,» arXiv , 2022. Dirección: https://arxiv.org/abs/2203.13590.
dc.relation.referencesM. Fabien, «Introduction to Automatic Speech Recognition (ASR),» Medium , 2019. Dirección: https://maelfabien.github.io/machinelearning/speech_reco/
dc.relation.referencesA. University, «8.3. Speech Recognition,» Aalto University , 2020. Dirección: https: //speechprocessingbook.aalto.fi/Recognition/Speech_Recognition.html.
dc.relation.referencesJ. Hui, «Speech Recognition Feature Extraction MFCC & PLP,» Medium , 2019. Dirección: https://jonathan-hui.medium.com/speech-recognition-feature-extraction- mfcc-plp-5455f5a69dd9.
dc.relation.referencesJ. Hui, «Speech Recognition Feature Extraction MFCC & PLP,» Medium , 2019. Dirección: https://jonathan-hui.medium.com/speech-recognition-feature-extraction- mfcc-plp-5455f5a69dd9.
dc.relation.referencesG. Hinton, L. Deng y D. Yu, «Deep Neural Networks for Acoustic Modeling in Speech Recognition,»
dc.relation.referencesB. Lutkevich, «Analizador sintáctico,» jul. de 2022. Dirección: https://www-techtarget- com.translate.goog/searchapparchitecture/definition/parser?_x_tr_sl=en&_x_tr_ tl=es&_x_tr_hl=es&_x_tr_p.
dc.relation.referencesD. Jurafsky y J. H. Martin, Speech and Language Processing , 3rd ed. (draft). Prentice Hall, 2023, Chs. on semantics, SLU, and NER.
dc.relation.referencesM. W. Spong, S. Hutchinson y M. Vidyasagar, Robot Modeling and Control . John Wiley & Sons, 2006.
dc.relation.referencesK. M. Lynch y F. C. Park, Modern Robotics: Mechanics, Planning, and Control . Cambridge University Press, 2017.
dc.relation.referencesC. Carissoli, L. Negri, M. Bassi, F. A. Storm y A. Delle Fave, «Mental Workload and Human-Robot Interaction in Collaborative Tasks: A Scoping Review,» International Journal of Human-Computer Interaction , vol. 40, n. o 20, págs. 6458-6477, 2024. doi : 10.1080/10447318.2023.2254639.
dc.relation.referencesJ. Sweller, «Cognitive Load During Problem Solving: Effects on Learning,» Cognitive Science , vol. 12, n. o 2, págs. 257-285, 1988. doi : 10.1207/s15516709cog1202_4.
dc.relation.referencesM. A. Goodrich y A. C. Schultz, «Human-Robot Interaction: A Survey,» Foundations and Trends in Human-Computer Interaction , vol. 1, n. o 3, págs. 203-275, 2007. doi : 10.1561/1100000005.
dc.relation.referencesC. D. Wickens, «Multiple Resources and Mental Workload,» Human Factors , vol. 50, n. o 3, págs. 449-455, 2008. doi : 10.1518/001872008X288394.
dc.relation.referencesC. Murad, H. Candello y C. Munteanu, «What’s The Talk on VUI Guidelines? A Meta-Analysis of Guidelines for Voice User Interface Design,» en Proceedings of the ACM Conference on Conversational User Interfaces , New York, NY, USA: ACM, 2023, págs. 1-15. doi : 10.1145/3571884.3597129.
dc.relation.referencesM. H. Cohen, J. P. Giangola y J. Balogh, Voice User Interface Design . Boston, MA: Addison-Wesley Professional, 2004, isbn : 9780321185761.
dc.relation.referencesElephant Robotics. «MyCobot 280 M5.» Recuperado el 28 de septiembre de 2025. Dirección: https://www.elephantrobotics.com/en/mycobot-en.
dc.relation.referencesHarman International Industries, Inc. «Quantum Stream Talk Especificaciones técnicas,» Harman International Industries, Inc. Dirección: https : / / www . jbl . com / QUANTUM-STREAM-TALK.html.
dc.relation.referencesO. F. López Godínez, Proyecto_Modulo_Reconocimiento_de_voz , Repositorio en GitHub, 2025. Dirección: https://github.com/Olopez12/Proyecto_Modulo_Reconocimiento_ de_voz/blob/main/speech_parser.py.
dc.rights.accessrightsinfo:eu-repo/semantics/openAccess
dc.rights.coarhttp://purl.org/coar/access_right/c_abf2
dc.rights.licenseAtribución-NoComercial-SinDerivadas 4.0 Internacional (CC BY-NC-ND 4.0)
dc.rights.urihttps://creativecommons.org/licenses/by-nc-nd/4.0/
dc.subject.armarcRobotics
dc.subject.armarcRobótica
dc.subject.armarcTechnological innovations
dc.subject.armarcInnovaciones tecnológicas
dc.subject.armarcSpeech processing systems
dc.subject.armarcSpeech perception -- Guatemala
dc.subject.armarcVoz-Reconocimiento automático
dc.subject.armarcInteracción hombre-computador
dc.subject.armarcSistemas de procesamiento de la voz
dc.subject.armarcHuman-robot interaction -- Guatemala
dc.subject.armarcNatural language processing (Computer science)
dc.subject.ddc620 - Ingeniería y operaciones afines::629 - Otras ramas de la ingeniería
dc.subject.ocde2. Ingeniería y Tecnología
dc.subject.odsODS 9: Industria, innovación e infraestructura. Construir infraestructuras resilientes, promover la industrialización inclusiva y sostenible y fomentar la innovación
dc.subject.proposalMyCobot280spa
dc.subject.proposalControl robóticospa
dc.subject.proposalReconocimiento de vozspa
dc.subject.proposalConversión de voz a textospa
dc.subject.proposalInterfaz gráfica de usuariospa
dc.subject.proposalInteracción humano-robotspa
dc.subject.proposalAnálisis sintáctico-semánticospa
dc.titleDiseño e implementación de un módulo de reconocimiento de voz para el control de sistemas robóticosspa
dc.title.translatedDesign and implementation of a speech recognition module for robotic system control
dc.typeTrabajo de grado - Pregrado
dc.type.coarhttp://purl.org/coar/resource_type/c_7a1f
dc.type.coarversionhttp://purl.org/coar/version/c_970fb48d4fbd8a85
dc.type.contentText
dc.type.driverinfo:eu-repo/semantics/bachelorThesis
dc.type.versioninfo:eu-repo/semantics/publishedVersion
dc.type.visibilityPublic Thesis
dspace.entity.typePublication

Archivos

Bloque original

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
Oscar Fernando López Godínez.pdf
Tamaño:
5.54 MB
Formato:
Adobe Portable Document Format

Bloque de licencias

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
license.txt
Tamaño:
14.49 KB
Formato:
Item-specific license agreed upon to submission
Descripción: