Publicación: Diseño e implementación de un módulo de reconocimiento de voz para el control de sistemas robóticos
| dc.contributor.advisor | Rivera, Luis Alberto | |
| dc.contributor.author | López Godínez, Oscar Fernando | |
| dc.contributor.jury | Rivera, Luis Alberto | |
| dc.contributor.jury | Esquit Hernández, Carlos Alberto | |
| dc.contributor.jury | Zea Arenales, Miguel Enrique | |
| dc.date.accessioned | 2026-08-05T16:50:56Z | |
| dc.date.issued | 2026 | |
| dc.description | Formato PDF digital — 60 páginas — incluye gráficos, tablas y referencias bibliográficas. | |
| dc.description.abstract | Este trabajo presenta el diseño, la implementación y la validación de un módulo de reconocimiento de voz en español para el control de sistemas robóticos. El objetivo principal es reducir la carga operativa relacionada al uso de las manos. El sistema se estructura de cuatro componentes. El primero transcribe en tiempo real la voz del usuario a texto legible mediante un servicio speech-to-text en modalidad de streaming . El segundo, un ánalizador sintáctico-semántico, toma las expresiones transcritas y las traduce a intenciones entendibles para el sistema robótico. El tercero es un controlador capaz de ejecutar órdenes sobre el robot que considere los límites articulares y los comandos con instrucciones preestablecidas. El cuarto es una interfaz gráfica que permite visualizar las transcripciones, las acciones interpretadas, el estado y los ángulos del robot. Para la validación experimental de este proyecto, se utilizó un myCobot280 M5 y un micrófono JBL Quantum, con 8 sujetos de prueba. El sistema alcanzó un desempeño de comandos interpretados y ejecutados de 96.25 % en modo absoluto y 95 % en modo relativo, lo que demuestra su viabilidad operativa. | spa |
| dc.description.abstract | This work presents the design, implementation, and validation of a Spanish voice recognition module for the control of robotic systems. The main objective is to reduce the operational burden associated with the use of hands when manipulating robotic systems. The system is structured into four components. The first component transcribes the user’s speech into readable text in real time through a streaming speech-to-text service. The second component is a syntacticsemantic analyzer that processes the transcribed expressions and translates them into system-understandable robotic commands. The third component is a controller capable of executing robot commands while enforcing joint limits and predefined instruction constraints. The fourth component is a graphical user interface that displays live transcriptions, interpreted actions, system status, and robot joint angles. For experimental validation, a myCobot280 M5 manipulator and a JBL Quantum microphone were used, involving eight test subjects. The system achieved a correctly interpreted and executed command rate of 96.25% in absolute mode and 95% in relative mode, demonstrating operational feasibility. | |
| dc.description.degreelevel | Pregrado | |
| dc.description.degreename | Licenciado en Ingeniería Mecatrónica | |
| dc.format.extent | 60 p. | |
| dc.format.mimetype | application/pdf | |
| dc.identifier.uri | https://repositorio.uvg.edu.gt/handle/123456789/6783 | |
| dc.language.iso | spa | |
| dc.publisher | Universidad del Valle de Guatemala | |
| dc.publisher.branch | Campus Central | |
| dc.publisher.faculty | Facultad de Ingeniería | |
| dc.publisher.place | Guatemala | |
| dc.publisher.program | Licenciatura en Ingeniería Mecatrónica | |
| dc.relation.references | S. Boch, «Optimización de la herramienta de procesamiento de imágenes para el sistema Brainlab de Humana, Fase IV,» Trabajo de graduación de licenciatura, Universidad del Valle de Guatemala, 2024. | |
| dc.relation.references | R. Hernández, «Facultad de ciencias de la computación,» Trabajo de graduación de licenciatura, Benémerita Universidad Autónoma de Puebla, ago. de 2018. Dirección: https://hdl.handle.net/20.500.12371/7914. | |
| dc.relation.references | J. Piza, «Parser semántico basado en redes neuronales de tipo Deep Q-Network,» ago. de 2024. doi : https://doi.org/10.48713/10336_43268. | |
| dc.relation.references | E. Verbit, «Reconocimiento automático de voz (ASR),» sep. de 2023. Dirección: https: //verbit-ai.translate.goog/transcription/automatic-speech-recognition-asr/?_x_tr_ sl=en&_x_tr_tl=es&_x_tr_hl=es&_x_tr_p. | |
| dc.relation.references | J. Banesty, «Springer Handbook of Speech,» 2008. doi : https://doi.org/10.1007/978- 3-540-49127-9. | |
| dc.relation.references | R. Maher, A Tutorial on Acoustical Transducers: Microphones and Loudspeakers . Dirección: https://www.montana.edu/rmaher/ee417/transducer_tutorial.pdf. | |
| dc.relation.references | II.5. Micrófonos: tipos y manejo , es, II.5, Circular técnica, Serie CIRCULARES. Documento descargable sin metadatos de fecha ni autor individual., Taller de Alabanza, Área Técnica, n.d. | |
| dc.relation.references | M. Brandstein y D. Ward, eds., Microphone Arrays: Signal Processing Techniques and Applications . Springer, 2001. doi : 10.1007/978-3-662-04619-7. | |
| dc.relation.references | J. Sohn, N. S. Kim y W. Sung, «A Statistical Model-Based Voice Activity Detection,» IEEE Signal Processing Letters , vol. 6, n. o 1, págs. 1-3, 1999. doi : 10.1109/97.736233. Dirección: https : / / www . researchgate . net / publication / 3342424 _ A _ statistical _ model_based_voice_activity_detector. | |
| dc.relation.references | S. B. Davis y P. Mermelstein, «Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences,» IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. ASSP-28, n. o 4, págs. 357-366, 1980. doi : 10.1109/TASSP.1980.1163420. Dirección: https://doi.org/10.1109/TASSP.1980. 1163420. | |
| dc.relation.references | S. Singh, «Impact of Dataset on Acoustic Models for Automatic Speech Recognition,» arXiv , 2022. Dirección: https://arxiv.org/abs/2203.13590. | |
| dc.relation.references | M. Fabien, «Introduction to Automatic Speech Recognition (ASR),» Medium , 2019. Dirección: https://maelfabien.github.io/machinelearning/speech_reco/ | |
| dc.relation.references | A. University, «8.3. Speech Recognition,» Aalto University , 2020. Dirección: https: //speechprocessingbook.aalto.fi/Recognition/Speech_Recognition.html. | |
| dc.relation.references | J. Hui, «Speech Recognition Feature Extraction MFCC & PLP,» Medium , 2019. Dirección: https://jonathan-hui.medium.com/speech-recognition-feature-extraction- mfcc-plp-5455f5a69dd9. | |
| dc.relation.references | J. Hui, «Speech Recognition Feature Extraction MFCC & PLP,» Medium , 2019. Dirección: https://jonathan-hui.medium.com/speech-recognition-feature-extraction- mfcc-plp-5455f5a69dd9. | |
| dc.relation.references | G. Hinton, L. Deng y D. Yu, «Deep Neural Networks for Acoustic Modeling in Speech Recognition,» | |
| dc.relation.references | B. Lutkevich, «Analizador sintáctico,» jul. de 2022. Dirección: https://www-techtarget- com.translate.goog/searchapparchitecture/definition/parser?_x_tr_sl=en&_x_tr_ tl=es&_x_tr_hl=es&_x_tr_p. | |
| dc.relation.references | D. Jurafsky y J. H. Martin, Speech and Language Processing , 3rd ed. (draft). Prentice Hall, 2023, Chs. on semantics, SLU, and NER. | |
| dc.relation.references | M. W. Spong, S. Hutchinson y M. Vidyasagar, Robot Modeling and Control . John Wiley & Sons, 2006. | |
| dc.relation.references | K. M. Lynch y F. C. Park, Modern Robotics: Mechanics, Planning, and Control . Cambridge University Press, 2017. | |
| dc.relation.references | C. Carissoli, L. Negri, M. Bassi, F. A. Storm y A. Delle Fave, «Mental Workload and Human-Robot Interaction in Collaborative Tasks: A Scoping Review,» International Journal of Human-Computer Interaction , vol. 40, n. o 20, págs. 6458-6477, 2024. doi : 10.1080/10447318.2023.2254639. | |
| dc.relation.references | J. Sweller, «Cognitive Load During Problem Solving: Effects on Learning,» Cognitive Science , vol. 12, n. o 2, págs. 257-285, 1988. doi : 10.1207/s15516709cog1202_4. | |
| dc.relation.references | M. A. Goodrich y A. C. Schultz, «Human-Robot Interaction: A Survey,» Foundations and Trends in Human-Computer Interaction , vol. 1, n. o 3, págs. 203-275, 2007. doi : 10.1561/1100000005. | |
| dc.relation.references | C. D. Wickens, «Multiple Resources and Mental Workload,» Human Factors , vol. 50, n. o 3, págs. 449-455, 2008. doi : 10.1518/001872008X288394. | |
| dc.relation.references | C. Murad, H. Candello y C. Munteanu, «What’s The Talk on VUI Guidelines? A Meta-Analysis of Guidelines for Voice User Interface Design,» en Proceedings of the ACM Conference on Conversational User Interfaces , New York, NY, USA: ACM, 2023, págs. 1-15. doi : 10.1145/3571884.3597129. | |
| dc.relation.references | M. H. Cohen, J. P. Giangola y J. Balogh, Voice User Interface Design . Boston, MA: Addison-Wesley Professional, 2004, isbn : 9780321185761. | |
| dc.relation.references | Elephant Robotics. «MyCobot 280 M5.» Recuperado el 28 de septiembre de 2025. Dirección: https://www.elephantrobotics.com/en/mycobot-en. | |
| dc.relation.references | Harman International Industries, Inc. «Quantum Stream Talk Especificaciones técnicas,» Harman International Industries, Inc. Dirección: https : / / www . jbl . com / QUANTUM-STREAM-TALK.html. | |
| dc.relation.references | O. F. López Godínez, Proyecto_Modulo_Reconocimiento_de_voz , Repositorio en GitHub, 2025. Dirección: https://github.com/Olopez12/Proyecto_Modulo_Reconocimiento_ de_voz/blob/main/speech_parser.py. | |
| dc.rights.accessrights | info:eu-repo/semantics/openAccess | |
| dc.rights.coar | http://purl.org/coar/access_right/c_abf2 | |
| dc.rights.license | Atribución-NoComercial-SinDerivadas 4.0 Internacional (CC BY-NC-ND 4.0) | |
| dc.rights.uri | https://creativecommons.org/licenses/by-nc-nd/4.0/ | |
| dc.subject.armarc | Robotics | |
| dc.subject.armarc | Robótica | |
| dc.subject.armarc | Technological innovations | |
| dc.subject.armarc | Innovaciones tecnológicas | |
| dc.subject.armarc | Speech processing systems | |
| dc.subject.armarc | Speech perception -- Guatemala | |
| dc.subject.armarc | Voz-Reconocimiento automático | |
| dc.subject.armarc | Interacción hombre-computador | |
| dc.subject.armarc | Sistemas de procesamiento de la voz | |
| dc.subject.armarc | Human-robot interaction -- Guatemala | |
| dc.subject.armarc | Natural language processing (Computer science) | |
| dc.subject.ddc | 620 - Ingeniería y operaciones afines::629 - Otras ramas de la ingeniería | |
| dc.subject.ocde | 2. Ingeniería y Tecnología | |
| dc.subject.ods | ODS 9: Industria, innovación e infraestructura. Construir infraestructuras resilientes, promover la industrialización inclusiva y sostenible y fomentar la innovación | |
| dc.subject.proposal | MyCobot280 | spa |
| dc.subject.proposal | Control robótico | spa |
| dc.subject.proposal | Reconocimiento de voz | spa |
| dc.subject.proposal | Conversión de voz a texto | spa |
| dc.subject.proposal | Interfaz gráfica de usuario | spa |
| dc.subject.proposal | Interacción humano-robot | spa |
| dc.subject.proposal | Análisis sintáctico-semántico | spa |
| dc.title | Diseño e implementación de un módulo de reconocimiento de voz para el control de sistemas robóticos | spa |
| dc.title.translated | Design and implementation of a speech recognition module for robotic system control | |
| dc.type | Trabajo de grado - Pregrado | |
| dc.type.coar | http://purl.org/coar/resource_type/c_7a1f | |
| dc.type.coarversion | http://purl.org/coar/version/c_970fb48d4fbd8a85 | |
| dc.type.content | Text | |
| dc.type.driver | info:eu-repo/semantics/bachelorThesis | |
| dc.type.version | info:eu-repo/semantics/publishedVersion | |
| dc.type.visibility | Public Thesis | |
| dspace.entity.type | Publication |
