AI-Powered Multilingual Content Localization Engine for  Skill Courses

Main Article Content

Pankaj Vaishnav

Abstract

This paper presents an AI-powered multilingual content localization engine designed to bridge the language accessibility gap in India's vocational and skilling ecosystem. The proposed system integrates multiple AI technologies including Automatic Speech Recognition (ASR) using Whisper, Neural Machine Translation (NMT) using MarianMT, domain- specific glossary alignment via Retrieval-Augmented Generation (RAG), and Text-to-Speech (TTS) synthesis using Coqui TTS into a unified pipeline for both text-based course content localization and video speech dubbing. The engine targets all 22 scheduled Indian languages and incorporates a modular microservice architecture with LMS integration through FastAPI REST endpoints. Key components include a vocal isolation layer using Spleeter, structured PDF and text parsing, structured assessment localization, and a feedback-driven quality improvement loop. The framework addresses challenges of domain terminology misalignment, regional voice naturalness, audio-video synchronization, and scalability for national-level deployment. Evaluation is designed around BLEU scores for translation, Word Error Rate for transcription, and Mean Opinion Score for TTS naturalness. This work contributes a feasible, cost-efficient, and scalable architecture for automated multilingual localization of vocational training material with significant implications for learner inclusion and certification success across India's diverse linguistic communities

Article Details

Section

Articles

Author Biography

Pankaj Vaishnav

Department of Computer Science and Engineering Geetanjali Institute of Technical Studies (GITS) Dabok, Udaipur (Raj.), India

References

[1] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, "Attention is all you need," in Advances in Neural Information Processing Systems, vol. 30, 2017.

[2] J. Tiedemann and S. Thottingal, "OPUS-MT - Building open translation services for the World," in Proc. 22nd Annual Conference of the European Association for Machine Translation, 2020, pp. 479-480

.

[3] A. Kunchukuttan, P. Mehta, and P. Bhattacharyya, "The IIT Bombay English-Hindi Parallel Corpus," in Proc. 11th International Conference on Language Resources and Evaluation, 2018, pp. 2828-2835.

[4] Tiwari, K., Patel, M. (2020). Facial Expression Recognition Using Random Forest Classifier. In: Mathur, G., Sharma, H., Bundele, M., Dey, N., Paprzycki, M. (eds) International Conference on Artificial Intelligence: Advances and Applications 2019. Algorithms for Intelligent Systems. Springer, Singapore. https://doi.org/10.1007/978-981-15-1059-5_15.

[5] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, "Robust speech recognition via large-scale weak supervision," in Proc. 40th International Conference on Machine Learning, 2023.

[6] V. Peddinti, G. Chen, V. Manohar, T. Ko, D. Povey, and S. Khudanpur, "JHU ASpIRE system: Robust LVCSR with time delay neural networks and IVECTOR adaptation," in Proc. IEEE ASRU Workshop, 2015, pp. 539-546.

[7] Shekhawat, V.S., Tiwari, M., Patel, M. (2021). A Secured Steganography Algorithm for Hiding an Image and Data in an Image Using LSB Technique. In: Singh, V., Asari, V.K., Kumar, S., Patel, R.B. (eds) Computational Methods and Data Engineering. Advances in Intelligent Systems and Computing, vol 1257. Springer, Singapore. https://doi.org/10.1007/978-981-15-7907-3_35

[8] Patel, Mayank, Neelam Badi, and Amit Sinhal. "The role of fuzzy logic in improving accuracy of phishing detection system." International Journal of Innovative Technology and Exploring Engineering (2019): 3162-3164.

[9] E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Golge, and M. A. Ponti, "YourTTS: Towards zero-shot multi-speaker TTS and zero- shot voice conversion for everyone," in Proc. 39th International Conference on Machine Learning, 2022.

[10] A. Baby, A. P. Thomas, N. Nishanthi, and H. A. Murthy, "Resources for Indian languages," in Proc. 1st Workshop on Technologies for MT of Low Resource Languages, 2016, pp. 46-54.

[11] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Kuttler, M. Lewis, W. Yih, T. Rocktaschel, S. Riedel, and D. Kiela, "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Advances in Neural Information Processing Systems, vol. 33, 2020,

pp. 9459-9474.