Seeing Beyond Text: A Visual-Linguistic Dataset and Multimodal Framework for English-Hindi Video-Guided Translation

Authors

  • Binnu Paul Department of CSE, NIT Agartala 799 046, India
  • Dwijen Rudrapal Department of CSE, NIT Agartala 799 046, India
  • Kunal Chakma Department of CSE, NIT Agartala 799 046, India
  • Anupam Jamatia Department of CSE, NIT Agartala 799 046, India

DOI:

https://doi.org/10.56042/jsir.v85i4.22732

Keywords:

Artificial intelligence, Cross-lingual learning, Indian languages, Machine learning, Natural language processing

Abstract

Despite the progress in Neural Machine Translation (NMT), translating ambiguous and context rich content remains a major challenge, especially in low-resource language pairs like English-Hindi. Traditional NMT systems often fail to solve these challenges due to their reliance on textual data alone. Multimodal approaches, particularly those incorporating visual context, offer promising solutions to the task by resolving linguistic ambiguities. To address this, a novel solution is introduced through a Visual Scene-Aware Hindi Subtitles Dataset (VISA-HIN), designed specifically for English-Hindi Video-Guided Multimodal Machine Translation (VMMT). This dataset aligns English subtitles with corresponding video frames and provides Hindi translations. Alongside the dataset, this study propose a video-guided MMT framework that leverages visual cues to enhance translation quality. The results of the experiments show the potential of scene aware information to improve contextual understanding and fluency in English-to-Hindi translation, paving the way for more robust and accurate multimodal translation systems in low-resource settings.

Author Biographies

  • Dwijen Rudrapal, Department of CSE, NIT Agartala 799 046, India

    Dwijen Rudrapal is presently working as an Assistant Professor at National Institute of Technology Agartala, India. He obtained his Ph.D. from Department of Computer Science and Engineering, National Institute of Technology Agartala. His research interests are human language processing, social media text and artificial intelligence. He is a reviewer of reputed journals like Natural Language Engineering, BioMed Research International and premier international conferences like ACL, RANLP, ICON, CiCLing.

  • Kunal Chakma, Department of CSE, NIT Agartala 799 046, India

    Kunal Chakma is a NLP researcher and currently working as an Assistant Professor at National Institute of Technology Agartala, India. He obtained his Ph.D. from Department of Computer Science and Engineering, National Institute of Technology Agartala. His research interests are information retrieval, code-mixed social media text and artificial intelligence. He is a reviewer of reputed journals like Natural Language Engineering and reputed international conferences like ACL, ICON, RANLP, FIRE, CiCLing.

  • Anupam Jamatia, Department of CSE, NIT Agartala 799 046, India

    Anupam Jamatia is a NLP researcher who is presently working as an Assistant Professor in the Department of Computer Science and Engineering at National Institute of Technology, Agartala, Tripura, India. His research interests span all aspects of Natural Language Processing, Computational Linguistis, Language Technology, Text Processing, Computational Social Sciences. He is a reviewer of reputed journals like Natural Language Engineering, Journal of Electrical and Computer Engineering, Journal of Intelligent Systems etc and reputed international conferences like ACL, ICON, RANLP, FIRE, CiCLing.

Downloads

Published

29.07.2026

Issue

Section

Computer Sciences, Communication and Information Technology

How to Cite

Seeing Beyond Text: A Visual-Linguistic Dataset and Multimodal Framework for English-Hindi Video-Guided Translation. (2026). Journal of Scientific & Industrial Research (JSIR), 85(4), 312-324. https://doi.org/10.56042/jsir.v85i4.22732

Similar Articles

1-10 of 197

You may also start an advanced similarity search for this article.