Seeing Beyond Text: A Visual-Linguistic Dataset and Multimodal Framework for English-Hindi Video-Guided Translation
DOI:
https://doi.org/10.56042/jsir.v85i4.22732Keywords:
Artificial intelligence, Cross-lingual learning, Indian languages, Machine learning, Natural language processingAbstract
Despite the progress in Neural Machine Translation (NMT), translating ambiguous and context rich content remains a major challenge, especially in low-resource language pairs like English-Hindi. Traditional NMT systems often fail to solve these challenges due to their reliance on textual data alone. Multimodal approaches, particularly those incorporating visual context, offer promising solutions to the task by resolving linguistic ambiguities. To address this, a novel solution is introduced through a Visual Scene-Aware Hindi Subtitles Dataset (VISA-HIN), designed specifically for English-Hindi Video-Guided Multimodal Machine Translation (VMMT). This dataset aligns English subtitles with corresponding video frames and provides Hindi translations. Alongside the dataset, this study propose a video-guided MMT framework that leverages visual cues to enhance translation quality. The results of the experiments show the potential of scene aware information to improve contextual understanding and fluency in English-to-Hindi translation, paving the way for more robust and accurate multimodal translation systems in low-resource settings.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Scientific & Industrial Research (JSIR)

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.