In today's increasingly global and digital learning environments, language barriers remain a critical challenge to effective communication and accessibility. This competition focuses on building a real-time multilingual classroom translation system that enables seamless interaction between instructors and students across diverse linguistic backgrounds.
Participants are tasked with developing an end-to-end pipeline that converts spoken English lectures into multilingual outputs, including both translated subtitles and synthesized speech in the listener's native language. The system must operate under strict real-time constraints, ensuring minimal latency while maintaining high accuracy and contextual integrity. This problem involves NLP tasks such as combining automatic speech recognition (ASR), machine translation (MT), and text-to-speech (TTS) into a unified, streaming-based architecture. Solutions should be robust to variations in accents, background noise, and domain-specific terminology commonly found in educational settings.
The competition emphasizes: - Low-latency processing for real-time usability - High-quality, context-aware translation - Natural and synchronized speech synthesis - Scalability across multiple Indian and international languages
Top-performing participants may be considered for internships and collaboration opportunities on INICAI-driven projects.
Key Challenges
Evaluation
Submissions are evaluated using a composite metric that measures transcription accuracy, translation quality, speech synthesis quality, and end-to-end latency. The objective is to balance linguistic correctness with real-time system performance. The final score considers Word Error Rate (WER) for transcription, BLEU/COMET/BERTScore for translation quality, automated perceptual metrics for speech synthesis naturalness, and normalized latency scores.