Multimodal speech recognition combines acoustic, visual and other sensory inputs to transcribe spoken language more reliably than audio-only approaches. Early systems employed statistical models such ...