Dynamic Bayesian networks for audio-visual speech recognition
Summary: The use of visual features in audio-visual speech recognition (AVSR) is justified by both the speech generation mechanism, which is essentially bimodal in audio and visual representation, and by the need for features that are invariant to acoustic noise perturbation. As a result, current AVSR systems demonstrate significant accuracy improvements in environments affected by acoustic noise. In this paper, we describe the use of two statistical models for audio-visual integration, the coupled HMM (CHMM) and the factorial HMM (FHMM), and compare the performance of these models with the existing models used in speaker dependent audio-visual isolated word recognition. The statistical properties of both the CHMM and FHMM allow to model the state asynchrony of the audio and visual observation sequences while preserving their natural correlation over time. In our experiments, the CHMM performs best overall, outperforming all the existing models and the FHMM.
- scientific article; zbMATH DE number 1717105 (Why is no real title available?)
- scientific article; zbMATH DE number 1717109 (Why is no real title available?)
- scientific article; zbMATH DE number 2089701 (Why is no real title available?)
- scientific article; zbMATH DE number 2089703 (Why is no real title available?)
- Use of missing and unreliable data for audiovisual speech recognition
- A primer on coupled state-switching models for multiple interacting time series
- scientific article; zbMATH DE number 2013336 (Why is no real title available?)
- scientific article; zbMATH DE number 1927208 (Why is no real title available?)
- scientific article; zbMATH DE number 2087273 (Why is no real title available?)
- Multi Channel Sequence Processing
- Computer Vision - ECCV 2004
- A coupled HMM approach to video-realistic speech animation
This page was built for publication: Dynamic Bayesian networks for audio-visual speech recognition
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1424537)