Audio classification in speech and music: A comparison between a statistical and a neural approach
Summary: We focus on the problem of audio classification in speech and music for multimedia applications. In particular, we present a comparison between two different techniques for speech/music discrimination. The first method is based on zero crossing rate and Bayesian classification. It is very simple from a computational point of view and gives good results in case of pure music or speech. The simulation results show that some performance degradation arises when the music segment contains also some speech superimposed on music, or strong rhythmic components. To overcome these problems, we propose a second method that uses more features and that is based on neural networks (specifically a multi-layer perceptron). In this case we obtain better performance, at the expense of a limited growth in the computational complexity. In practice, the proposed neural network is simple to implement if a suitable polynomial is used as the activation function, and a real-time implementation is possible even if low-cost embedded systems are used.
- Identifying the classical music composition of an unknown performance with wavelet dispersion vector and neural nets
- A physiologically inspired method for audio classification
- Real-time Speech and Music Classification by Large Audio Feature Space Extraction
- Classification of general audio data for content-based retrieval
- Online speech/music segmentation based on the variance mean of filter bank energy
This page was built for publication: Audio classification in speech and music: A comparison between a statistical and a neural approach
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1607684)