phoneme (Q6034067)

OpenML dataset with id 1489

Language	Label	Description	Also known as
English	phoneme	OpenML dataset with id 1489

Statements

instance of

data set

0 references

dataset version identifier

1

0 references

description

**Author**: Dominique Van Cappel, THOMSON-SINTRA \N**Source**: [KEEL](http://sci2s.ugr.es/keel/dataset.php?cod=105#sub2), [ELENA](https://www.elen.ucl.ac.be/neural-nets/Research/Projects/ELENA/databases/REAL/phoneme/) - 1993 \N**Please cite**: None \N\NThe aim of this dataset is to distinguish between nasal (class 0) and oral sounds (class 1). Five different attributes were chosen to characterize each vowel: they are the amplitudes of the five first harmonics AHi, normalised by the total energy Ene (integrated on all the frequencies): AHi/Ene. The phonemes are transcribed as follows: sh as in she, dcl as in dark, iy as the vowel in she, aa as the vowel in dark, and ao as the first vowel in water. \N\N### Source\N\NThe current dataset was formatted by the KEEL repository, but originally hosted by the [ELENA Project](https://www.elen.ucl.ac.be/neural-nets/Research/Projects/ELENA/elena.htm#stuff). The dataset originates from the European ESPRIT 5516 project: ROARS. The aim of this project was the development and the implementation of a real time analytical system for French and Spanish speech recognition. \N\N### Relevant information\N\NMost of the already existing speech recognition systems are global systems (typically Hidden Markov Models and Time Delay Neural Networks) which recognizes signals and do not really use the speech\Nspecificities. On the contrary, analytical systems take into account the articulatory process leading to the different phonemes of a given language, the idea being to deduce the presence of each of the\Nphonetic features from the acoustic observation.\N\NThe main difficulty of analytical systems is to obtain acoustical parameters sufficiantly reliable. These acoustical measurements must :\N\N - contain all the information relative to the concerned phonetic feature.\N - being speaker independent.\N - being context independent.\N - being more or less robust to noise.\N\NThe primary acoustical observation is always voluminous (spectrum x N different observation moments) and classification cannot been processed directly.\N\NIn ROARS, the initial database is provided by cochlear spectra, which may be seen as the output of a filters bank having a constant DeltaF/F0, where the central frequencies are distributed on a\Nlogarithmic scale (MEL type) to simulate the frequency answer of the auditory nerves. The filters outputs are taken every 2 or 8 msec (integration on 4 or 16 msec) depending on the type of phoneme\Nobserved (stationary or transitory). \N\NThe aim of the present database is to distinguish between nasal and\Noral vowels. There are thus two different classes:\N\N- Class 0 : Nasals \N- Class 1 : Orals \N\NThis database contains vowels coming from 1809 isolated syllables (for example: pa, ta, pan,...). Five different attributes were chosen to characterize each vowel: they are the amplitudes of the five first harmonics AHi, normalised by the total energy Ene (integrated on all the frequencies): AHi/Ene. Each harmonic is signed: positive when it corresponds to a local maximum of the spectrum and negative otherwise.\N\NThree observation moments have been kept for each vowel to obtain 5427 different instances: \N\N - the observation corresponding to the maximum total energy Ene. \N \N - the observations taken 8 msec before and 8 msec after the observation corresponding to this maximum total energy.\N\NFrom these 5427 initial values, 23 instances for which the amplitude of the 5 first harmonics was zero were removed, leading to the 5404 instances of the present database. The patterns are presented in a random order.\N\N### Past Usage \N\NAlinat, P., Periodic Progress Report 4, ROARS Project ESPRIT II- Number 5516, February 1993, Thomson report TS. ASM 93/S/EGS/NC/079 \N \NGuerin-Dugue, A. and others, Deliverable R3-B4-P - Task B4: Benchmarks, Technical report, Elena-NervesII "Enhanced Learning for Evolutive Neural Architecture", ESPRIT-Basic Research Project Number 6891, June 1995 \N\NVerleysen, M. and Voz, J.L. and Thissen, P. and Legat, J.D., A statistical Neural Network for high-dimensional vector classification, ICNN'95 - IEEE International Conference on Neural Networks, November 1995, Perth, Western Australia. \N \NVoz J.L., Verleysen M., Thissen P. and Legat J.D., Suboptimal Bayesian classification by vector quantization with small clusters. ESANN95-European Symposium on Artificial Neural Networks, April 1995, M. Verleysen editor, D facto publications, Brussels, Belgium. \N \NVoz J.L., Verleysen M., Thissen P. and Legat J.D., A practical view of suboptimal Bayesian classification, IWANN95-Proceedings of the International Workshop on Artificial Neural Networks, June 1995, Mira, Cabestany, Prieto editors, Springer-Verlag Lecture Notes in Computer Sciences, Malaga, Spain

0 references

upload date

25 May 2015

0 references

full work available at URL

https://api.openml.org/data/v1/download/1592281/phoneme.arff

0 references

https://sci2s.ugr.es/keel/dataset.php?cod=105#sub2

0 references

default target attribute

Class

0 references