Abstract
This article presents the original results of Polish language statistical analysis, based on the orthographic and phonemic language corpus. Phonemic language corpus for Polish was developed by using automatic grapheme-to-phoneme conversion of the source orthographic language corpus, obtained from the National Corpus of Polish (NCP). The corpus contains the most frequently used Polish words, written with the use of phonemic notation. Performed statistical analysis of Polish language based on phonemic language corpus, includes frequency of occurrence calculation of the orthographic and phonemic language components, as well as their sequence. Statistical language data, obtained as a result of performed statistical analysis, enable to develop statistical word-based and phoneme-based language models for Polish. Applying these language models can effectively contribute to efficiency improvement of automatic speech recognition for Polish.
| Original language | English |
|---|---|
| Article number | 5 |
| Journal | Eurasip Journal on Audio, Speech, and Music Processing |
| Volume | 2017 |
| Issue number | 1 |
| DOIs | |
| Publication status | Published - 1 Dec 2017 |
Keywords
- Automatic grapheme-to-phoneme conversion
- Automatic speech recognition
- Language corpus
- Language modelling
- Language statistical analysis
ASJC Scopus subject areas
- Acoustics and Ultrasonics
- Electrical and Electronic Engineering
Fingerprint
Dive into the research topics of 'Statistical analysis of orthographic and phonemic language corpus for word-based and phoneme-based Polish language modelling'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver