Abstrakt
This article presents the original results of Polish language statistical analysis, based on the orthographic and phonemic language corpus. Phonemic language corpus for Polish was developed by using automatic grapheme-to-phoneme conversion of the source orthographic language corpus, obtained from the National Corpus of Polish (NCP). The corpus contains the most frequently used Polish words, written with the use of phonemic notation. Performed statistical analysis of Polish language based on phonemic language corpus, includes frequency of occurrence calculation of the orthographic and phonemic language components, as well as their sequence. Statistical language data, obtained as a result of performed statistical analysis, enable to develop statistical word-based and phoneme-based language models for Polish. Applying these language models can effectively contribute to efficiency improvement of automatic speech recognition for Polish.
| Język oryginału | angielski |
|---|---|
| Numer artykułu | 5 |
| Czasopismo | Eurasip Journal on Audio, Speech, and Music Processing |
| Tom | 2017 |
| Numer wydania | 1 |
| Identyfikatory DOI | |
| Status publikacji | Opublikowano - 1 gru 2017 |
Obszary tematyczne ASJC Scopus
- Akustyka i ultradźwięki
- Inżynieria elektryczna i elektroniczna
Fingerprint
Zanurz się w tematy badawcze publikacji „Statistical analysis of orthographic and phonemic language corpus for word-based and phoneme-based Polish language modelling”. Razem tworzą niepowtarzalny odcisk palca.Cytowanie
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver