Skip to main navigation Skip to search Skip to main content

Statistical analysis of orthographic and phonemic language corpus for word-based and phoneme-based Polish language modelling

Research output: Contribution to journalArticlepeer-review

12 Citations (Scopus)

Abstract

This article presents the original results of Polish language statistical analysis, based on the orthographic and phonemic language corpus. Phonemic language corpus for Polish was developed by using automatic grapheme-to-phoneme conversion of the source orthographic language corpus, obtained from the National Corpus of Polish (NCP). The corpus contains the most frequently used Polish words, written with the use of phonemic notation. Performed statistical analysis of Polish language based on phonemic language corpus, includes frequency of occurrence calculation of the orthographic and phonemic language components, as well as their sequence. Statistical language data, obtained as a result of performed statistical analysis, enable to develop statistical word-based and phoneme-based language models for Polish. Applying these language models can effectively contribute to efficiency improvement of automatic speech recognition for Polish.

Original languageEnglish
Article number5
JournalEurasip Journal on Audio, Speech, and Music Processing
Volume2017
Issue number1
DOIs
Publication statusPublished - 1 Dec 2017

Keywords

  • Automatic grapheme-to-phoneme conversion
  • Automatic speech recognition
  • Language corpus
  • Language modelling
  • Language statistical analysis

ASJC Scopus subject areas

  • Acoustics and Ultrasonics
  • Electrical and Electronic Engineering

Fingerprint

Dive into the research topics of 'Statistical analysis of orthographic and phonemic language corpus for word-based and phoneme-based Polish language modelling'. Together they form a unique fingerprint.

Cite this