TY - GEN
T1 - Comparison of Speech Recognition and Natural Language Understanding Frameworks for Detection of Dangers with Smart Wearables
AU - Mrozek, Dariusz
AU - Kwaśnicki, Szymon
AU - Sunderam, Vaidy
AU - Małysiak-Mrozek, Bożena
AU - Tokarz, Krzysztof
AU - Kozielski, Stanisław
N1 - Publisher Copyright:
© 2021, Springer Nature Switzerland AG.
PY - 2021
Y1 - 2021
N2 - Wearable IoT devices that can register and transmit human voice can be invaluable in personal situations, such as summoning assistance in emergency healthcare situations. Such applications would benefit greatly from automated voice analysis to detect and classify voice signals. In this paper, we compare selected Speech Recognition (SR) and Natural Language Understanding (NLU) frameworks for Cloud-based detection of voice-based assistance calls. We experimentally test several services for speech-to-text transcription and intention recognition available on selected large Cloud platforms. Finally, we evaluate the influence of the manner of speaking and ambient noise on the quality of recognition of emergency calls. Our results show that many services can correctly translate voice to text and provide a correct interpretation of caller intent. Still, speech artifacts (tone, accent, diction), which can differ even for each individual in various situations, significantly influences the performance of speech recognition.
AB - Wearable IoT devices that can register and transmit human voice can be invaluable in personal situations, such as summoning assistance in emergency healthcare situations. Such applications would benefit greatly from automated voice analysis to detect and classify voice signals. In this paper, we compare selected Speech Recognition (SR) and Natural Language Understanding (NLU) frameworks for Cloud-based detection of voice-based assistance calls. We experimentally test several services for speech-to-text transcription and intention recognition available on selected large Cloud platforms. Finally, we evaluate the influence of the manner of speaking and ambient noise on the quality of recognition of emergency calls. Our results show that many services can correctly translate voice to text and provide a correct interpretation of caller intent. Still, speech artifacts (tone, accent, diction), which can differ even for each individual in various situations, significantly influences the performance of speech recognition.
KW - Cloud computing
KW - Intention recognition
KW - Internet of Things
KW - Natural language processing
KW - Older adults
KW - Speech recognition
KW - Wearable sensors
UR - https://www.scopus.com/pages/publications/85111429858
U2 - 10.1007/978-3-030-77970-2_36
DO - 10.1007/978-3-030-77970-2_36
M3 - Conference contribution
AN - SCOPUS:85111429858
SN - 9783030779696
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 471
EP - 484
BT - Computational Science – ICCS 2021 - 21st International Conference, Proceedings
A2 - Paszynski, Maciej
A2 - Kranzlmüller, Dieter
A2 - Kranzlmüller, Dieter
A2 - Krzhizhanovskaya, Valeria V.
A2 - Dongarra, Jack J.
A2 - Sloot, Peter M.A.
A2 - Sloot, Peter M.A.
A2 - Sloot, Peter M.A.
PB - Springer Science and Business Media Deutschland GmbH
T2 - 21st International Conference on Computational Science, ICCS 2021
Y2 - 16 June 2021 through 18 June 2021
ER -