TY - GEN
T1 - Media-Text
T2 - 38th Annual European Simulation and Modelling Conference, ESM 2024
AU - Kalisz, Seweryn
AU - Marczyk, Michał
AU - Fagas, Rafał
AU - Polańska, Joanna
N1 - Publisher Copyright:
© 2024 Modelling and Simulation 2024 - 38th Annual European Simulation and Modelling Conference 2024, ESM 2024. All rights reserved.
PY - 2024
Y1 - 2024
N2 - Detection of text in digital images used in the media industry is challenging, mostly due to the lack of specialized datasets and the unknown performance of existing models. Here, a unique Media-Text dataset, comprising 400 images of banners, posters, and covers, annotated with 7,744 text instances is introduced. Two state-of-the-art deep learning models trained on three publicly available datasets were evaluated on the Media-Text dataset to test their generalization capabilities. Additionally, the application of transfer learning techniques with synthetic data and their impact on improving model performance was tested. Regardless of the model and training dataset used, a significant decrease in text detection performance on the media-type dataset was observed, in comparison to hold-out test sets of public data. These results indicate a strong need for media-specialized datasets to enhance the robustness and adaptability of existing text detection models. Future work will explore the use of diverse data sources and advanced model architectures, such as transformers, to potentially improve model generalization. The dataset and code are available at https://github.com/ZAEDPolSl/MediaText.
AB - Detection of text in digital images used in the media industry is challenging, mostly due to the lack of specialized datasets and the unknown performance of existing models. Here, a unique Media-Text dataset, comprising 400 images of banners, posters, and covers, annotated with 7,744 text instances is introduced. Two state-of-the-art deep learning models trained on three publicly available datasets were evaluated on the Media-Text dataset to test their generalization capabilities. Additionally, the application of transfer learning techniques with synthetic data and their impact on improving model performance was tested. Regardless of the model and training dataset used, a significant decrease in text detection performance on the media-type dataset was observed, in comparison to hold-out test sets of public data. These results indicate a strong need for media-specialized datasets to enhance the robustness and adaptability of existing text detection models. Future work will explore the use of diverse data sources and advanced model architectures, such as transformers, to potentially improve model generalization. The dataset and code are available at https://github.com/ZAEDPolSl/MediaText.
KW - Scene text detection
KW - artificial intelligence
KW - deep learning
KW - scene text dataset
UR - https://www.scopus.com/pages/publications/85210267333
M3 - Conference contribution
AN - SCOPUS:85210267333
T3 - Modelling and Simulation 2024 - 38th Annual European Simulation and Modelling Conference 2024, ESM 2024
SP - 138
EP - 144
BT - Modelling and Simulation 2024 - 38th Annual European Simulation and Modelling Conference 2024, ESM 2024
A2 - Nunez-Gonzalez, Jose David
A2 - Grana Romay, Manuel
A2 - Geril, Philippe
PB - EUROSIS
Y2 - 23 October 2024 through 25 October 2024
ER -