TY - GEN
T1 - Efficient and Robust Scene Text Classification
T2 - 39th Annual European Simulation and Modelling Conference, ESM 2025
AU - Kalisz, Seweryn
AU - Marczyk, Michał
AU - Fagas, Rafał
N1 - Publisher Copyright:
© 2025 Modelling and Simulation 2025 - 39th Annual European Simulation and Modelling Conference 2025, ESM 2025. All rights reserved.
PY - 2025
Y1 - 2025
N2 - This paper addresses the task of classifying detected scene text in images as either natural or artificial, a key capability for enabling computer vision systems to interpret text contextually. We extended the original MediaText dataset by adding detailed annotations and constructed a dedicated training set. In the original dataset, we annotated 6,775 legible instances (5,304 artificial, 1,471 natural), and added 9,037 new (5,350 artificial, 3,687 natural). We evaluated several most popular backbones combined with a lightweight classification module. Hyperparameters were optimized using a Random Search strategy. EfficientNet-based models achieved the best results on the independent test set (balanced accuracy equal to 69.7% and AUC equal to 0.76). Despite fewer parameters than ResNet50, it outperformed them, confirming the effectiveness of compact architectures with minimal classification heads. These findings suggest that lightweight models can achieve competitive performance in scene text classification tasks, offering a favorable trade-off between accuracy and computational efficiency. The updated dataset is available at GitHub repository: https://github.com/ZAEDPolSl/MediaText.
AB - This paper addresses the task of classifying detected scene text in images as either natural or artificial, a key capability for enabling computer vision systems to interpret text contextually. We extended the original MediaText dataset by adding detailed annotations and constructed a dedicated training set. In the original dataset, we annotated 6,775 legible instances (5,304 artificial, 1,471 natural), and added 9,037 new (5,350 artificial, 3,687 natural). We evaluated several most popular backbones combined with a lightweight classification module. Hyperparameters were optimized using a Random Search strategy. EfficientNet-based models achieved the best results on the independent test set (balanced accuracy equal to 69.7% and AUC equal to 0.76). Despite fewer parameters than ResNet50, it outperformed them, confirming the effectiveness of compact architectures with minimal classification heads. These findings suggest that lightweight models can achieve competitive performance in scene text classification tasks, offering a favorable trade-off between accuracy and computational efficiency. The updated dataset is available at GitHub repository: https://github.com/ZAEDPolSl/MediaText.
KW - Scene text classification
KW - artificial intelligence
KW - deep learning
KW - image classification
KW - scene text dataset
UR - https://www.scopus.com/pages/publications/105024232850
M3 - Conference contribution
AN - SCOPUS:105024232850
T3 - Modelling and Simulation 2025 - 39th Annual European Simulation and Modelling Conference 2025, ESM 2025
SP - 69
EP - 74
BT - Modelling and Simulation 2025 - 39th Annual European Simulation and Modelling Conference 2025, ESM 2025
A2 - Bhonsale, Satyajeet S.
A2 - Polanska, Monika E.
A2 - Impe, Jan F.M. Van
PB - EUROSIS
Y2 - 22 October 2025 through 24 October 2025
ER -