Abstract
The paper presents a deep-learning model that may be used to calculate saliency maps for video content. Classic saliency algorithms take into account only spatial information obtained from an image to generate a saliency map showing regions of higher importance-hence, regions where people are more likely to turn their gaze and attention. Algorithms for video stimuli add temporal information about the movements of objects on a frame-to-frame basis, resulting in more complex three-dimensional architectures. The paper analyses existing models and proposes a model based on one of them. The model's performance is compared with the literature using four widely accepted measures (AUC, NSS, SIM, and CC). It is comparable and, in many cases, even better than already published models. Additionally, because of some improvements in the architecture, it is significantly faster in terms of the number of frames processed per second.
| Original language | English |
|---|---|
| Pages (from-to) | 2922-2932 |
| Number of pages | 11 |
| Journal | Procedia Computer Science |
| Volume | 246 |
| Issue number | C |
| DOIs | |
| Publication status | Published - 2024 |
| Event | 28th International Conference on Knowledge Based and Intelligent information and Engineering Systems, KES 2024 - Seville, Spain Duration: 11 Nov 2022 → 12 Nov 2022 |
Keywords
- deep learning
- eye movements
- saliency maps
- video stimuli
ASJC Scopus subject areas
- General Computer Science
Fingerprint
Dive into the research topics of 'Creating saliency maps for video stimuli and comparing with eye movement data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver