Przeskocz do nawigacji głównej Przeskocz do wyszukiwania Przeskocz do głównej treści

Robot-Mediated Explainability for CNN-Based Complex Emotion Recognition: Aligning Visual Attributions With Social Robot Behavior for Trust-Aware HRI

  • Warsaw University of Technology

Wyniki badań: Wkład do czasopismaArtykułrecenzja

Abstrakt

In recent years, the development of socially interactive robots has heightened the need for inter-pretable and trustworthy models capable of recognizing complex human emotions in real time. While deep convolutional neural networks achieve strong performance in facial affect recognition, their decision processes remain largely opaque, which limits user trust and safe deployment in human–robot interaction (HRI). This challenge becomes particularly critical when recognizing complex or subtle emotional states under domain shift, where both performance robustness and explanation stability are required. A robot-mediated explainability framework for complex emotion recognition links the visual attribut of a convolutional neural network (CNN) to social-robot behaviors in real time. The perception module augments a CNN with an Attention Map Alignment Layer (AMAL) that regularizes the internal focus toward semantically meaningful facial regions during training. Model-agnostic eXplainable AI (XAI) methods (SHAP and LIME) provide per-instance rationales at inference and are audited for temporal stability and faithfulness under streaming conditions. Explanations are embodied through a behavior-generation layer that converts saliency into deictic head motion and gaze, optionally governed by an adaptive emphasis policy that modulates when and how strongly to highlight evidence as a function of model confidence, explanation stability, and user-state signals. Evaluation spans AffectNet and EMOTIC, plus an out-of-distribution robot-domain set (OhBot) to probe domain shift. A human–robot interaction study uses a 3×2×2 between-subjects design crossing embodiment (no robot, still gaze, head motion), visual XAI (no heatmap vs heatmap), and emphasis policy (fixed vs adaptive). Across in-distribution and robot-domain tests, AMAL improves accuracy and macro-F1 with lower expected calibration error, and yields higher attribution–landmark overlap with greater temporal stability. In interaction, head motion increases acceptance of system recommendations relative to still-gaze and no-robot baselines, while heatmaps improve task performance irrespective of embodiment. Adaptive emphasis delivers additional gains primarily under well-calibrated predictions and attenuates guidance under low confidence, limiting over-reliance. Overall, aligning internal attention, surfacing faithful rationales, and embodying those rationales through socially legible behaviors enhances both objective performance and calibrated trust. Next steps include on-device deployment on embedded accelerators, longitudinal evaluation of trust dynamics, and integration with language-model-based social-norm explanation modules as a complementary, normative layer.

Język oryginałuangielski
Strony (od–do)42887-42896
Liczba stron10
CzasopismoIEEE Access
Tom14
Identyfikatory DOI
Status publikacjiOpublikowano - 2026

Obszary tematyczne ASJC Scopus

  • Informatyka ogólna
  • Materiałoznawstwo ogólne
  • Inżynieria ogólna

Fingerprint

Zanurz się w tematy badawcze publikacji „Robot-Mediated Explainability for CNN-Based Complex Emotion Recognition: Aligning Visual Attributions With Social Robot Behavior for Trust-Aware HRI”. Razem tworzą niepowtarzalny odcisk palca.

Cytowanie