The AIxHMI 2026 will take place on October 7, 2026 starting 10:30 at the University of Perugia.
The exact venue is the room Pittura 4 at the Accademia di Belle Arti, Piazza San Francesco al Prato, 5, 06123 Perugia PG.
Rediscovering the Human in Human–Machine Interaction: Affective Computing, Ethics, and Responsible AI
Affective Computing invites us to rethink AI from the perspective of human experience: emotions, vulnerability, care, context, and relationship. This talk discusses how affective technologies can support ethical human–machine interaction by preserving relational spaces, supporting human judgement, and helping the human dimension emerge more fully within intelligent systems.
Zhe Sun
Digital Brain Foundation Models: Integrating MRI, Brain Dynamics and Generative AI for Human-Machine Interaction
The talk will present our recent work on multimodal MRI foundation models, digital brain construction, large-scale brain dynamics, and generative AI, together with their potential applications to next-generation human-machine interaction and human-centered AI.
From ERP Screens to Conversational Intelligence: The Evolution of Human-Machine Interaction in Enterprise Decision-Making
Bismit Pratapsingh
Human-machine interaction (HMI) within enterprise environments has undergone significant transformation over the past three decades. Early enterprise systems relied primarily on structured forms, menu-driven interfaces, and transaction-oriented workflows that required users to navigate complex application screens and manually interpret business information. The subsequent emergence of business intelligence platforms, analytics-driven decision support systems, and predictive technologies expanded enterprise interaction toward information-oriented and analytics-oriented forms of decision support. More recently, advances in Generative Artificial Intelligence (GenAI), Large Language Models (LLMs), and conversational interfaces have further expanded these capabilities by enabling natural language dialogue, intelligent recommendations, and more iterative forms of human-AI interaction.
This paper examines the evolution of human-machine interaction in enterprise decision-making through five cumulative and overlapping interaction paradigms: Transaction-Oriented Interaction, Information-Oriented Interaction, Analytics-Oriented Interaction, AI-Augmented Interaction, and Conversational Collaboration. Rather than representing a strict replacement sequence, these paradigms can coexist within contemporary enterprise environments and support different users, tasks, and decision contexts. The study examines how advances in artificial intelligence are reshaping user roles and decision-support mechanisms while emphasizing the continuing importance of human judgment, trust, transparency, explainability, and oversight.
The paper further develops a conceptual perspective on conversational intelligence that distinguishes natural language interaction from genuine human-AI collaboration. Collaborative interaction requires not only dialogue-based interfaces but also complementary contributions from AI-generated analysis and human contextual judgment, validation, and decision authority. The study identifies opportunities and limitations associated with conversational intelligence and emphasizes that its benefits depend on task characteristics, system capabilities, user expertise, and appropriate human oversight. Future research is proposed to empirically validate the conceptual model and operationalize human-centered factors across different enterprise contexts.
Designing Human-mediated AI Guidance: Ready Together for Personalized Family Emergency Preparedness
Nini Kurashvili, Yana Ivanchenko, Greta Schiavo, Cansu Koyuturk and Dimitri Ognibene
Artificial intelligence (AI) systems are increasingly used across domains to provide personalized information, recommendations, and decision support. However, in some contexts, AI-generated information may not be suitable for direct delivery to the final recipient. Instead, it may need to be interpreted, adapted, and communicated by a human who understands the recipient’s needs, emotional state, and situational context. Human–AI interaction research has given less attention to situations in which a more knowledgeable human acts as an intermediary between an AI system and a less experienced or less informed recipient. We introduce \textit{the human-mediated AI guidance framework} and explore it through Ready Together, an AI-supported family emergency preparedness system in which parents mediate AI-generated content for their children. The system is designed to provide personalized guidance and support parents in making emergency preparedness more interactive and understandable through guided activities and family-centered learning. The system design was informed by a qualitative, design-oriented research process involving semi-structured interviews and co-design activities. Findings identified challenges in family emergency preparedness, including difficulty discussing emergencies with children, uncertainty about providing appropriate explanations, and a preference for interactive learning activities. These findings informed the design of an interactive prototype, subsequently evaluated through a pilot study and a heuristic evaluation. Participants responded positively to the personalized recommendations and practical activities. Preliminary findings suggest that human-mediated AI guidance may support context-sensitive family preparedness while preserving parents' responsibility for interpreting, adapting, and communicating AI-generated information.
When Attention Makes Flowers Bloom: An Electroencephalography-Based Neurofeedback System for Attention Enhancement in Individuals with ADHD
Đorđe Veljković, Linn Håberg Dimmen, Julianne Brendeløkken and Marta Molinas
This study investigated electroencephalography (EEG) patterns in adolescents and young adults with (n=4) and without (n=7) ADHD diagnoses during three active neurofeedback sessions, each consisting of four runs, with the goal of developing and verifying a gamified ADHD neurofeedback user interface. Participants received feedback based on two different EEG features: {the first two runs using} the theta-beta ratio (TBR), and the last two runs using the slope of the aperiodic component of the power spectrum. The neurofeedback consisted of a single sunflower growing with speed determined by predefined thresholds. The results show that some participants, particularly in the ADHD group, were able to modulate the feedback signals during individual runs, although most struggled in both groups. Offline analyses revealed individual differences in metrics both within and between participants, indicating the need for future work with adapting thresholds to individuals across sessions, and during runs. Although little evidence of systematic improvement across runs and sessions was observed, the work presents a solid foundation to build upon in further studies of neurofeedback applied to ADHD.
Feature-Space Data Augmentation for Exploratory RMD Classification Using Smartphone-Based TUG Recordings
Judith Urbina, David Arcos and Jordi Solé-Casals
Rheumatic and musculoskeletal diseases (RMDs) may be associated with frailty, mobility impairments, and an increased risk of falls. The Timed Up and Go (TUG) test is widely used to assess functional mobility, and smartphone-embedded inertial measurement units (IMUs) can provide movement information beyond test completion time. However, limited clinical datasets may constrain the development of data-driven classification models. This exploratory study investigated whether feature-space data augmentation could improve RMD-status classification using acceleration-norm features recorded by two smartphones during the TUG. Linear-acceleration recordings from 57 participants were preprocessed, and their Euclidean norms were calculated separately for smartphones carried in a pocket and a shoulder bag. Temporal, spectral, nonlinear, and participant-level gait features were extracted, and ReliefF was used for feature selection. Logistic regression (LR), a linear-kernel SVM (SVM-linear), and a radial-basis-function SVM (SVM-RBF) were evaluated without augmentation and with 19 synthetic observations per class generated using SMOTE or ADASYN. Augmentation affected the classifiers and metrics differently. Overall, SVM-RBF obtained the highest mean performance, and its ADASYN configuration achieved the highest mean area under the ROC curve (AUC), (0.712). For SVM-RBF, augmentation yielded higher mean AUC, accuracy, and specificity but lower F1-score and sensitivity. However, the Nadeau--Bengio corrected resampled (t)-tests did not support a reliable improvement in mean AUC. The results also did not show that the observed effects were caused specifically by model nonlinearity. The small sample, absence of external validation, and uncertain biomechanical plausibility of the synthetic feature vectors limit the conclusions.
Understanding Emotions in Music through Audio-Language Models: from Acoustic and Linguistic Characteristics to Semantic Enriched Audio Embeddings
Aurora Saibene, Luca Colombo, Andrea Orsenigo and Francesca Gasparini
Music Emotion Recognition (MER) is a difficult task as audio and lyrics convey affective information in different ways and as emotion can be perceived, induced, or intended. In this work, we propose a refinement of audio-based emotion representation by exploiting audio-language models, namely Contrastive Language-Audio Pretraining (CLAP) and MuLan. We benchmark on Wav2Vec 2.0 and Hidden unit BERT (HuBERT) as speech-based self-supervised learning strategies and find that our proposed approaches improve MER performances, balancing the interpretation of MERGE dataset songs falling in the quadrants of the valence/arousal plane.
Particularly, we achieve 0.83 Macro F1 on the four-class MER task by using MuLan as an audio feature extractor and combining it with an eXtreme Gradient Boosting (XGBoost) classifier.
This suggests that prompts introducing textual information to align audio with perceived emotions, complements acoustic and linguistic information. Combining our best audio-representation with lyrics embeddings obtained using Robustly Optimized BERT Pre-Training Approach (RoBERTa) and metadata, we further improve the performances (Macro F1 equal to 0.85) confirming that MER requires the understanding of emotion considering all the elements constituting a song.
Addressing the Small Sample Size Problem in Stroke Rehabilitation: Feature Relevance Analysis of IMU Signals
Francesca Gasparini, Sara Nocco, Nadim El Hanafi, Amelia Gioscia and Stefano Mazzoleni
Quantitative assessment of upper limb motor impairment using wearable Inertial Measurement Units (IMUs) holds significant promise for personalised stroke rehabilitation. However, clinical machine learning models frequently suffer from the "curse of dimensionality", where small clinical sample sizes paired with high-dimensional inertial sensor data lead to severe model overfitting and poor generalizability.
To address this critical data scarcity, this study introduces a novel framework utilising data from a group of healthy subjects, who have also simulated upper limb motor impairments, to rigorously analyse feature relevance and perform dimensionality reduction prior to clinical deployment.
Our strategy leverages a hybrid knowledge-driven and data-driven approach. First, domain experts pre-selected a comprehensive pool of clinically meaningful features related to specific joint angles and segments derived from IMU data during multiple motor tasks. Second, a multi-task classification pipeline was deployed to evaluate the sensitivity and relevance of these features across different tasks and simulated impairments.
This framework allows to identify a subset of core kinematic biomarkers that may be useful in characterising upper-limb motor impairments, and demonstrate that healthy and simulated-impairment data offer a scalable, controlled, and reliable alternative for establishing feature relevance, effectively mitigating the small sample size problem in stroke rehabilitation.
In this way, we laid the groundwork for a pipeline that not only can be successfully translated to real clinical stroke cohorts, but can also characterize patient-specific upper-limb motor impairments.
Beyond Valence and Arousal: Perceptual and Semantic Diversity for Multimodal Affective Stimulus Selection
Claudia Rabaioli, Aleph Campos da Silveira, Alessandra Grossi, Giulia Rizzi, Veikko Surakka, Roope Raisamo and Francesca Gasparini
Affective stimulus selection is a fundamental step in emotion elicitation studies, yet stimulus assignment is often performed through manual procedures, random selection strategies, or study-specific scripts, with limited support for participant-specific and procedural constraints. As a result, factors such as familiarity, stimulus repetition, duration requirements, or individual exclusions are frequently addressed only after data collection, potentially introducing biases in emotional responses and reducing experimental control. Participant-tailored stimulus selection approaches address this issue by enabling reproducible customized playlist generation that jointly considers affective targets, participant-specific constraints, and experimental requirements.
This work investigates the transferability of this selection strategy across different types of stimulus, including film clips, visual artwork, and classical music excerpts. By extending the analysis beyond a single media domain, we explore the potential of a unified methodology for selecting heterogeneous affective stimuli.
The cross-domain analysis further reveals a limitation of current affective selection practices. Traditional selection procedures often rely on random sampling or affective matching alone. Stimuli satisfying the same valence--arousal targets and experimental constraints may still exhibit recurring perceptual and semantic characteristics, potentially reducing variability and increasing familiarity or habituation effects.
To address these issues, we discuss a theoretical extension toward multimodal affective stimulus selection. The proposed framework introduces additional dimensions derived from audiovisual, speech-related, semantic, and multimodal representations, allowing future selection strategies to jointly consider affective consistency, participant constraints, and perceptual-semantic diversity.
This work lays the foundation for diversity-aware affective stimulus selection frameworks capable of balancing emotional validity, participant constraints, and perceptual-semantic diversity across heterogeneous media modalities.
A Hybrid Approach to Board Games for Autistic Spectrum Disorders
Samuel De Benedictis, Giovanna Daverio and N. Alberto Borghese
We introduce a hybrid version of “The town of unforeseen events”, a physical board game co-designed with therapists, teachers and adolescents with Autism Spectrum Disorders (ASD) at the Vocational Training and Job Integration Center (CFPIL) of Varese, Italy. This game is inspired to the “Royal goose game” and provides a set of activities targeted to train social competences of children and adolescents with Autism Spectrum Disorders (ASD). The hybrid version extends the original physical version in several directions. First, the board can be personalized, selecting the mix of activities most suitable to the actual patients cohort. We also introduce digital videos inside some activities, to provide a more challenging and ecological framework, that is known to be more effective. Lastly, we provide a natural way to hide from some of the players the solution of an activity, while all other players can see it. The game is organized in a client-server architecture with two types of clients: one for the main shared display, that can be a large screen or an overhead projector, and one for the smart phone of each player. The platform follows a Human-in-the-Loop model, with the trainer retaining central control over game flow and evaluation. We also discuss how Artificial Intelligence can be used to help the trainer in personalize the game board and the game difficulty to the actual players cohort.
Integrating Conversational AI and Mixed Reality for Personalized Health Monitoring in Smart homes for Older Adults
Atieh Mahroo, Daniela Micucci, Paolo Napoletano and Marco Sacco
Population aging is increasing the need for integrated technologies that support independent living, safety, and personalized well-being for older adults. This paper presents a smart home-Mixed Reality architecture with a conversational AI interaction layer for personalized older adult support. The proposed architecture brings together MR-based smart home interaction and conversational AI within a broader framework designed to incorporate IoT-based environmental monitoring, wearable sensing, and adaptive physical and cognitive exercise delivery. The MR interface enables users to interact with smart home functions and exercise activities through spatially anchored virtual elements while maintaining awareness of their physical environment. The conversational AI agent serves as an accessible user-facing interface that simplifies interaction, provides step-by-step guidance, collects user feedback, and supports personalized adaptation of system behavior and exercise parameters. By combining multimodal sensing, immersive interaction, and AI-driven personalization, the framework aims to reduce technological barriers, improve engagement, and support continuous health-oriented assistance in aging-in-place contexts. The proposed architecture contributes toward more cohesive, adaptive, and human-centered smart home environments for older adults.