26th ACM International Conference on Multimodal Interaction
(4-8 Nov 2024)

Home

Registration

Important dates

Keynote Speakers

Conference Program

Proceedings

Companion Proceedings

Awards

Workshops

Grand Challenges

Tutorials

Special Sessions

Call for Papers

Doctoral Consortium

Blue Sky

Demos

Late-Breaking Results

New Initiatives

Presentation Guidelines

Camera Ready Instructions

Author Guidelines

Reviewer Guidelines

Call for Sponsors

People

Travel, Accomodation and Venue

About Costa Rica

For Families

Visa Information

Student Volunteers

Steering Committee

Platinum Sponsor
Silver Sponsors
blank
blank
Bronze Sponsor
blank
blank
Blue Sky Sponsor
blank
Institutional Sponsor
blank
blank
blank
blank

Workshops

GENEA: Generation and Evaluation of Non-verbal Behaviour for Embodied Agents

Click here to go to the workshop site.

Embodied Social Artificial Intelligence in the form of conversational virtual humans and social robots are becoming key aspects of human-machine interaction. For several decades, researchers from varying fields such as human-computer interaction and robotics, have been proposing methods and models to generate non-verbal behaviour for conversational agents in the form of facial expressions, gestures, and gaze. This workshop aims at bringing together these researchers. The aim of the workshop is to stimulate discussions on how to improve both generation methods and the evaluation of the results, and spark an exchange of ideas and cue possible collaborations.

Organisers:
  • Youngwoo Yoon, ETRI, South Korea
  • Alice Delbosc, DAVI-Les Humaniseurs, France
  • Taras Kucherenko, SEED – Electronic Arts, Sweden
  • Teodor Nikolov, Motorica AI, Sweden
  • Rajmund Nagy, KTH Royal Institute of Technology, Sweden
  • Gustav Eje Henter, KTH Royal Institute of Technology, Sweden

First Multimodal Banquet: Exploring Innovative Technology for Commensality and Human-Food Interaction

Click here to go to the workshop site.

Commensality, the act of eating together, offers a rich multisensory and social experience that can be enhanced through technology. Dining involves interactions with food, where smells, colors, sounds, and textures contribute to a multisensory experience and the table becomes a focal point for social interaction, with nonverbal cues and conversations being indispensable elements of commensal experience. This workshop aims to explore how interactive technology can enrich dining experiences. The other aim is to build an interdisciplinary community around the topics of related to commensality and human-food interaction, with special focus on the role of multimodal interaction among commensal partners sharing food, being humans or artificial dining companions (such as social robots). We aim to collect novel contributions that explore how interactive technology can enhance, facilitate, or make these experiences more enjoyable.

Organisers:
  • Radoslaw Niewiadomski, DIBRIS, University of Genoa, Italy
  • Ferran Altarriba Bertran, Escola Università ria ERAM (Girona), Spain
  • Christopher Dawes, Department of Computer Science, Multi-Sensory Devices (MSD) Research Group, UCL, UK
  • Marianna Obrist, Department of Computer Science, Multi-Sensory Devices (MSD) Research Group, UCL, UK
  • Maurizio Mancini, Department of Computer Science, Sapienza University of Rome, Italy

HumanEYEze: Eye Tracking for Multimodal Human-Centric Computing

Click here to go to the workshop site.

Over the last 20 years, eye tracking has evolved from being a diagnostic tool to a powerful input modality for real-time interactive systems. This was partly driven by advances in eye tracking hardware concerning the devices’ affordability, availability, performance, and form factor. Eye tracking was first used in niche applications in the ’80s and ’90s and then gathered significant attention through research on gaze-based interaction and gaze-supported multimodal interaction. In the last 10-15 years, a third promising direction has emerged: eye-based user and context modeling, i.e., seeing the eyes as an additional modality that provides rich information about user (interactive) behavior and their (interaction) context. The eyes reveal information about visual activities, personality, user intents and goals, attention, expertise, and other cognitive abilities, just to name a few. With that, eye tracking bears great potential for the development of human-centered multimodal AI systems. Gaze-based multimodal user models can be used to, e.g., generate direct feedback to steer the training of AI systems or trigger explicit feedback requests if the user seems to disagree with the output of an AI system. The goal of this workshop is to bring together researchers from eye tracking, multimodal human-computer interaction, and artificial intelligence.

Organisers:
  • Michael Barz, German Research Center for Artificial Intelligence (DFKI)
  • Roman Bednarik, University of Eastern Finland
  • Andreas Bulling, University of Stuttgart
  • Cristina Conati, University of British Columbia
  • Daniel Sonntag, University of Oldenburg & DFKI

Multimodal Co-Construction of Explanations with XAI

Click here to go to the workshop site.

This workshop aims to foster collaboration between Explainable AI (XAI) and multimodal interaction research. XAI aims to make AI systems transparent and understandable, yet current methods primarily cater to experts. Meanwhile, multimodal interaction research focuses on enabling meaningful communication between humans and AI agents. The workshop aims to facilitate cross-fertilization between these fields, recognizing the need for situated and multimodal explanations in XAI. It will explore interactive and social processes involved in the co-construction of explanations, emphasizing multimodal communication channels such as conversational speech, facial expressions, and gestures. By fostering collaboration, the workshop aims to identify key research directions and approaches, improve networking among researchers, and advance understanding of XAI explanations as a multimodal, interactive challenge.

Organisers:
  • Hendrik Buschmeier, Digital Linguistics Lab, Faculty of Linguistics and Literary Studies, Bielefeld University
  • Stefan Kopp, Social Cognitive Systems Group, Faculty of Technology, Bielefeld University
  • Teena C. Hassan, Institute for AI and Autonomous Systems, University of Applied Science Bonn-Rhein-Sieg