Beyond Captioning: Weaving Visual and Social Dialogue for Robotic Receptionists

Combining Visual and Social Dialogue for Human-Robot Interaction

2021-10-15
Nancie Gunson, Daniel Hernández García, Jose L. Part, Yanchao Yu, Weronika Sieinska, Christian Dondrup, Oliver Lemon
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a multimodal conversational AI prototype designed for a hospital receptionist robot. It integrates the award-winning Alana social bot with a visual perception system to handle task-based navigation, medical check-ins, and visually-grounded dialogue alongside open-domain social interaction.

TL;DR

Researchers from Heriot-Watt University have unveiled a multimodal prototype that transforms robots from simple tools into socially-aware assistants. By combining the Alana social bot with visual perception and a Petri-Net-based interaction planner, this system allows a hospital receptionist robot to handle everything from medical check-ins and "Where is my bag?" queries to discussing the latest news and coronavirus quizzes.

Background: The Gap in Situated Interaction

In the world of AI, vision and language are often treated as separate silos or linked through static tasks like "Image Captioning." However, in a real-world setting like a hospital waiting room, a robot needs to be situated. It must understand the physical space (Visual Dialogue) while maintaining the social fabric of the interaction (Social Dialogue).

The SPRING project addresses this by creating a Socially Assistive Robot (SAR) that doesn't just answer questions but acts as a proactive, empathetic participant in a shared environment.

The "Brain" of the System: Multi-Bot Architecture

The core innovation lies in the system's modularity. Rather than a single monolithic model, it uses an ensemble of specialized "bots" managed by a priority-based Dialogue Manager.

1. The Dialogue System (The Green Blocks)

Built on the Alana v2 framework, the system includes:

  • Task/Direction Bots: Handles hospital-specific logic like navigation and check-ins.
  • Visual Dialogue Bot: Connects language to the physical world.
  • Social/Quiz Bots: Prevents the interaction from feeling "robotic" by providing entertainment and information.

2. Social Interaction Planner (The Blue Blocks)

This module acts as the conductor. It solves the "Multi-Threaded Dialogue" problem. If a user asks a question that requires looking at the room, the planner executes a Petri-Net Plan (PNP). This allows the robot to manage state, wait for sensor data, and resume conversation without losing the context of the interaction.

System Architecture Figure 1: The architecture shows the interplay between the Social Interaction Planner, the Alana-based Dialogue System, and the ROS Vision Action Server.

Seeing the World: Visual Grounding

The robot uses Detectron2 for scene segmentation, which is then translated into a Scene Graph. While the current prototype uses manually refined scene graphs, it enables the robot to answer spatial queries like "Is there a seat available near the door?" or "I left my jacket on the chair, can you see it?"

Web Interface Example Figure 2: The web-based demonstration interface showing a user checking in while the robot maintains the dialogue state.

Critical Insight: Why This Matters

Most Task-Oriented Dialogue (TOD) systems fail because they are too rigid—one "off-script" comment and the system breaks. By nesting domain-specific bots within an open-domain social framework (Alana), this research provides a "safety net." If the robot doesn't understand a specific medical check-in intent, it can fall back to social conversation to maintain rapport while it attempts to resolve the task error.

Conclusion and Future Outlook

This work represents a "first step" toward a truly autonomous hospital assistant. The researchers are moving toward:

  • Automatic Scene Graph Generation: Removing the need for manual scene rules.
  • Multi-party Interaction: Enabling the robot to handle a group of people in a waiting room, not just a 1-on-1 chat.

As these systems move from web interfaces to physical platforms like the ARI robot, the boundary between "computer vision" and "human conversation" will continue to blur, leading to robots that truly understand both the words we say and the world we inhabit.

Find Similar Papers

Try Our Examples

  • Search for recent papers on Socially Assistive Robots (SAR) that combine Petri-Net Plans with Large Language Models for task management.
  • Which paper first proposed the Alana social bot framework, and how does the current multimodal extension differ from the original version?
  • Explore how vision-language models like CLIP or Grounding DINO are currently being integrated into ROS-based robotic dialogue systems.
Contents
Beyond Captioning: Weaving Visual and Social Dialogue for Robotic Receptionists
1. TL;DR
2. Background: The Gap in Situated Interaction
3. The "Brain" of the System: Multi-Bot Architecture
3.1. 1. The Dialogue System (The Green Blocks)
3.2. 2. Social Interaction Planner (The Blue Blocks)
4. Seeing the World: Visual Grounding
5. Critical Insight: Why This Matters
6. Conclusion and Future Outlook