Gabriel Skantze
Professor in Speech Technology
Department of Speech Music and Hearing
School of Computer Science and Communication
KTH Royal Institute of Technology

Young Academy of Sweden

Videos and Demos

A Map Task dialogue system

This is a completely automatic dialogue system based on the Map Task - a typical experimental setting for studying human-human dialogue. The user is told that she should follow the route with the mouse cursor for logging purposes. However the real purpose of this is that the system knows what she is actually talking about - there is no speech recognition involved, only speech detection. The purpose of the system is to study feedback behaviour in human-computer dialogue.

Incremental dialogue processing

This is a 45 minute lecture from a dialogue system course, in which I present the work I have done on incremental dialogue processing. (Best viewed in HD and full screen).

Furhat - a talking robot head on exhibition

This is the robot head Furhat exhibited at Tällberg Forum, talking to me (who developed the software for situated multimodal multi-party interaction) and Samer Al Moubayed (who developed the head). The clip is from Swedish Television - the introduction is in Swedish but the dialogue is in English.

The Incremental Turk

This is a fun application of incremental processing that I implemented, using Jindigo and a back-projected head. The Mechanical Turk was a fake chess-playing machine constructed in the late 18th century. Here is an updated version, which is not fake and which uses incremental speech understanding and speech generation, as well as a back-projected head (FurHat). While the robot is pondering the next move, it generates the speech command incrementally.

Incremental speech generation in DEAL

This is a demonstration of how speech can be generated incrementally in dialogue systems. A Wizard-of-Oz setup is used where all modules except the ASR are running. Even if it takes time for the Wizard to transcribe the user's speech, the system may start to speak using fillers, cue phrases and self-repairs. The video shows an interaction with a real user who has never used the system before.

The Higgins spoken dialogue system

This is a demonstration of the initial domain for the Higgins spoken dialogue system: pedestrian navigation. The user is walking around in a virtual city and the system is guiding user. Note that the system does not have any information about the user's position; it has to rely on the user's descriptions of the environment. This is a real run of the system.

The Numbers spoken dialogue system

This is a demonstration of incrementality in spoken dialogue systems. This is a very simple domain, where the user is dictating numbers to the system. However, the system processes the user's speech incrementally and uses prosodic analysis, allowing it to respond very rapidly. Note that the user does not get any visual feedback from the system and that the output shown here is solely for monitoring the internal state of the system.

Speech controlled chess

This is a demonstration of speech controlled chess implemented with the Higgins framework. As you can see, the system's performance is reflected in its emotions.

The MonAmi Reminder

This is a demonstration of a reminder application for the EU-project MonAmi (researching technology for elderly and disabled persons). It illustrates the use of digital pen and paper in combination with speech technology. The Higgins framework is used here as well.

Singing synthesis

The Unreal Group. This is mostly for fun. I did it together with Jonas Beskow.

Song and animation (in Swedish) is completely automatically generated:
1. Notes with text have been scanned into bitmap.
2. Bitmap images have been transformed by SharpEye into MusicXML.
3. We have implemented a program that transforms MusicXML into song and animation.