|
Gabriel Skantze
Professor in Speech Technology
|
Department of Speech Music and Hearing
School of Computer Science and Communication
KTH Royal Institute of Technology
Young Academy of Sweden |
This is a completely automatic dialogue system based on the Map Task - a typical experimental setting for studying human-human dialogue. The user is told that she should follow the route with the mouse cursor for logging purposes. However the real purpose of this is that the system knows what she is actually talking about - there is no speech recognition involved, only speech detection. The purpose of the system is to study feedback behaviour in human-computer dialogue.
This is a 45 minute lecture from a dialogue system course, in which I present the work I have done on incremental dialogue processing. (Best viewed in HD and full screen).
This is the robot head Furhat exhibited at Tällberg Forum, talking to me (who developed the software for situated multimodal multi-party interaction) and Samer Al Moubayed (who developed the head). The clip is from Swedish Television - the introduction is in Swedish but the dialogue is in English.
This is a fun application of incremental processing that I implemented, using Jindigo and a back-projected head. The Mechanical Turk was a fake chess-playing machine constructed in the late 18th century. Here is an updated version, which is not fake and which uses incremental speech understanding and speech generation, as well as a back-projected head (FurHat). While the robot is pondering the next move, it generates the speech command incrementally.
This is a demonstration of how speech can be generated incrementally in dialogue systems. A Wizard-of-Oz setup is used where all modules except the ASR are running. Even if it takes time for the Wizard to transcribe the user's speech, the system may start to speak using fillers, cue phrases and self-repairs. The video shows an interaction with a real user who has never used the system before.
This is a demonstration of the initial domain for the Higgins spoken dialogue system: pedestrian navigation. The user is walking around in a virtual city and the system is guiding user. Note that the system does not have any information about the user's position; it has to rely on the user's descriptions of the environment. This is a real run of the system.
This is a demonstration of incrementality in spoken dialogue systems. This is a very simple domain, where the user is dictating numbers to the system. However, the system processes the user's speech incrementally and uses prosodic analysis, allowing it to respond very rapidly. Note that the user does not get any visual feedback from the system and that the output shown here is solely for monitoring the internal state of the system.
This is a demonstration of speech controlled chess implemented with the Higgins framework. As you can see, the system's performance is reflected in its emotions.
This is a demonstration of a reminder application for the EU-project MonAmi (researching technology for elderly and disabled persons). It illustrates the use of digital pen and paper in combination with speech technology. The Higgins framework is used here as well.
The Unreal Group. This is mostly for fun. I did it together with Jonas Beskow.
Song and animation (in Swedish) is completely automatically generated:
1. Notes with text have been scanned into bitmap.
2. Bitmap images have been transformed by SharpEye into MusicXML.
3. We have implemented a program that transforms MusicXML into song and animation.