Research into expressive virtual assistants
Summary
A 2020 research programme and closed-beta prototype.
The research explored conversational memory, expressive speech and action selection in a virtual assistant.
A prototype combining dialogue models, a knowledge graph, speech synthesis and reinforcement-learning components for interaction.
Tech Stack
- AWS
- GCP
- PyTorch
- Python
- React Native
- TensorFlow
- Scala
Tech Challenge
- The research explored controllable synthetic voices using voice-actor recordings.
- Voice has to be generated with 15 different emotions and intensity levels.
- The research explored low-latency speech generation for responsive conversation.
- The system is able to chat with the user, have some common sense on things, know and remember a lot of different information, be disconnected from the internet (we don’t want to create one more Google Assistant).
- The research explored natural dialogue, conversational memory and expressive responses.
Solution
- For the voice, seq2seq with attention architecture was used, called tacotron2. It was implemented almost from scratch, by that time there weren't good implementations.
- For voice emotion control GST tokens were applied, which utilizes multi-head self-attention mechanism.
- For faster than real-time inference WaveGlow post-processor was used, which is basically parallelized version of WaveNet.
- For the best possible quality of voice more data was gathered from voice actors. Special technical task was created to reach the best possible quality.
- For the dialog system, Blender + Large Attention + PPLM is used with custom conditioning mechanism, which gives us control over different Dialogue Acts, Emotions, Topics, Contexts and Tasks (Q&A, chit-chatting, generation).
- Utterances were processed with SOA and named-entity recognition, and conversation-related information was represented as a graph in Neo4j.
- In order to solve the dialog management problem, we apply DDPG, Rainbow and A3C + Imagination & Curiosity blocks with Dialogue Acts, Emotions, Topics, Contexts and Tasks as an action space (Reinforcement Learning).
- In order to solve Reinforcement Learning the WebSocket interface was created for Amazon Mechanical Turk to receive rewards for dialog from users.
- Dialogue-rating experiments provided feedback on conversational quality during research.
Impact
A research prototype combining dialogue, memory, expressive speech and action selection.
Published
