Research into expressive virtual assistants

Summary

A 2020 research programme and closed-beta prototype.

The research explored conversational memory, expressive speech and action selection in a virtual assistant.

A prototype combining dialogue models, a knowledge graph, speech synthesis and reinforcement-learning components for interaction.

Tech Stack

  • AWS
  • GCP
  • PyTorch
  • Python
  • React Native
  • TensorFlow
  • Scala

Tech Challenge

  • The research explored controllable synthetic voices using voice-actor recordings.
  • Voice has to be generated with 15 different emotions and intensity levels.
  • The research explored low-latency speech generation for responsive conversation.
  • The system is able to chat with the user, have some common sense on things, know and remember a lot of different information, be disconnected from the internet (we don’t want to create one more Google Assistant).
  • The research explored natural dialogue, conversational memory and expressive responses.

Solution

  • For the voice, seq2seq with attention architecture was used, called tacotron2. It was implemented almost from scratch, by that time there weren't good implementations.
  • For voice emotion control GST tokens were applied, which utilizes multi-head self-attention mechanism.
  • For faster than real-time inference WaveGlow post-processor was used, which is basically parallelized version of WaveNet.
  • For the best possible quality of voice more data was gathered from voice actors. Special technical task was created to reach the best possible quality.
  • For the dialog system, Blender + Large Attention + PPLM is used with custom conditioning mechanism, which gives us control over different Dialogue Acts, Emotions, Topics, Contexts and Tasks (Q&A, chit-chatting, generation).
  • Utterances were processed with SOA and named-entity recognition, and conversation-related information was represented as a graph in Neo4j.
  • In order to solve the dialog management problem, we apply DDPG, Rainbow and A3C + Imagination & Curiosity blocks with Dialogue Acts, Emotions, Topics, Contexts and Tasks as an action space (Reinforcement Learning).
  • In order to solve Reinforcement Learning the WebSocket interface was created for Amazon Mechanical Turk to receive rewards for dialog from users.
  • Dialogue-rating experiments provided feedback on conversational quality during research.

Impact

A research prototype combining dialogue, memory, expressive speech and action selection.

Published

DRL Team · Ivan Didur