Adaptive game experimentation

Summary

Reinforcement learning for an A/B testing platform.

A game team needed to explore interacting parameters while campaign goals and product conditions changed.

A system that pre-trains on historical data, runs smaller campaigns, learns from results and iterates against product-owner goals.

Tech Stack

  • Firebase
  • OpenAI Gym
  • Python
  • TensorFlow

Project workstreams

  1. 01

    1 Week

    Solution Architecture Design

    Solution Architect
  2. 02

    2 Weeks

    Hypothesis Generation & Validation

    Deep Learning Researcher
  3. 03

    4 Weeks

    RL Environment Development

    Data Engineer,Deep Learning Engineer,Deep Learning Researcher
  4. 04

    2 Weeks

    RL Algorithms Development

    Deep Learning Researcher
  5. 05

    6 Weeks

    Training & Tuning Cycle pt.1

    Deep Learning Researcher
  6. 06

    2 Weeks

    Data Labelling & Processing & Integration into RL Environment

    Data Engineer
  7. 07

    2 Weeks

    Training & Tuning Cycle pt.2

    Deep Learning Researcher
  8. 08

    2 Weeks

    Integration & A/B Testing & Deployment

    Backend Developer,Dev Ops

Tech Challenge

  • When tested product goes through various stages of its life-cycle, goals might swap priorities.
  • Most of the tests are created manually and managed by humans, with only few parameters changed at once not to introduce over-complexity into results’ interpretation.
  • Different parameters, competing goals, lots of statistical events are the factors that introduce additional complexity and need automatic insights extraction.
  • Finding complex interaction between the entire list of A/B parameters and real-world state is extremely technically challenging.

Solution

  • Our team used recently introduced techniques from reinforcement and deep learning to model complex parameters interaction, time dependent external state and A/B campaign events stream.
  • As a result, we have developed a self-learning solution which, firstly, pre-trains itself using historical data and then takes control over the campaigns. Product owner remains in control of goals and parameters.
  • System creates hundreds of micro-campaigns with different goals and fast-checks hypotheses on a user base sample. Then, it explores the resulting space with deeper tests to maximize the outcomes in each cycle.
  • Campaign results are integrated back via learning process and checked against updated goals; the cycle repeats.
  • AI explores sophisticated dependencies in parameters/state spaces and their relations with the list of goals, provides opportunities beyond conventional approach to A/B testing, while the rest is controlled by product owner in an intuitive way.

Impact

The approach was tested on a mobile game and compared with a Bayesian optimisation baseline.

Published

DRL Team · Ivan Didur