Recommendations for a health app
Summary
Personalised discovery across a large content catalogue.
The app needed to match users with relevant content while balancing relevance, content quality and response time.
A recommendation pipeline using user and content embeddings, quality scoring and regular retraining, integrated with the app’s serving infrastructure.
Tech stack
- AWS
- ECS Fargate
- Lambda
- MSK (Kafka)
- OpenSearch
- Milvus
- Python
- Gunicorn
- aiohttp
- OpenSearch-py
- boto3
- PyTorch
- Transformers
- Scikit-Learn
- Pandas
- Numpy
- Terraform
- AWS CloudFormation
- GitHub
- CodePipelines
Project workstreams
- 01
3 weeks
Collection of Initial Requirements
Solution ArchitectSolution Architecture Design
Solution Architect, System Engineer - 02
6 weeks
Engineering Development
System Engineer, NLP Engineer, Dev OpsData Integration With the App (via Kafka)
Data Engineer, Dev Ops - 03
4 weeks
Data Cleaning & Preprocessing
Data Engineer, NLP EngineerModel Training
Data Engineer, NLP EngineerPerformance Dashboards Setup
NLP Engineer - 04
Every 4 weeks
Review the recommendations' performance
Data Engineer, NLP EngineerModify / Retrain recommendation model and tune algorithms if needed
System Engineer, Data Engineer, NLP Engineer
Tech Challenge
-
Audio content, with its multifaceted attributes, poses a significant challenge for recommendation systems. The answer itself, along with the question's topic, relevance, provided value, and wisdom, all play crucial roles in the process of suggesting.
-
All the inputs are usually recorded via phone. Thus, many things might contribute to overall audio quality — mic sensitivity, environmental noise, speech loudness, speed, pitch, etc.
- Response time was an important serving constraint. The system needed to return recommendations promptly while keeping personalisation relevant to current user context.
-
All recommendation systems have the cold start problem. It is usually related to new users who don't have historical activity. Such users should still be able to get hints right after signup and eventually should get a more personalized experience.
Solution
-
We trained an ML model to create embeddings for answers and users. The model features include audio transcription, question text, author description, and historical stats like listen rate, likes, comments, etc. Every day, the model is automatically retrained to tune weights and adjust embeddings based on the latest user's activity.
-
All answers on the platform were moderated by humans and scored according to rules. Those scores reflected how good the answers were, how precisely they answered the topic, and how much value they gave to the listener.
-
At the later stage of the project, we implemented an auto-scoring mechanism to reduce the amount of human work on the platform. We developed a comprehensive set of OpenAI ChatGPT prompts to evaluate the content score by the same criteria as humans did.
-
Additionally, we developed an audio processing pipeline to analyze the audio quality based on a set of features like speech ratio, short-time objective intelligibility (STOI), perceptual evaluation of speech quality (PESQ), signal-to-noise ratio (SNR), etc.
-
The system is primarily built using Python tech stack and scalable databases and is deployed on AWS Cloud. All the services are running on Lambdas and ECS. External integration with the app is performed via Kafka. Utilizing the OpenSearch and Milvus databases allows the creation of on-the-fly recommendations of 20-50 items in ~100-300ms time.
Outcome
A recommendation engine integrated with the app, combining user relevance with content-quality signals.
The project reports approximately 100–300 ms to return 20–50 recommendations; this describes the recommendation response, not a health outcome.
Published
