Audio and motion for a pet wearable

Summary

Sensor models and cloud inference for a connected device.

A connected wearable needed to interpret speech, environmental audio and motion while serving many concurrent sessions.

Specialised audio and IMU models, keyword detection, streaming speech recognition and response generation, connected through a cloud inference pipeline.

Tech Stack

  • Python
  • AWS cloud stack
  • Starlette websockets
  • FastConformer
  • FastFit
  • Gemini API
  • ElevenLabs API
  • Twilio
  • Voice Activity Detection
  • Speaker diarization
  • Keyword spotting
  • Audio classification
  • IMU-based activity recognition
  • Triton Inference Server
  • ONNX
  • TensorRT
  • Postgres
  • Redis
  • Prometheus
  • MongoDB

Tech Challenge

  • Low-latency dialogue flow. Processing of the caregiver's speech, replicas, and response selection should occur with minimal latency, reminding real human-to-human communication.

  • Realistic, natural, and emotional voice. The TTS model should provide a highly expressive voice according to the original voice actor chosen by the user.

  • Few-shot keyword spotting. To save battery life, the system should be activated only after the keyword, the pet's name. However, there will be access to only a few samples of the user's voice recorded during onboarding.

  • High-quality speech recognition in noisy environments. Most of the use cases for collars are outdoor walks with caregivers. Moreover, a collar on a dog, even in a quiet room, is a scenario with additional noise as the dog breathes, moves, whines, or barks.

  • Sound recognition. Developing a robust pet sound recognition system required overcoming the lack of open-source audio datasets and tuning models to capture high-frequency patterns often present in dog vocalizations. This involved custom data collection and complex feature extraction.

  • Emotion recognition. The system includes emotional recognition for pets and caregivers. It uses an audio signal and text context as input to a designed and trained neural network to produce labels to supplement the pet-human interaction context.

  • IMU-based activity detection and analysis. Predicted activities such as resting, running, or jumping should trigger relevant events that drive the pet's well-being analytics. This requires training in a self-hosted IMU-based classification model supported by an ETL pipeline for sensor data collection and preprocessing.

  • Data Collection and Labelling. One of the key challenges was data gathering, which included specific pet-related sounds and IMU sensor readings.

  • Real-time ML models inference at scale. Models should run fast enough to serve predictions instantly (typically in milliseconds) while handling thousands of requests per second. It requires efficient use of computing resources (like GPUs) and the design of scalable, fault-tolerant services.

  • Event-driven alert system. Based on short-term and long-term events and predictions from ML models, the collar should inform the caregiver about nutrition, dehydration, missed activity, sleep, or potential danger to the pet.

Solution

  • For selected interaction scenarios, we used IBM FastFit, a few-shot text-classification method. The project reports an average latency of 20 ms for this classifier, not for the full voice pipeline.
  • For a generic scenario, we used the Gemini API with streaming to generate a response. Our prompt engineer created a golden pet's character so that the answers match the persona selected during onboarding.

  • Voice generation used the ElevenLabs API. Experiments with voice-actor samples informed the voice settings; responses were streamed using the Flash v2.5 text-to-speech model.
  • To handle the keyword spotting challenges, we adapted the open-source PLiX model. It aligns audio and text representations via contrastive learning. So, it enables the recognition of new keywords with just a few audio samples without retraining.

  • The Trillson2 model was trained on a custom dataset for voice activity detection. The project reports 98% accuracy on its validation set; this is a task-specific evaluation rather than an end-to-end system metric.
  • As a speech-to-text model, we use self-hosted NVIDIA FastConformer-CTC XLarge with streaming support. According to the results of our experiments, this model outperformed the well-known proprietary STT solutions currently on the market.

  • To solve the data shortage problem, we set up a pipeline for collecting and labelling IMU sensor data and sound data by assembling a separate team.

  • A BEATs model was trained on a custom dataset for environmental-sound classification. The project reports 90% accuracy on that validation set.
  • For emotion recognition, we use the emotion2vec+ model (classification based on audio) and Spacy NER (extracting emotion entities such as anger, laughter, excitement, sadness, etc. from text utterances).

  • Based on the dataset collected through our pipeline, we trained an effective RNN-based model for classifying IMU sensor-based activities, successfully recognizing different movement patterns.

  • To monitor performance and ensure data quality, key metrics were visualized in real-time using Prometheus and Grafana.

  • Our team used TensorRT SDK to optimize and accelerate trained models for sound, emotion, and IMU classification, ensuring low-latency execution on GPU hardware. These models, as well as the STT model, are then deployed via Triton Inference Server, which provides dynamic batching, concurrent model execution, ensuring scalable and efficient serving.

  • To manage and orchestrate real-time data flow, we utilized Redis streams as a lightweight message broker, making it easy to send, receive, and track inference requests and results across services. This stack allows us to achieve a robust, low-latency inference pipeline capable of handling high loads in production environments.

Impact

  • The developed system allows the collar to recognize the pet's name, interpret behaviors, and respond intelligently through natural voice. This laid the foundation for real-time communication between pets and humans.

  • Load testing exercised more than 10,000 concurrent users. This measures tested system capacity and is distinct from the hundreds of users described during early launch.
  • The product attracted hundreds of users at the early-stage launch, demonstrating strong market interest and the value of our solution.

  • The alert system, which has personalized voices, reminders, and notifications, allows users to stay connected with their pets anywhere. Beyond the product itself, our work sets a new standard for smart collars and encourages further innovation in pet care technology.

Published

DRL Team