Wine knowledge and cellar search

Summary

A conversational interface to wine knowledge and stock records.

The product needed to answer general wine questions and precise cellar queries using different kinds of data.

Two paths: retrieval over wine literature for general questions, and intent classification plus entity extraction to query structured cellar records.

Tech Stack

  • Python
  • AWS cloud stack
  • Deberta LM
  • Spacy NER
  • Milvus
  • Postgres
  • OpenAI API

Project workstreams

  1. 01

    2 weeks

    Solution Architecture Design

    Solution Architect

    Deployment of DRL's Internal Knowledge Bot

    Dev Ops
  2. 02

    4 weeks

    Data Integration Pipelines Development

    Data Engineer, Dev Ops

    Data Cleaning & Preprocessing

    2x NLP Engineers
  3. 03

    3 weeks

    Agent Case-Specific Customization and Training

    System Engineer, Data Engineer, NLP Engineer

    Search Query Optimization

    Data Engineer, NLP Engineer
  4. 04

    1 weeks

    Integration with a Metaverse Avatar, Testing & Deployment

    System Engineer, NLP Engineer, Data Engineer, Dev Ops

Tech Challenge

  • Wine knowledge is a complex science, with each bottle presenting unique characteristics. The Conversational Agent should precisely navigate the database from each wine cellar, answering questions from wine types and flavor profiles to the geography of vineyards. It should analyze such factors as grape varietals, terroir influences, and aging processes. As a result, the virtual wine consultant would provide qualified replies and personalized recommendations.

  • Furthermore, the Agent should be capable of drawing parallels between different wines. Whether comparing wines based on similar flavor profiles, terroir characteristics, or aging process, the system needs to excel in comparative wine analysis. It should give dynamic and engaged replies, as a real wine consultant would do.

  • The major challenge was to find a way to match the user question represented within the natural language with a structured query in the cellar database and answer it in a human-like format after extracting information.

Solution

  • The system must handle 2 main cases of user questions: questions about wines that the user already has in his cellar (specific details about the chosen bottle, information about the quantity or presence of the selected wine, listing wines according to given filters like color, region, classification, etc.) and questions related to the wine topic in general.

  • To understand the user's intent, we classify the question with Deberta on the dataset created by our team. Further, we augmented it using ChatGPT to reach the data volume enough to fine-tune the selected model.

  • To handle generic questions about wines, we have created a vast knowledge base containing domain information. To implement it, we processed different wine-related books to create a database of short topic-specific textual pieces of information. All this processed information is uploaded to the Milvus vector database, which allows users to search for relevant information by question. The vector search is based on embeddings produced by the gte-large model. After the most relevant pieces of information are found, they are passed into GPT-3.5 Turbo to summarize it and generate a human-like response for the user.

  • To handle questions about wines that the user already has in his cellar, we created a wide range of templates used to map user questions into predefined SQL queries to select the necessary information from the structured cellar data. All the cellar information is stored under the Postgres database. To efficiently map the user question into a relevant template, we trained the Spacy Roberta-based NER model to extract specific wine-related entities (such as color, region, classification, vintage, vineyard, etc.). This allows us to reduce the search space for the relevant templates. When the most relevant templates are found, the corresponding queries are executed, and the information from the Postgres database is passed into GPT-3.5 Turbo to summarize it and generate a human-like response for the user.

  • Operating solely via APIs enabled the adoption of a serverless architecture, leveraging the AWS toolbox to optimize cloud costs and seamless scalability during load spikes, all under a 100% usage-based charging model.

Outcome

A wine assistant combining domain knowledge with queries against the company’s cellar database.

Published

DRL Team