Wine knowledge and cellar search
Summary
A conversational interface to wine knowledge and stock records.
The product needed to answer general wine questions and precise cellar queries using different kinds of data.
Two paths: retrieval over wine literature for general questions, and intent classification plus entity extraction to query structured cellar records.
Tech Stack
- Python
- AWS cloud stack
- Deberta LM
- Spacy NER
- Milvus
- Postgres
- OpenAI API
Project workstreams
- 01
2 weeks
Solution Architecture Design
Solution ArchitectDeployment of DRL's Internal Knowledge Bot
Dev Ops - 02
4 weeks
Data Integration Pipelines Development
Data Engineer, Dev OpsData Cleaning & Preprocessing
2x NLP Engineers - 03
3 weeks
Agent Case-Specific Customization and Training
System Engineer, Data Engineer, NLP EngineerSearch Query Optimization
Data Engineer, NLP Engineer - 04
1 weeks
Integration with a Metaverse Avatar, Testing & Deployment
System Engineer, NLP Engineer, Data Engineer, Dev Ops
Tech Challenge
-
Wine knowledge is a complex science, with each bottle presenting unique characteristics. The Conversational Agent should precisely navigate the database from each wine cellar, answering questions from wine types and flavor profiles to the geography of vineyards. It should analyze such factors as grape varietals, terroir influences, and aging processes. As a result, the virtual wine consultant would provide qualified replies and personalized recommendations.
-
Furthermore, the Agent should be capable of drawing parallels between different wines. Whether comparing wines based on similar flavor profiles, terroir characteristics, or aging process, the system needs to excel in comparative wine analysis. It should give dynamic and engaged replies, as a real wine consultant would do.
-
The major challenge was to find a way to match the user question represented within the natural language with a structured query in the cellar database and answer it in a human-like format after extracting information.
Solution
-
The system must handle 2 main cases of user questions: questions about wines that the user already has in his cellar (specific details about the chosen bottle, information about the quantity or presence of the selected wine, listing wines according to given filters like color, region, classification, etc.) and questions related to the wine topic in general.
-
To understand the user's intent, we classify the question with Deberta on the dataset created by our team. Further, we augmented it using ChatGPT to reach the data volume enough to fine-tune the selected model.
-
To handle generic questions about wines, we have created a vast knowledge base containing domain information. To implement it, we processed different wine-related books to create a database of short topic-specific textual pieces of information. All this processed information is uploaded to the Milvus vector database, which allows users to search for relevant information by question. The vector search is based on embeddings produced by the gte-large model. After the most relevant pieces of information are found, they are passed into GPT-3.5 Turbo to summarize it and generate a human-like response for the user.
-
To handle questions about wines that the user already has in his cellar, we created a wide range of templates used to map user questions into predefined SQL queries to select the necessary information from the structured cellar data. All the cellar information is stored under the Postgres database. To efficiently map the user question into a relevant template, we trained the Spacy Roberta-based NER model to extract specific wine-related entities (such as color, region, classification, vintage, vineyard, etc.). This allows us to reduce the search space for the relevant templates. When the most relevant templates are found, the corresponding queries are executed, and the information from the Postgres database is passed into GPT-3.5 Turbo to summarize it and generate a human-like response for the user.
-
Operating solely via APIs enabled the adoption of a serverless architecture, leveraging the AWS toolbox to optimize cloud costs and seamless scalability during load spikes, all under a 100% usage-based charging model.
Outcome
A wine assistant combining domain knowledge with queries against the company’s cellar database.
Published
