What is the problem with LLMs and safety?

Have you ever faced the problem of LLM generating unsafe or deceptive but very believable stuff? – Google does: after their newly presented GPT-like language model was caught falsifying facts in its demo, Alphabet Inc. lost $100 billion in market value.

What is NeMo Guardrails?

NeMo Guardrails is an open-source toolkit by NVIDIA that aims to solve safety.

With NeMo Guardians, developers can guide LLMs with programmable rules defining user interactions and making the generation more reliable.

The toolkit is compatible with in-house trained LLMs (e.g., from HuggingFace hub) and API-based (e.g., OpenAI’s ChatGPT), providing templates and patterns to build LLM-powered applications.

NVIDIA NeMo Guardrails Runtime
Fig. 1: NeMo Guardrails processing flow.

Why is it important?

The growing adoption of language models in various applications has raised industry-wide concerns about the safety and trustworthiness of the generated information.

LLMs can sometimes produce misleading, incorrect, or offensive content, posing significant user risks. To address these challenges, developers, organizations, and AI researchers prioritize refining models, and using reinforcement learning from human feedback to align AI systems with human values, thus ensuring user trust, ethical standards, and responsible AI technology use across industries and applications.

How does NVIDIA NeMo Guardrails work?

NeMo Guardrails are programmable constraints that monitor and dictate LLM user interactions, keeping conversations on track and within desired parameters.

Also, it's built on Colang, a modeling language developed by NVIDIA for conversational AI, offering a human-readable and extensible interface for users to define dialog flows.

Developers can use a Python library to describe user interactions and integrate these guardrails into applications.

The toolkit supports three categories of guardrails:

  • Topical – ensures that conversations stay focused on a particular topic and prevent them from veering off into undesired areas;

  • Safety – ensures that interactions with an LLM do not result in misinformation, toxic responses, or inappropriate content;

  • Security – prevent an LLM from executing malicious code or calls to an external application in a way that poses security risks.

How to use NVIDIA NeMo Guardrails?

Incorporating NeMo Guardrails into the NeMo framework allows developers to train and fine-tune language models using their domain expertise and datasets.

The developer configures dialog flows, user and bot utterances, and limitations using Colang, a human-readable modeling language.

Showcase: Here are some simple examples of using the NeMo Guardrails:

Conclusions

In conclusion, NVIDIA's NeMo Guardrails toolkit offers a solution to improve the safety and reliability of large language models (LLMs) by providing programmable rules and constraints for user interactions.

Although the toolkit is not a complete solution for all safety issues, it demonstrates the potential to keep conversations on-topic, filter out harmful content, and prevent security risks.

As the adoption of language models in various applications grows, ensuring user trust and responsible AI technology use remains a priority for developers and researchers alike.

← Back to blog