Skip to main content
This page provides information about the training data used in Raven, in line with PolyAI’s commitment to transparency, responsible AI development, and applicable regulatory expectations. The details below describe the provenance, composition, processing, and intended use of the datasets used to develop Raven v3 and v3.5.

System overview

Dataset summary

Intellectual property considerations

Personal and consumer data

Synthetic data usage

Data processing and preparation

The datasets used for Raven v3 and v3.5 have undergone multiple processing steps for quality, safety, and suitability for training customer service agents.

Types of data used

Purpose and intended use

Ongoing governance

PolyAI regularly reviews its data practices against current legal, regulatory, and ethical standards. Dataset composition and processing methods may be updated over time to reflect improvements in safety, coverage, and system performance.
Last modified on June 18, 2026