This page provides information about the training data used in Raven, in line with PolyAI’s commitment to transparency, responsible AI development, and applicable regulatory expectations.
The details below describe the provenance, composition, processing, and intended use of the datasets used to develop Raven v3 and v3.5.
System overview
Dataset summary
Intellectual property considerations
Personal and consumer data
Synthetic data usage
Data processing and preparation
The datasets used for Raven v3 and v3.5 have undergone multiple processing steps for quality, safety, and suitability for training customer service agents.
Types of data used
Purpose and intended use
Ongoing governance
PolyAI regularly reviews its data practices against current legal, regulatory, and ethical standards. Dataset composition and processing methods may be updated over time to reflect improvements in safety, coverage, and system performance. Last modified on June 18, 2026