EVALUATE Generative AI Behavioral Safety

Evaluating how AI influences human judgment, trust, and behavior.

white wind turbine under blue sky during daytime

Overview

EVALUATE develops methods for assessing how AI systems influence human judgment, trust, decision-making, and behavior.

Traditional AI evaluation often focuses on model performance, accuracy, or technical safety. Behavioral AI evaluation extends this perspective by examining what happens when people interact with AI systems—including persuasion, overreliance, trust formation, dependency, and manipulation.

Our goal is to develop practical frameworks that help organizations identify behavioral risks before AI systems are deployed at scale.

Research Approach

Our research combines behavioral science, experimental methods, and AI evaluation to measure how people respond to and make decisions with AI systems.

We examine behavioral outcomes across human–AI interactions, including changes in trust, reliance, judgment, risk perception, and decision-making. This includes evaluating not only individual AI responses, but also how behavioral effects emerge across extended conversations and repeated interactions.

These insights inform the development of behavioral evaluation protocols, benchmarks, and Behavioral Red Teaming methods designed to identify risks that conventional technical evaluations may overlook.

Expected Inpact

EVALUATE aims to expand AI safety evaluation beyond model behavior to include measurable human outcomes.

By identifying how AI systems influence trust, judgment, reliance, and decision-making, Behavioral AI evaluation can help organizations detect risks earlier, compare systems more systematically, and design safeguards around real human behavior.

Over time, this work aims to contribute practical benchmarks and evaluation methods for safer, more trustworthy, and more human-centered AI systems.

EVALUATE Generative AI Behavioral Safety

Evaluating how AI influences human judgment, trust, and behavior.

white wind turbine under blue sky during daytime

Overview

EVALUATE develops methods for assessing how AI systems influence human judgment, trust, decision-making, and behavior.

Traditional AI evaluation often focuses on model performance, accuracy, or technical safety. Behavioral AI evaluation extends this perspective by examining what happens when people interact with AI systems—including persuasion, overreliance, trust formation, dependency, and manipulation.

Our goal is to develop practical frameworks that help organizations identify behavioral risks before AI systems are deployed at scale.

Research Approach

Our research combines behavioral science, experimental methods, and AI evaluation to measure how people respond to and make decisions with AI systems.

We examine behavioral outcomes across human–AI interactions, including changes in trust, reliance, judgment, risk perception, and decision-making. This includes evaluating not only individual AI responses, but also how behavioral effects emerge across extended conversations and repeated interactions.

These insights inform the development of behavioral evaluation protocols, benchmarks, and Behavioral Red Teaming methods designed to identify risks that conventional technical evaluations may overlook.

Expected Inpact

EVALUATE aims to expand AI safety evaluation beyond model behavior to include measurable human outcomes.

By identifying how AI systems influence trust, judgment, reliance, and decision-making, Behavioral AI evaluation can help organizations detect risks earlier, compare systems more systematically, and design safeguards around real human behavior.

Over time, this work aims to contribute practical benchmarks and evaluation methods for safer, more trustworthy, and more human-centered AI systems.