Hero Cover

The Missing Human Layer in AI Safety

Human-Centered AI

This article is based on a keynote I delivered at Tokyo AI on designing AI systems that remain safe after deployment.

Many discussions about AI safety focus on models, evaluations, governance, and reliability.

These are all essential.

Yet many AI systems fail for a different reason:

Human behavior.

The pattern I keep seeing

Across my career—in government platforms, consumer technology, and enterprise AI—I have observed the same failure mode, again and again.

The system works.
The outcome doesn’t.

A technically excellent AI system can still fail if the organization underestimates how people will use it, trust it, misuse it, or adapt to it.

Technology creates capabilities.
Human behavior determines outcomes.

Four questions every AI team should ask before launch

1. Will people trust it—appropriately?

Trust is not binary. Too little leads to rejection. Too much leads to overreliance.

The goal is calibrated trust: users who understand when to rely on the system and when not to. This is harder to design for than most teams expect.

2. What behaviors are we rewarding?

People respond to incentives, often in ways designers never anticipated.

At a ride-sharing platform, a dynamic pricing model successfully increased supply during high-demand periods. But during severe weather events, the same incentive structure created an unintended tension: drivers were financially motivated to stay on the road in unsafe conditions.

The technical system worked exactly as designed.

The behavioral outcome required careful redesign.

3. Will it create cognitive overload?

One of the most overlooked barriers to AI adoption is cognitive burden.

In a workplace safety reporting system, even well-designed processes can fail when reporting becomes too complex. Generative AI offered an opportunity to reverse the interaction: instead of asking employees to navigate lengthy forms, the system could guide them through a conversation.

The goal wasn’t simply automation.

It was reducing cognitive load.

The safest systems are often the easiest systems to use.

4. How will people adapt?

People change their behavior once a system exists.

Bad actors adapt even faster.

At a large consumer platform, we introduced a premium membership tier designed to signal quality and attract higher-intent users. Instead, it became a visible signal of wealth—attracting fraudsters who targeted those accounts, and driving away the very users the feature was meant to serve.

The feature worked.
Human adaptation changed the outcome.

Safety is not a one-time feature. It is an ongoing process of learning and adaptation.

Safety by Design

Traditional product development often follows a familiar sequence:

Build → Launch → Fix

Safety by Design reframes the development process.

It treats the four questions above not as post-launch concerns, but as design inputs—asked before a single line of code is written.

Anticipate → Design → Launch → Monitor

The goal is not to predict every possible outcome. That is impossible.

The goal is to identify foreseeable behavioral risks before development and design safeguards accordingly.

The surprises will still come. But they will come on top of a foundation, not instead of one.

One final thought

As AI systems become more capable, organizations will continue investing in governance, evaluation, and reliability—and they should.

But technical excellence alone is not enough.

Organizations must also understand how people will trust, misuse, adapt to, and ultimately shape AI systems after deployment.

The most important question is often not:

Will the model work?

It is:

How will people behave once it does?

In my experience, the organizations that succeed are not those that simply build better technology. They are the ones that design for real human behavior from the very beginning.

Technology creates capabilities.

Human behavior creates outcomes.

At Behavioral AI Lab, we believe this human layer is essential to building AI systems that are not only more capable, but also safer, more trustworthy, and more resilient.

Presentation slides from the Tokyo AI conference are available here: [View Slides]

If your organization is exploring AI safety, human-centered AI, or behavioral evaluation, we'd love to hear from you.


Hero Cover

The Missing Human Layer in AI Safety

Human-Centered AI

This article is based on a keynote I delivered at Tokyo AI on designing AI systems that remain safe after deployment.

Many discussions about AI safety focus on models, evaluations, governance, and reliability.

These are all essential.

Yet many AI systems fail for a different reason:

Human behavior.

The pattern I keep seeing

Across my career—in government platforms, consumer technology, and enterprise AI—I have observed the same failure mode, again and again.

The system works.
The outcome doesn’t.

A technically excellent AI system can still fail if the organization underestimates how people will use it, trust it, misuse it, or adapt to it.

Technology creates capabilities.
Human behavior determines outcomes.

Four questions every AI team should ask before launch

1. Will people trust it—appropriately?

Trust is not binary. Too little leads to rejection. Too much leads to overreliance.

The goal is calibrated trust: users who understand when to rely on the system and when not to. This is harder to design for than most teams expect.

2. What behaviors are we rewarding?

People respond to incentives, often in ways designers never anticipated.

At a ride-sharing platform, a dynamic pricing model successfully increased supply during high-demand periods. But during severe weather events, the same incentive structure created an unintended tension: drivers were financially motivated to stay on the road in unsafe conditions.

The technical system worked exactly as designed.

The behavioral outcome required careful redesign.

3. Will it create cognitive overload?

One of the most overlooked barriers to AI adoption is cognitive burden.

In a workplace safety reporting system, even well-designed processes can fail when reporting becomes too complex. Generative AI offered an opportunity to reverse the interaction: instead of asking employees to navigate lengthy forms, the system could guide them through a conversation.

The goal wasn’t simply automation.

It was reducing cognitive load.

The safest systems are often the easiest systems to use.

4. How will people adapt?

People change their behavior once a system exists.

Bad actors adapt even faster.

At a large consumer platform, we introduced a premium membership tier designed to signal quality and attract higher-intent users. Instead, it became a visible signal of wealth—attracting fraudsters who targeted those accounts, and driving away the very users the feature was meant to serve.

The feature worked.
Human adaptation changed the outcome.

Safety is not a one-time feature. It is an ongoing process of learning and adaptation.

Safety by Design

Traditional product development often follows a familiar sequence:

Build → Launch → Fix

Safety by Design reframes the development process.

It treats the four questions above not as post-launch concerns, but as design inputs—asked before a single line of code is written.

Anticipate → Design → Launch → Monitor

The goal is not to predict every possible outcome. That is impossible.

The goal is to identify foreseeable behavioral risks before development and design safeguards accordingly.

The surprises will still come. But they will come on top of a foundation, not instead of one.

One final thought

As AI systems become more capable, organizations will continue investing in governance, evaluation, and reliability—and they should.

But technical excellence alone is not enough.

Organizations must also understand how people will trust, misuse, adapt to, and ultimately shape AI systems after deployment.

The most important question is often not:

Will the model work?

It is:

How will people behave once it does?

In my experience, the organizations that succeed are not those that simply build better technology. They are the ones that design for real human behavior from the very beginning.

Technology creates capabilities.

Human behavior creates outcomes.

At Behavioral AI Lab, we believe this human layer is essential to building AI systems that are not only more capable, but also safer, more trustworthy, and more resilient.

Presentation slides from the Tokyo AI conference are available here: [View Slides]

If your organization is exploring AI safety, human-centered AI, or behavioral evaluation, we'd love to hear from you.