

Why We're Building a Global Partnership Against AI-Enabled Manipulation
Behavioral AI
Today, we're announcing the launch of the Global Partnership for Detecting and Preventing AI-Enabled Manipulation and Scams — an initiative listed in the AI Dialogue Partnerships Hub, part of the United Nations Global Dialogue on AI Governance, with Behavioral AI Lab serving as the Implementing Entity. For the official announcement, you can read our press release here.
But today's launch is really the first step of something longer. I want to explain why we're building this partnership — and why we're building it the way we are.
The Scams Are Getting Smarter. So Should Our Defenses.
Throughout my career in behavioral science and Trust & Safety, I've watched online manipulation evolve from clumsy, easy-to-spot scripts into something far more sophisticated. Cryptocurrency investment fraud. Romance scams that unfold over months. Impersonation attacks that borrow a real person's voice, face, or writing style. According to the FBI's Internet Crime Complaint Center (IC3), losses related to cryptocurrency fraud exceeded USD 11 billion in 2025 — including USD 7.2 billion from cryptocurrency investment fraud alone.
When I worked in Trust & Safety — at Uber and Match Group — I kept seeing the same pattern repeat itself, no matter which platform or industry I was in. Technology kept changing. Human psychology did not. The channel shifted from email to dating apps to marketplaces to AI chat. The underlying playbook — build trust, manufacture urgency, isolate the target, escalate the ask — stayed remarkably constant.
Most detection systems still rely on keyword filtering and reactive moderation: flag the suspicious word, block the known pattern, react after the report comes in. That approach was never fully adequate, and it is becoming less adequate by the month. AI now allows bad actors to generate persuasive, personalized, multilingual manipulation at a scale and speed that keyword lists simply cannot keep up with.
The problem isn't just technical. It's behavioral.
What We Found
This initiative builds directly on research we've been doing for some time: an analysis of more than 15,000 scammer messages drawn from more than 150 documented cryptocurrency romance scam cases, spanning multiple languages and cultural contexts.
What we found reinforced something I've believed throughout my career in behavioral science: the signal that matters most isn't a suspicious word or a flagged phrase. It's the underlying psychology — how trust is built, how urgency is manufactured, how isolation is engineered, how persuasion escalates over time.
Focusing on those behavioral and psychological patterns, rather than surface-level keywords, gives us a far more robust and generalizable way to detect manipulation as it adapts.
That's the core thesis this partnership is built on.
Why Behavioral Science?
Behavioral science has spent decades studying how people make decisions, whom they trust, how emotions influence judgment, and how manipulation works. It is a mature, rigorous field with a long track record of explaining—and predicting—human behavior.
Yet surprisingly little of that knowledge has been incorporated into how we evaluate and safeguard AI systems. Most AI safety work today is led by machine learning researchers, which makes sense given the challenges the field has traditionally faced. But as AI systems increasingly operate at the boundary between technology and human judgment, that boundary is exactly where behavioral science belongs.
AI systems are ultimately designed to interact with people. If we want AI to be safe, it isn't enough to understand how models generate language—we also need to understand how humans interpret it, trust it, and act on it.
That gap, between what models do and how people actually respond, is where I've spent my career, and it's the gap Behavioral AI Lab exists to close.
We believe AI safety needs to evolve—and behavioral science must be part of that evolution.


Building the Foundation
Behavioral AI Lab leads this initiative as the Implementing Entity, but meaningful progress in AI safety cannot be achieved by any single organization alone. Shisa.AI contributes expertise in multilingual Small Language Models, safety-focused post-training, dataset development, and automated model evaluation. As the initiative grows, we expect additional research institutions, technology companies, financial organizations, and public-sector partners to join.
Together, we're building the foundation for a behavioral approach to AI safety. This work includes:
Behavioral Manipulation Taxonomy v1, with implementation guidance
Multilingual behavioral-risk datasets
Behavioral AI evaluation benchmarks
A Safety by Design framework for preventing AI-enabled manipulation
Capacity-building workshops across multiple countries
Pilot implementations with partner organizations
A global knowledge-sharing network connecting researchers, industry, and policymakers
These efforts are designed to move beyond academic research. They will support practical evaluation methods, implementation frameworks, and technologies that help organizations build, evaluate, and deploy AI systems more safely and responsibly.
Our long-term vision is to make behavioral intelligence a core component of AI safety—supporting AI developers in building more trustworthy systems, helping Trust & Safety teams identify emerging behavioral risks, enabling financial institutions to combat AI-enabled fraud, and providing policymakers with evidence-based approaches to AI governance.
Behavioral Science on the Global AI Agenda
One reason this partnership is particularly meaningful to us is that it reflects a broader shift in how AI safety is being understood.
For many years, AI safety discussions have focused primarily on technical challenges such as model alignment, robustness, and harmful outputs. Those challenges remain essential. But as AI systems increasingly influence human decisions, trust, and behavior, behavioral risks also need to become part of the conversation.
Being included in the United Nations Global Dialogue on AI Governance's Partnerships Hub is meaningful not simply because it recognizes our work, but because it signals that behavioral, human-centered approaches to AI safety are becoming an important part of the global AI governance agenda.
We hope this initiative helps bring researchers, developers, Trust & Safety professionals, policymakers, and industry leaders together to advance AI safety from both technical and behavioral perspectives.
What Comes Next
This is a multi-year initiative, and today is only the beginning. Over the coming months, we'll publish the Behavioral Manipulation Taxonomy, expand behavioral evaluation datasets and benchmarks, and begin the first capacity-building workshops with partner organizations across Asia-Pacific and beyond.
If your organization works in AI safety, Trust & Safety, financial fraud prevention, or AI governance—and shares our belief that behavioral science has a critical role to play—we'd love to hear from you.
AI systems will continue becoming more intelligent.
The question is whether they'll also become more trustworthy.
That's the challenge we're building this partnership to address.
Behavioral science has an essential role to play in making that happen.
Continue Reading
BLOG & INSIGHTS
Exploring the Human Side
of Artificial Intelligence.


Why We're Building a Global Partnership Against AI-Enabled Manipulation
Behavioral AI
Today, we're announcing the launch of the Global Partnership for Detecting and Preventing AI-Enabled Manipulation and Scams — an initiative listed in the AI Dialogue Partnerships Hub, part of the United Nations Global Dialogue on AI Governance, with Behavioral AI Lab serving as the Implementing Entity. For the official announcement, you can read our press release here.
But today's launch is really the first step of something longer. I want to explain why we're building this partnership — and why we're building it the way we are.
The Scams Are Getting Smarter. So Should Our Defenses.
Throughout my career in behavioral science and Trust & Safety, I've watched online manipulation evolve from clumsy, easy-to-spot scripts into something far more sophisticated. Cryptocurrency investment fraud. Romance scams that unfold over months. Impersonation attacks that borrow a real person's voice, face, or writing style. According to the FBI's Internet Crime Complaint Center (IC3), losses related to cryptocurrency fraud exceeded USD 11 billion in 2025 — including USD 7.2 billion from cryptocurrency investment fraud alone.
When I worked in Trust & Safety — at Uber and Match Group — I kept seeing the same pattern repeat itself, no matter which platform or industry I was in. Technology kept changing. Human psychology did not. The channel shifted from email to dating apps to marketplaces to AI chat. The underlying playbook — build trust, manufacture urgency, isolate the target, escalate the ask — stayed remarkably constant.
Most detection systems still rely on keyword filtering and reactive moderation: flag the suspicious word, block the known pattern, react after the report comes in. That approach was never fully adequate, and it is becoming less adequate by the month. AI now allows bad actors to generate persuasive, personalized, multilingual manipulation at a scale and speed that keyword lists simply cannot keep up with.
The problem isn't just technical. It's behavioral.
What We Found
This initiative builds directly on research we've been doing for some time: an analysis of more than 15,000 scammer messages drawn from more than 150 documented cryptocurrency romance scam cases, spanning multiple languages and cultural contexts.
What we found reinforced something I've believed throughout my career in behavioral science: the signal that matters most isn't a suspicious word or a flagged phrase. It's the underlying psychology — how trust is built, how urgency is manufactured, how isolation is engineered, how persuasion escalates over time.
Focusing on those behavioral and psychological patterns, rather than surface-level keywords, gives us a far more robust and generalizable way to detect manipulation as it adapts.
That's the core thesis this partnership is built on.
Why Behavioral Science?
Behavioral science has spent decades studying how people make decisions, whom they trust, how emotions influence judgment, and how manipulation works. It is a mature, rigorous field with a long track record of explaining—and predicting—human behavior.
Yet surprisingly little of that knowledge has been incorporated into how we evaluate and safeguard AI systems. Most AI safety work today is led by machine learning researchers, which makes sense given the challenges the field has traditionally faced. But as AI systems increasingly operate at the boundary between technology and human judgment, that boundary is exactly where behavioral science belongs.
AI systems are ultimately designed to interact with people. If we want AI to be safe, it isn't enough to understand how models generate language—we also need to understand how humans interpret it, trust it, and act on it.
That gap, between what models do and how people actually respond, is where I've spent my career, and it's the gap Behavioral AI Lab exists to close.
We believe AI safety needs to evolve—and behavioral science must be part of that evolution.


Building the Foundation
Behavioral AI Lab leads this initiative as the Implementing Entity, but meaningful progress in AI safety cannot be achieved by any single organization alone. Shisa.AI contributes expertise in multilingual Small Language Models, safety-focused post-training, dataset development, and automated model evaluation. As the initiative grows, we expect additional research institutions, technology companies, financial organizations, and public-sector partners to join.
Together, we're building the foundation for a behavioral approach to AI safety. This work includes:
Behavioral Manipulation Taxonomy v1, with implementation guidance
Multilingual behavioral-risk datasets
Behavioral AI evaluation benchmarks
A Safety by Design framework for preventing AI-enabled manipulation
Capacity-building workshops across multiple countries
Pilot implementations with partner organizations
A global knowledge-sharing network connecting researchers, industry, and policymakers
These efforts are designed to move beyond academic research. They will support practical evaluation methods, implementation frameworks, and technologies that help organizations build, evaluate, and deploy AI systems more safely and responsibly.
Our long-term vision is to make behavioral intelligence a core component of AI safety—supporting AI developers in building more trustworthy systems, helping Trust & Safety teams identify emerging behavioral risks, enabling financial institutions to combat AI-enabled fraud, and providing policymakers with evidence-based approaches to AI governance.
Behavioral Science on the Global AI Agenda
One reason this partnership is particularly meaningful to us is that it reflects a broader shift in how AI safety is being understood.
For many years, AI safety discussions have focused primarily on technical challenges such as model alignment, robustness, and harmful outputs. Those challenges remain essential. But as AI systems increasingly influence human decisions, trust, and behavior, behavioral risks also need to become part of the conversation.
Being included in the United Nations Global Dialogue on AI Governance's Partnerships Hub is meaningful not simply because it recognizes our work, but because it signals that behavioral, human-centered approaches to AI safety are becoming an important part of the global AI governance agenda.
We hope this initiative helps bring researchers, developers, Trust & Safety professionals, policymakers, and industry leaders together to advance AI safety from both technical and behavioral perspectives.
What Comes Next
This is a multi-year initiative, and today is only the beginning. Over the coming months, we'll publish the Behavioral Manipulation Taxonomy, expand behavioral evaluation datasets and benchmarks, and begin the first capacity-building workshops with partner organizations across Asia-Pacific and beyond.
If your organization works in AI safety, Trust & Safety, financial fraud prevention, or AI governance—and shares our belief that behavioral science has a critical role to play—we'd love to hear from you.
AI systems will continue becoming more intelligent.
The question is whether they'll also become more trustworthy.
That's the challenge we're building this partnership to address.
Behavioral science has an essential role to play in making that happen.
Continue Reading
BLOG & INSIGHTS
