STORIES & INSIGHTS

Can AI help us hear communities more clearly and efficiently?

A competition held by UNHCR, UNICEF, the World Bank, the UN Behavioural Science Group, and Stanford University is testing how faithfully AI can reproduce real humanitarian survey responses, to help close the simulacra gap. 

Mary MacLennan (Senior Advisor on Behavioural Science and Lead of the United Nations Behavioural Science Group) and Andreas Haupt (Digital Fellow at Stanford University)
A man in a white coat stands next to a mother and child
A new competition tests whether the AI models being used to fill humanitarian data gaps are competent at capturing diversity and acknowledging uncertainty. Photo: UNICEF.

Further coauthored by: Sanmi Koyejo (Associate Professor of Computer Science at Stanford University), Ahmed Abukhashaba (DIMA Coordinator, UNHCR Regional Bureau for Middle East and North Africa), Rifat Zahir (Data Scientist, UNHCR Regional Bureau for Middle East and North Africa), and Rebeca Moreno Jiménez (Lead Data Scientist, UNHCR Innovation)

Effective and efficient humanitarian and development programming depends on high-quality data about the communities that programming is intended to support. Surveys are an essential way of gathering that data. A survey that identifies the barriers young people face in finding work can help a country office decide where employment opportunities are lacking, what support or training is needed, and which groups may be getting left behind. A survey that captures whether refugees intend to return home, what conditions would make return possible, and what barriers they face can help teams focus protection efforts, cash assistance, and reintegration support. 
 

The challenge is that collecting this information is often slow, costly, and, at times, hard to complete. People may hang up before a phone survey ends. They may skip certain questions. They may answer one round of a survey and disappear before the next. In crisis settings, the people most difficult to reach are often also those whose needs are most important to understand. This creates a partial picture — and partial pictures can lead to inefficient or poorly targeted programming.  
 

Artificial intelligence (AI) is increasingly being explored to fill these data gaps for hard-to-reach populations, though how well it works in humanitarian settings is still an open question. Are the AI models being used competent at capturing diversity and acknowledging uncertainty? Can the synthetic data they generate be relied upon as a complement to direct engagement with communities? SimulacraBench is a competition launching in August 2026, developed through a partnership involving UNHCR Innovation, which is designed to test just that. 
 

The promise and pitfalls of AI

Modern AI models can generate plausible responses to questions and are already being used to do so — for example, to forecast the macroeconomy, to simulate public-opinion survey respondents conditioned on real demographic profiles, and to estimate consumers' willingness to pay in market research. 
 

Statisticians use several related terms to refer to such AI-generated information. “Synthetic data” usually refers to artificially generated data designed to resemble real data. “Simulated respondents” are AI-generated stand-ins for people who might answer survey questions.  
 

The promise of AI in this space is efficiency and impact. If these tools worked well, they could help teams fill gaps in survey data, ask fewer and smarter questions, test programme ideas before spending scarce resources, and design more targeted interventions. They could support better decision-making in settings where time, money, and attention are limited. However, there is an important caveat. Plausible is not the same as accurate. 


An AI-generated response may sound reasonable while still being wrong. It may flatten differences between people, making everyone sound more like the average. It may reproduce the perspectives most visible online, rather than the perspectives of communities that are underrepresented in digital data. It may appear confident in precisely the cases where it should be uncertain. That is why it is essential to crosscheck the data generated by these AI models against real data, drawn from communities. 
 

The question is not whether AI can produce synthetic survey responses. It already can. The question is whether those responses are faithful enough, diverse enough, and well-calibrated enough to be useful in humanitarian and development decision-making. 
 

A faithful model should be able to predict answers it has never seen. It should not simply memorise public data. It should capture meaningful variation across people and contexts. It should also be able to say, in effect, “I am not sure,” when the evidence is weak. 

 

A clean test 

A joint initiative of UNHCR, UNICEF, the World Bank, the UN Behavioural Science Group, and Stanford University, SimulacraBench is running as part of NeurIPS 2026, one of the world’s largest AI research conferences. The idea is simple. Take responses from four real UN survey programmes that have never been released to the public — from UNICEF, UNHCR’s Middle East and North Africa Regional Bureau, and the World Bank, covering about tens of thousands of people across 15 countries — hide the answers, and ask participating teams to develop AI systems to predict these answers. 
 

The held-out survey data acts like an exam. Because the submitted models are graded on answers they have never seen, memorisation earns it nothing. The only way to score well is to capture something real about how people and communities vary. 
 

There is one design choice that makes the competition possible despite the sensitivity of the data: participants never get to see the underlying survey responses. Instead, teams submit their model’s code to the organisers, who run it on secure computers and return only a score. 
 

Models are also rewarded for being honestly uncertain. A system that says an answer is “70% likely” should be right about 70% of the time. That’s because, for decisions about youth employment, refugee returns, cash assistance, or programme design, knowing how much to trust a prediction matters just as much as the prediction itself. 

A woman in a UNHCR vest holds a clipboard out to a man
Responses from four real UN survey programmes, including from UNHCR's Middle East and North Africa Regional Bureau, are used as a benchmark for the competition. Photo: UNHCR


 

What better synthetic data could enable 

Suppose some teams do well, and we begin to see AI models that genuinely capture the diversity of a real population. What could change on the ground? 
 

Filling gaps without skewing the picture. When people skip questions or drop out of a survey, a faithful model could estimate what they may have said, so final results better reflect the whole community, not only the people who completed every question. 
 

Asking fewer, smarter questions. Instead of putting every question to every person, a good model could help identify the questions that would teach us the most about a given community. That could shorten a phone call or camp visit while preserving much of the value of the data. 
 

Designing more targeted interventions. If models can help identify which barriers matter most for which groups, programmes can be better tailored. For vaccination, that might mean distinguishing between communities where trust is the main issue and communities where access, timing, or transport are the larger barriers. 
 

Using resources more efficiently. Humanitarian and development organizations operate under serious resource constraints. Better modelling could help teams decide where additional data collection is most needed, where existing data is already strong enough, and where programme adjustments are most likely to matter. 
 

Testing ideas before spending on them. A reliable stand-in could let a team rehearse a survey, message, or campaign before committing to large-scale implementation. It could help spot confusing questions, identify subgroups that may respond differently, and improve design before resources are spent in the field. 
 

Every one of these uses only helps if the model is trustworthy. A confident-but-wrong stand-in does not save money. It launders a bad guess into an official-looking number and risks sending help to the wrong place. That is why measurement matters.  
 

SimulacraBench is designed around real problems the UN routinely faces. The goal is not only to rank AI systems, but to identify approaches that can be used in practice. The winning algorithms will be published under a non-commercial license, and hence will be freely available to the UN and other entities working on humanitarian and development challenges to help understand communities, target support, and make good use of limited resources. 
 

In that sense, the competition is not just about better AI. It is about a more responsive, effective, and efficient UN. 

 

The convergence of behavioural science and AI 

SimulacraBench is a technical challenge, but it is also a behavioural one. It asks whether AI can model not only what people say, but how stated beliefs, intentions, fears, needs, and behaviours vary across real populations and contexts. 
 

This is where behavioural science and AI are beginning to converge. Behavioural science brings understanding and careful measurement of human judgement, decision-making, context, actions and lived experience. AI brings new tools for prediction, simulation, and pattern detection at scale. Together, they open up new possibilities for understanding populations more efficiently, while also requiring much better tests of accuracy, uncertainty, and bias. 
 

AI cannot replace listening to communities, but it might be able to help identify where data is missing, where predictions are uncertain, and where further listening or human judgement is most needed. That is what SimulacraBench aims to explore.
 

The family deciding whether it is safe to return home, the young person trying to find work, the community waiting to be heard — these are the people behind the benchmark. 
 

SimulacraBench launches publicly in August 2026. The starter kit, tutorials, and rules are at simulacrabench.org. The competition is open to AI researchers, social scientists, and anyone interested in developing models to address these questions. We are particularly interested in submissions from the Global South. 
 

For those interested in real-world behavioural and humanitarian applications, the Behavioural Science and AI Exchange offers a forum to engage with emerging use cases, open questions, and potential collaborators. The community brings together researchers, practitioners, governments, civil society, and technology organisations working on AI, behaviour, and deployment in complex settings. Previous sessions have featured speakers from Anthropic, OpenAI, the World Bank, UNICEF, and leading academic experts. Sign up here