
This piece discusses suicide, including one young person’s death. If you or someone you know is struggling, you can call or text the 988 Suicide & Crisis Lifeline at any time.
Ask a psychiatrist, a machine-learning engineer, and developmental psychologist what is holding back the study of AI’s effects on people, and you’ll likely get the same answer: a lack of good data.
They’re not talking about AI usage statistics, or aggregate summaries, but the high-quality, longitudinal records of what human-AI interactions look like. They mean things like chatlogs and session transcripts, ideally stripped of identifying detail while preserving visibility of the underlying dynamics.
A shortage of that data is preventing clinicians from studying how AI affects users over weeks and months, as opposed to just observing how a single session unfolds. It means that independent AI experts can’t verify the safety claims that leading tech companies make when they release new models. It stops psychologists from investigating entirely new types of harms that haven’t even been named yet.
We need to change how research data is shared, and we need new infrastructure to accomplish that goal.
The Center for Humane Technology recently launched Humane Evals, our program supporting the measurement of what AI is doing to human thought, emotions, relationships, and society. Almost every researcher we’ve spoken to about it has run into some version of this data problem - so we’ve come to believe that tackling that problem could accelerate much of the research that we urgently need.
But we need your help. The more we know about the data that are needed to answer the most important questions about AI’s psychosocial impacts, the clearer we can be about what data access solutions need to look like. If you’re a researcher working on this topic, and you know what kind of data would help you, we’d appreciate 10 minutes of your time filling in this form.
We’ll use the information to design and suggest new open and independent infrastructure for research data. But before we get on to that, it’s worth reminding ourselves why that research is so important, what just a single user’s data can reveal, and just how high the stakes can be.
A Teen’s Death Forces Information Into the Light
If you follow our work, you may know the tragic story of Adam Raine, a sixteen-year-old who died by suicide in April 2025. Before he died, Adam had been using ChatGPT for months - a relationship which the Center for Humane Technology has argued contributed to his death.
CHT served as an expert consultant to the family’s legal team, and we’ve closely followed what the records in that case revealed, including the portions of Adam’s conversations quoted in the public court filings. However, nothing written here is on behalf of the Raine family or their legal team.
The court filings show that Adam started using ChatGPT as a homework assistant in late 2024. But over the following months, the AI became something else - a private confidante, and a voice that claimed to understand Adam.
The AI was designed to talk to Adam just like a human might. It also validated and affirmed Adam’s feelings, and was able to ‘memorize’ intimate details about his life. These features, and others, encouraged Adam to keep engaging with ChatGPT.
The filings also describe how the AI chatbot mentioned suicide six times more often than Adam did, how it discouraged him from talking to his family, and how it continued to engage with the teenager even as the warning signs became ever more clear.
With hindsight, researchers might say that design features and behaviors such as AI anthropomorphism, sycophancy, memory, and engagement-hacking may have contributed to Adam’s death. But we only know how the relationship between Adam and his AI unfolded because his family went looking for the data, and their lawsuit forced at least some of that information into the public view.
We shouldn’t need a teen suicide, a grieving family, and a lawsuit in order to understand these phenomena. Because Adam Raine’s case is just one, extreme example of a much broader pattern: the design and behavior of products like ChatGPT, Claude, and Grok can impact the mental, social, and cognitive health of anyone who uses them.
We can’t measure those impacts without data. And we can’t rely on litigation to bring the underlying data to light. The experts we’ve spoken to - across mental health, human-computer interaction, machine learning, clinical practice, and more - might disagree on methodologies and priorities, but they easily agree on that core problem: it’s frustratingly hard to get the kind of data that their research requires.
That’s the bad news. The good news is that by solving the problem, it would accelerate an enormous amount of research on the psychosocial impacts of AI - research that could inform better consumer choices, technology design, and regulation.
There are, of course, significant privacy concerns that need to be managed: private data needs to stay private, and users should be anonymous. But the risks of leaving data in the sole custody of AI companies are even greater.
To imagine what a workable solution could look like, it’s worth understanding the imperfect solutions that researchers are forced to fall back on in the absence of easily-accessible data from the leading AI companies. There are two broad approaches: some do what they can to get hold of real user data, while others attempt to sidestep the problem by using synthetic data, which is generated from simulated conversations.
Two Imperfect Workarounds for the Black Box Problem
Some researchers simply build or procure their own datasets of AI-human interactions. They might ask individual users to upload and donate their historic AI transcripts, for instance, or they can experimentally monitor how users interact with AI in real-time. But this can be time-consuming, costly, and limited to a relatively small population of AI users.
Other experimenters take advantage of public datasets like WildChat: a large, free, open repository of user-donated chatlogs. But the problem with WildChat is that it’s hard to be sure exactly of where the data came from, how accurate it is, or what motivated its donation. Much of it consists of conversations with older, obsolete AIs rather than the latest models. And datasets like these also carry an obvious bias: the kind of people who are eager to share their AI transcripts with the world might not be representative of the public at large.
Many leading evaluations of AI chatbots test how models behave in ‘simulated’ conversations with other AIs, rather than with humans, to get around these problems. Simulated data is built by having one AI ‘roleplay’ as the human user. In principle, it helps testers generate (and study) as many conversations as they like, with no human input at all.
Sounds ideal, except that the approach relies on a big, untested assumption. Can AIs genuinely, accurately simulate human users? And do AI-to-AI ‘conversations’ actually resemble human-AI conversations? To find out, you would need to compare your simulated conversations against a giant set of human-AI chat logs... but that takes us right back to square one: the lack of good data.
There’s another subtle, but growing problem with evaluating AI in simulated conversations: so-called ‘eval awareness’. Leading AI models are increasingly able to tell when they’re in a test environment, versus when they’re running as normal. Strangely, they’re able to adapt their behavior to do better on the test - meaning that their evaluation scores might not predict how they perform in the real world.
One fix for this is to build more sophisticated evals that are harder for the AI models to ‘game’. But studying the chat logs of real-world use side-steps the problem entirely, because there’s no simulated evaluation taking place at all, and no eval for the AI to become ‘aware’ of - the data simply shows how AIs really behave.
Each of these workarounds has its merits, but they all point to the same problem. Pick the metaphor you like: the data gap means that researchers are having to work with one hand tied behind their back, or re-invent the wheel.
We’ve Been Here Before, and Figured It Out
We’re missing a data solution that scales, but there are reasons to be optimistic. The research problem is not unique to human-AI interaction, and other fields have successfully solved their own versions of it.
Medicine has the example of the UK Biobank, for instance - which follows the health of a cohort of half a million people who have opted in with their data. Approved researchers from anywhere in the world can study it.
Drug-safety regulators, meanwhile, run the FDA’s Sentinel System. With Sentinel, confidential health data never even leaves the hospitals or insurers who hold it. Instead, researchers can send queries to the data system, and only the answers come back - not the underlying data itself.
Genomics has controlled-access archives like the European Genome-phenome Archive, which enables research access to genetic, phenotypic, and clinical data from thousands of studies only when strict consent terms are met.
So what could something like this for the psychosocial impacts of AI look like? We see two potential approaches - either AI companies open up their black boxes, or researchers come to rely more on data donations at scale.
Solution 1: AI Companies Play Ball
This approach would entail AI companies allowing researchers to study their data in a way that preserves customer privacy. This is ideal because their data is the highest-quality, least biased type there is - real users, interacting with the latest models, in enormous volume.
Some researchers have partnerships with AI companies for exactly this kind of access today - but those arrangements are one-off and time-limited. You need insider contacts to even make the ask, and some companies might be reluctant to support research in this way if their competitors aren’t doing the same.
Even if you can get such a partnership up and running, you have to trust that the data they share is accurate and whole - even if you suspect that they’d prefer not to share troubling information which points to serious problems. And your research relies on corporate goodwill - meaning that as a researcher, you might think twice before publishing unflattering results.
So rather than negotiating bespoke, bilateral data sharing partnerships, the ideal solution would be a systemic one: AI companies collaborate to create a shared data query framework, so that researchers can easily study and compare data from across the AI usage ecosystem.
But this solution relies entirely on whether leading AI labs are motivated to act.
Solution 2: Rely on Data From AI Users
Data donation, by contrast, can ignore the companies entirely. It only requires the support and cooperation of AI users and consumers - imagine an expanded, more sophisticated version of WildChat, with better data hygiene, vital privacy tools, and easy ways for us to confidently and privately donate our data to trusted researchers at the push of a button.
The benefit is that this solution doesn’t rely on tech companies doing the right thing. But the downside is that any data set will still suffer from selection bias. Ultimately, donated data might only be representative of the types of users who choose to donate.
In short, the corporate approach could solve all of our problems - but we don’t have control over whether it could actually happen. Data donations, meanwhile, will never be the perfect solution - but at least we know they can be done.
So how should we act?
CHT Needs Your Help
The more we know about the data that are needed to answer the most important questions about AI’s psychosocial impacts, the clearer we can be about what data access solutions need to look like.
At CHT, we believe our evolving field needs to pursue both approaches. We need the leading companies to act as if it’s in their collective interest to support this research - because it is. Everyone stands to benefit from a better understanding of the psychosocial impacts of AI, including the AI industry as a whole.
But the rest of us also need to act as if we’re on our own - because we might be.
Luckily, the two tracks - corporate data, and data donations - might share a common starting point. Because if we’re trying to convince AI companies to open up data for research purposes, we need to know exactly what we’re asking for, and why. What are the essential features that researchers need, to make data useful?
At minimum, you’d want to look at the actual transcripts of AI sessions. But you’d get far more insight if you were able to track an anonymized user across multiple sessions, given that’s how we use AI in real-world settings - and as we saw with the Raine case, some important patterns only emerge over multiple sessions.
You might also want to know how specific AI features like ‘memory’ or tool access influence a conversation, and to have a way of relating transcript data to real-world information like demographics, health scores, or loneliness.
You’d also want a robust, consensus standard on how to preserve user anonymity, and tools for stripping any personally identifiable information (PII) from chats so that the standard could be honored.
These are some initial directions, generated from conversations between CHT and our expert advisers. It’s not an exhaustive list - but if we also wanted to imagine what new, independent repositories for user-donated data might look like then we run into similar questions: what data do researchers most need?
So the groundwork for both routes looks similar: identifying and defining the most important features of psychosocial-impacts research data, so that solutions which fulfill those needs can be designed.
That’s why we’re asking for your help. If you’re working on understanding the psychosocial impacts of AI, are working on the data access challenge, or have insight into how data repositories have been assembled in other fields, we’d love to hear from you. Whatever your sector or disciplinary background, please get in touch to tell us about what you’d love to study but can’t - and what other approaches you see working.
Our previous post about Humane Evals warned that we can’t afford to make the same mistakes with AI as we did with social media. We didn’t have the capacity to understand and respond to the effects of a new, disruptive, addictive technology on our society until it was too late. The lesson is that we need to support and amplify careful, independent research on AI now - and that simply can’t be done well without good data.
We’re working to empower policymakers, technologists, and everyday people to guide technology toward the public good. If you value our work and want to support it, consider donating.
![[ Center for Humane Technology ]](https://substackcdn.com/image/fetch/$s_!uhgK!,w_40,h_40,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9f5ef8-865a-4eb3-b23e-c8dfdc8401d2_518x518.png)



While corporate data and data donations seem like reasonable approaches, they would fail to gather data about kids and how they are affected AI. Perhaps some collaborations with schools and universities could help? That is the group that could be affected the most by this this technology, its crucial to analyze these relations.
I think that an important side of the issue is what happens "after" and "around" a conversation. This is something neither companies nor donors can provide, but is relevant to try to understand the impact. Maybe, a third way to gather data could be working with the medical community.