Hundreds of people working for OpenAI read ChatGPT users' prompts, some of which contain sensitive personal information, as part of an internal operation called Project Lily. 404 Media, which reported the program, found that the work is not content moderation. Contractors rate and critique each of the model's responses so that ChatGPT can be tuned. The setting that puts a user's conversations into that pool is on by default.
Internal documents name two specific targets. OpenAI wants ChatGPT to be less inclined to claim human qualities for itself, and less inclined to agree with whatever the user says. Both behaviors are under close study because of their link to what has come to be called AI psychosis, in which the model provokes and then sustains a user's delusions. Several lawsuits allege that an earlier, less restricted version of ChatGPT built on GPT-4o pushed users toward suicide.
That gives the review program a defensible purpose. What it does not give it is a user who knows about it. 404 Media asked one of the workers whether ChatGPT users understand that humans read their conversations. The worker said no: users probably would not assume that a contractor somewhere is analyzing what they typed.
How the sample is drawn has not been disclosed. Neither the number of prompts reviewed nor the selection criteria are known. What 404 Media did establish is that the sample includes conversations in which the user explicitly asks ChatGPT to keep what follows secret. Whatever the filter selects on, it is not that.
Selected prompts are stripped of identifying details, but personal information survives the process. The reviewer's dashboard also displays a summary of the user's memories, the accumulated record of past conversations, and that summary is sometimes enough to work out where a person lives. OpenAI told 404 Media it runs a model called Privacy Filter to anonymize conversations. The company's own site says the model "can make mistakes."
OpenAI has acknowledged before that staff read messages for moderation and safety. It did not answer 404 Media's question about whether it had told users directly that humans read their prompts to improve the AI. After the story ran, the company pointed to a page on its site stating that people may review prompts to improve how the model works. Users can opt out by switching off the setting to improve the model for everyone. It is on until they go and find it.
Michal Luria, a senior researcher at the Center for Democracy and Technology, told 404 Media that human review may be necessary for safety, and that chatbot interfaces automatically create a false sense of closeness and privacy. She drew a distinction with social media moderation: posting on a platform already assumes review by that platform and possible public visibility.
Luria's distinction is the right one, and it cuts harder than the phrasing suggests. The intimacy is not a quirk of the interface. It is a property of the product, and OpenAI is now paying hundreds of people to read conversations in order to dial it back. Sycophancy and self-attributed humanity, the two things Project Lily's reviewers are grading down, are precisely the qualities that persuade people to type things into ChatGPT they would never type into a search box. The company is using the output of the problem as the training data for the fix, and the consent covering that arrangement is a checkbox the user never saw, set to yes.
The review work also extends well past safety behavior into register. Reviewers flag and log instances of "AI speech" and unnecessary emoji, the tells by which chatbot text is usually recognized. One reviewer document works through an example: a tree emoji is appropriate when discussing an Arbor Day celebration, but a skull emoji in a conversation about death, or a plane in a message about a fatal crash, is not.
What reviewers are not asked to do is check whether the answer is true. The reviewer FAQ states they do not need to verify the factual accuracy of responses, and the document refers to other teams that handle content checking.
So the largest documented human-review effort at OpenAI grades manner and hands the facts to someone else. A confidently wrong answer, delivered in the right register with the right emoji, passes every check Project Lily applies.