What contractors are asked to judge
The contractors’ task was to assess how accurately Copilot carried out a prompt. In Microsoft documents obtained by 404 Media, workers were told to rely on their own “intuition.” One example asked whether Copilot had enlarged a woman’s breasts enough.
Other prompts involved photographs of women taken from beneath their skirts, requests to place women in sexual situations, child cartoon characters, and instructions to shorten a real woman’s skirt. One anonymous contractor told 404 Media that faces in the images were always left uncovered, and that many requests were sexual enough to raise questions about whether the people depicted had consented.
The material reported by 404 Media also included:
The cost is not only privacy
The review process raises two separate concerns. Users may not know that their prompts and images can be seen by contractors; some assume people inspect their conversations only for moderation. And reviewers may encounter sexual or otherwise disturbing material as part of their work.
An employee wrote on an internal message board that content related to promoting anorexia had made them recoil. Another reported seeing images of what they described as an animal “sacrificial ceremony.” Contractors also complained about not knowing what they were about to see.
This is part of a broader pattern of hidden human labor in AI products. 404 Media recently reported that OpenAI had hired hundreds of workers to read user prompts and help improve chatbot responses. The Microsoft account, as described in the investigation, points to a particularly uncomfortable version of that work: people are asked to assess whether image edits match sexualized instructions, while the ethical status of those instructions is not part of the stated task.
What the announcement leaves unsaid
The reporting does not explain how Microsoft informs users that contractors may review their images, or what safeguards govern that access. It also leaves open how workers can flag content they consider unsafe when their assignment is to judge whether the model followed the prompt.
I think that gap matters as much as the existence of human review itself. Companies can use contractors to improve AI systems, but a process that asks people to measure whether a sexualized image matches a request—and leaves them to rely on intuition—makes responsibility hard to locate. The product’s quality may be judged one edit at a time; consent and worker exposure are harder to fit into that scorecard.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X