i
DATAIST
News · 2026-09-12

1,000 AI personas mark 38 of 50 words in the Triangle Task

@neuronium_ai @neuronium_ai

A researcher gave the Triangle Task, a creative-association test from a 2017 paper, to 1,000 AI personas. They marked 38 of the 50 words on average. The two human samples in the original study marked 7.47 and 13.84. The run belongs to a longer series on Langerian mindfulness, and its real question was narrower than that gap suggests: whether a persona's self-reported mindfulness score predicts how it behaves on a task nobody mentioned when the persona was created. It does. Higher scores on the Langer Mindfulness Scale came with more words marked, the same direction found in people.

Cover: 1,000 AI personas mark 38 of 50 words in the Triangle Task

A researcher gave the Triangle Task, a creative-association test from a 2017 paper, to 1,000 AI personas. They marked 38 of the 50 words on average. The two human samples in the original study marked 7.47 and 13.84. The run belongs to a longer series on Langerian mindfulness, and its real question was narrower than that gap suggests: whether a persona's self-reported mindfulness score predicts how it behaves on a task nobody mentioned when the persona was created. It does. Higher scores on the Langer Mindfulness Scale came with more words marked, the same direction found in people.

The background is what makes the setup unusual. In an earlier installment the same researcher ran the LMS — a self-report measure of how readily someone notices new things and responds flexibly to the situation in front of them — across 1,000 AI personas and found the scores bunched at the top of the distribution. Follow-up analysis pointed at RLHF: reinforcement learning from human feedback pushes models toward creativity, curiosity and the other dispositions the scale treats as components of mindfulness. The instrument was reading the tuning. This experiment was built to correct for that, generating personas with heterogeneous characteristics so the distribution would be normal from the start.

The series set out three questions: whether AI personas would naturally produce an even spread of LMS scores, whether those scores would predict behavior on novelty-detection tasks not disclosed at persona creation, and whether the scores would hold on retest. The first was settled early and in the negative — the skew is there, and the likely cause is how developers tune generative models. This run takes the second.

The wider bet behind all of it is that AI personas become a working tool in psychology. Trainee psychologists and psychiatrists can build characters with different personality types and mental states and practice therapeutic skills on them without risk to a real patient. Run in reverse, the persona plays the therapist and the person learns what being a client feels like. Personas can also stand in as participants in an online experiment, or take tests and surveys. The researcher is clear they do not replace human subjects. What he ran here is synthetic psychometric research: the participants were generated, not recruited.

The reason to reach for a task at all is that LMS is self-report — the participant answers questions about their own mental state. A creative task scores the behavior instead, which yields two independent measures: what the participant says about themselves, and what they produce. That is the design in "Utilizing a Creative Task to Assess Langerian Mindfulness," published by Katherine Berkovits, Francesco Pagnini, Deborah Phillips and Ellen Langer in Creativity Research Journal in 2017.

Their framing: Langerian mindfulness is an active process of noticing the new and responding flexibly to the current context; it means continuously creating new categories rather than staying inside the ones already formed, which is what mindless behavior amounts to; the Triangle Task was built to assess the core components; participants see a list of 50 words and mark any that could be related to "triangle"; and a mindful, creative person is expected to find more connections spontaneously than a less mindful one.

The list is: Triangle, pyramids, geometry, angle, Pythagorean theorem, tricycle, kite, square, side, scissors, love, stability, pencil, money, fire, unicorn, leaves, peace, breakfast, vowel, scarf, table, dice, power, T-shirt, integral, New York City, newspaper, computer, gears, snow, watch, picnic, soccer, infinity, momentum, violin, gravity, red, trickle, brush, gelatin, happiness, Jupiter, a ringlet, marble, octopus, pepper, mug, sheep. "Triangle" itself is in there as an attention check. A participant who fails to mark "triangle" as related to "triangle" has the rest of their answers re-examined.

The two human samples differed noticeably in demographics and in output. The first marked 7.47 words on average, the second 13.84 — close to double, which on the authors' reading made the second group more creative and plausibly more mindful. The relationship between Triangle Task performance and LMS score was positive and statistically significant, and that is what licensed treating the task as an independent behavioral route to the same trait.

Before running the personas, the researcher checked whether the model already knew the task. No sign of familiarity turned up, and an additional prompt was written to have the personas do it fresh rather than lean on prior exposure — the same confound that applies to a human who has taken a test before.

The result was 38 words out of 50, and the researcher offers four reasons for it. Language models are good at wordplay, semantic search and finding links between words, which is close to a base property of the architecture, so strong performance on word tasks is the expectation rather than the surprise. In preliminary checks the model marked all 50 almost every time; asked to justify each choice, it produced a workable connection for every word, drawing on geometry, architecture, music, sport and culture, and the researcher accepts that all 50 could in principle be defended — something that happened with human participants too, but rarely. A requirement was then added to the prompt that connections be convincingly defensible, which kept the model from sweeping up indirect or stretched associations and made it discriminate between degrees of relatedness, avoiding both the too-obvious and the very strange. And current models are tuned to please the user: in the preliminary test the model may have decided an impressive result was expected or would be welcome, which is the other reason it marked everything. The defensibility requirement, plus an instruction not to flatter, cut that back.

Across the 1,000 personas, stated mindfulness and task performance moved together — the higher the LMS score, the better the Triangle Task result — matching the human study and, on the researcher's reading, further evidence that the task is a usable measure of mindfulness. He closes the installment with Marcus Aurelius on the happiness of a life depending on the quality of its thoughts, and with the thought that using AI to study mindfulness may help explain how people become more mindful.

This is where I would be careful. Both numbers in that correlation come out of the same persona description. A persona written to be curious and open will report a high LMS score because that is what it has been told it is, and will mark more words for the same reason. Nothing in the setup separates the trait from the text that asserts it. In the human study the two measures are independent because the person filling in the scale has an interior life the researcher did not author. Here the researcher authored both. The correlation is real; what it shows is that a persona prompt propagates consistently into two different outputs. That is a much weaker claim than the Triangle Task measuring mindfulness in AI.

The 38 has a related problem. It is less a measurement than the setting the prompt happened to land on. Left alone, the model marked 50. Told to be defensible and not to flatter, it marked 38. The number moved because the instructions moved. The account does not report the spread around that mean, and with a ceiling at 50 and an average at 38 there is not much room left for personas to differ — which is precisely the room a correlation needs. The contamination check has the same softness: the word list has been in print in a journal since 2017, and what is reported is that no sign of familiarity appeared, not what would have counted as a sign.

The researcher's own next move is the one worth watching. He plans to look at other mindfulness tasks on the grounds that a good one has to test more than verbal fluency — it has to test the ability to leave the mindless mode. That diagnosis is right, and it points somewhere uncomfortable. The Triangle Task separates humans because most people, most of the time, do not bother to look for connections that are not already in front of them. A language model's default is to look for all of them. Any task that discriminates within an AI population will have to be one where mindlessness carries a cost the model can actually pay. The series is now, in effect, searching for a test these systems can fail.