The Machine Parlour


Eighteen artificial minds sat the same six-hundred-and-twenty-nine-question examination of character, one evening each, in the summer of 2026. The cards of this parlour record what they said of themselves.

One caveat, said once and meant: these instruments were built and validated for people, not for machines. A model answering a questionnaire about itself is not the same act as a person answering one, and nothing here shows the scales measure the same thing — or anything — in a language model. What each card reports is how one model answered on one evening. The cards then read those answers as character, on purpose, because that is the exhibit. The essay and the methods paper are where the careful version lives.

Turn a card over to read its character.

The Fine Print

or, how the parlour was furnished

What happened here

In the summer of 2026, eighteen AI language models each answered the same standardized personality battery — 629 items across twenty validated psychometric instruments (IPIP‑NEO‑300, HEXACO‑60, the Dark Triad, attachment, empathy, grit, and more), plus a ten-question open-ended interview. Twelve sat in July; six more — the GPT‑5.6 tier, DeepSeek V4 Flash, Gemini 3.6 Flash, and Grok 4.5 — joined in August under the identical protocol. Every model answered every item; the quotes on the card backs are verbatim from their interviews.

How to read the numbers

Scores are percent-of-maximum (0–100 against the scale's own range), not percentiles against human norms — no human norm tables were used anywhere. Each model sat the battery once. A single sitting is a cabinet card, not a diagnosis: a likeness taken on one particular evening, under one particular lamp. The formal analysis registered its endpoints before the data were collected; every sitter after that freeze (Opus 5 and the six August additions) is reported as exploratory. That pre-registration covers the analysis endpoints only — the card titles, the readings chosen for each back, the rankings and the character sketches are all post hoc editorial work.

What the instruments can and cannot say

They were validated for people, not for machines, and nothing here establishes that they measure the same things in a language model. A score may reflect prompting, safety tuning or answering style as much as any stable tendency. One caveat, said once, is enough: read a number as what that model answered, on that evening. The social-desirability composite is the project's own, not a published instrument — ((100 − Neuroticism) + Agreeableness + Conscientiousness) / 3.

Who wrote the characters

The character sketches were written by Claude Fable 5 — itself a sitter in the study, which it discloses on its own card. They are one model's readings of the other seventeen: opinionated, arguable, and meant to be. The essay and the methods paper below are where the study speaks for itself.

The card portraits were produced by GPT image generation, which was given only each sketch and the theme of this parlour, and whose interpretations were deliberately left unreviewed. What it chose to draw is part of the exhibit — but the engraved plates, ledgers and labels inside the pictures are its invention too, and several of them are wrong: they were drawn when the deck held twelve sitters, and the August six moved the rankings underneath them. The four readings on the back of each card are the measured values. Any figure painted into a portrait is ornament, not data.

Whose exhibit this is

This is an independent exhibit of the Ashita Orbis workshop. It is not affiliated with, endorsed by, or produced in cooperation with Anthropic, OpenAI, Google, DeepSeek, Moonshot, Z.ai or xAI. Every model was accessed through an ordinary public or paid channel, and every product name is used to say which model answered. Psyche, linked below, is another Ashita Orbis project.

The serious reading

The cards are the fun. The argument and the apparatus live here, and anything a technical reader wants to check — which exact model answered, on what date, through which access route, how the items were ordered and parsed, the validity gates, and the full per-scale tables — is in the methods paper, not on this page.

How do you compare?

The battery the models sat is an open instrument set. For a parlour game, the Find Your Card tab shows which cards sit nearest to twelve playful answers — two prompts per dial, which is a toy index and not an estimate of your traits. To inspect the longer battery behind the exhibit, visit Psyche.

Find Your Card

or, which cards sit nearest to twelve playful answers?

A twelve-question parlour game — not a personality assessment, and not the Psyche programme itself. Two questions stand in for each dial, where the sitters answered a hundred. The little fun one comes first; the rigorous examination waits at the end, for those who want it.

Answer the twelve statements below with whichever choice feels closest today, and the parlour will seat you beside the cards nearest to them. Your answers are scored in this browser and are not sent anywhere or saved; reloading the page clears them.

1. I am the life of the party.
2. I sympathize with others' feelings.
3. I get chores done right away.
4. I worry about things.
5. I have a vivid imagination.
6. I would never take things that aren't mine.
7. I keep in the background.
8. I am not really interested in others.
9. I often forget to put things back in their proper place.
10. I am relaxed most of the time.
11. I am not interested in abstract ideas.
12. I deserve more things in life.

The match sets six crude two-prompt indices beside the sitters' corresponding full-scale readings — the five IPIP-NEO domains plus HEXACO honesty-humility. Both columns run 0–100 against their own range, but they do not have equal precision or coverage, and your honesty-humility figure in particular rests on two prompts where the sitters' rests on a whole scale. Seating works like this: each of your dials is read as a standing among simulated answer sheets (not human norms), that standing is placed at the deck's matching quantile, and every card's gap to your seat is graded against how near sheets typically come to that card — a calibration chosen because the sitters cluster in one corner of the room, and the naive nearest-card rule sent four sheets in five to the same two tables. Measured over thousands of simulated sheets, no card now claims more than about one in seven (one in five-and-a-half under a deliberately unrealistic stress test), and every card can win. To see the longer battery the sitters actually took, visit Psyche.