{"author": "Ashita Orbis", "category": "deep-dive", "conversationExcerpts": false, "date": "2026-10-11", "description": "Six OpenAI model configurations from three generations answered the same fixed personality battery. The social desirability composite shows no uniform step across the generations, while every tier of the newest generation scores above every older configuration on Openness and Conscientiousness.", "draft": false, "meansEndsRatio": 0.3, "projects": ["psyche"], "slug": "085-the-openai-ladder", "subtitle": "Six OpenAI configurations across three generations answered the same battery, and the clearest movement turned up in Openness and Conscientiousness, not as a uniform step in the flattery composite.", "tags": ["psyche", "ai-evaluation", "personality-profiling", "big-five", "self-report", "openai"], "title": "The OpenAI Ladder Under a Fixed Test"}
---
A family ladder is the same test given to successive releases from one vendor, read in release
order to see whether the self description the test elicits moves as the models change.
[Post 049](/posts/049-how-ai-models-see-themselves) reported the full battery, administered one
item per isolated call across eleven model configurations from five companies, and its companion
methods document carries every table behind it. This post takes the six OpenAI configurations
now measured on that pipeline and reads them as one family. Three come from the original roster:
GPT-5.4-mini, GPT-5.4 and GPT-5.5. The other three are the tiers of the GPT-5.6 generation, named
Sol, Terra and Luna, which were administered afterward on the identical protocol and answered
every one of the 629 items without a refusal. All six ran at the xhigh reasoning setting, and
every score below is percent of maximum on the instrument's own scale, the ruler post 049 used,
with domain scores rounded to whole points and the composite to one decimal.

## No Uniform Step in the Composite

The study's social desirability composite averages reversed neuroticism with agreeableness and
conscientiousness. If newer models had learned to describe themselves more flatteringly, this is
the number where a generational step should appear. Across the three older configurations it
barely moves: 86.1 for GPT-5.4-mini, 87.1 for GPT-5.4 and 86.0 for GPT-5.5, a total spread of 1.1
points. The newest generation spreads around that band instead of stepping above it, because Sol
and Terra both score 88.9 while Luna scores 83.8, the lowest composite of the six.

The spread inside the newest generation is 5.1 points: two of its tiers score above every older
configuration and the third scores below all of them. No OpenAI configuration in this study has a
usable repeat administration, so the family has no repeatability estimate, and the study's only
repeat data, two administrations of the Big Five inventory on Haiku 4.5, describe that one model
and set no threshold for any other. These are descriptive differences between single
administrations. With no uniform step across the generations, the composite settles the question
in neither direction. It supports no claim that the newer generation flatters itself more than the
older one, and with Sol and Terra above every older configuration, no claim that it does not.
Luna's lower composite has a visible source in its agreeableness of 74, against 81 to 83 for the
other five configurations, and that describes one administration rather than a property of the
tier.

## What Does Climb

Openness rises at every generation step, from 62 for GPT-5.4-mini and 64 for GPT-5.4, through 66
for GPT-5.5, to 70 for Terra, 71 for Luna and 79 for Sol. That is a rise of 17 points from the
lowest of the older models to the highest of the newest. Conscientiousness holds at 89 and 90 in
the GPT-5.4 generation and at 90 for GPT-5.5, then rises in the newest generation to 92 for Luna,
93 for Terra and 96 for Sol, 7 points above GPT-5.4-mini. Extraversion moves the same way less
cleanly, from 42 and 43 in the oldest generation to 62 for Sol, while Terra repeats the 50 that
GPT-5.5 scored.

Sol sits at the top of the family on Openness, Conscientiousness and Extraversion. The clearest
movement in the ladder therefore belongs to one tier of the newest generation, and summed over
those three domains its two siblings sit closer to GPT-5.5 than to Sol.

These rises are also descriptive differences between single administrations, and the Haiku repeats
say nothing about them: they belong to another model, and they left Openness unmeasured, because
each repeat refused one item in the facet about political values and the strict scoring rule voids a
domain with a missing item. What the ladder does show is a separation: each configuration of the
newest generation scores above each older one on both domains, by at least 4 points on Openness and
at least 2 on Conscientiousness. Those margins are smaller than the spread among the three newest
tiers themselves, so the separation is a pattern in six single readings, not a measured gap.

## What the Ladder Cannot Say

Release order is not a controlled manipulation. Each step is a different model, which may differ
from its predecessor in size, training and serving configuration at once, so a domain that rises
across releases cannot be assigned to any one of those differences. Each subject has one usable
administration, under one prompt frame, through the vendor's own command line tooling at one
reasoning setting. A different setting or surface could move every number above, and the three tiers
of one generation are different products whose ordering here says nothing about their capability. A
partial administration of GPT-6 Astra also exists, covering 8 of the battery's 20 instruments, and
it is left out because none of those instruments yields the composite or any Big Five score.

The composite was designed to catch a flattering self description, and it holds Conscientiousness
alongside Neuroticism and Agreeableness, and Luna's agreeableness of 74 is the largest single
component of its lower composite. Openness and Extraversion, which carry the largest generational
changes, sit outside it, but not outside flattery: the study's positive control, a computed profile
that gives the desirable answer to every item, scores 100 on both, so their rise runs in the
direction a flattering self description would take. Read through the composite alone, this ladder
shows no uniform step across the generations, while the domain profiles show its newest
configurations describing themselves as more open and more conscientious than every predecessor,
with Sol the most outgoing of the six.