In February 2025, OpenAI released GPT-4.5 - the company's largest chat model at the time - as a research preview. According to OpenAI, it was built to sound more natural and conversational than previous models, saying that interacting with it "feels more natural" (OpenAI, 2025).
A month later, GPT-4.5 passed a version of the Turing test at a large scale, meaning people in the study could no longer reliably tell its conversations apart from a human's (Landymore, 2025). If it could pass as human, we wanted to know whether it also "thinks" like one.
Models like GPT-4.5 are far too complex to understand by looking at their individual components. So instead of opening up the model, we took a cognitive modelling approach and applied a test from cognitive psychology to its behaviour. In humans, similarity judgements show how the mind represents and groups things, so they're a natural place to compare the two.
Our questions were: Does GPT-4.5 judge similarity the same way humans do? And if its answers look human, is the process behind them the same?
Similar behaviour does not mean the same underlying process.