Can ChatGPT outperform humans in faking a personality assessment while avoiding detection?

Large language models (LLMs), such as ChatGPT, have reshaped opportunities and challenges across various fields, including human resources (HR). Concerns have arisen about the potential for personality assessment manipulation using LLMs, posing a risk to the validity of these tools. This threat is a...

Full description

Bibliographic Details
Main Authors: Robie, C., Phillips, J., Bourdage, J. S., Christiansen, N. D., Dunlop, Patrick, Risavy, S. D., Speer, A.
Format: Journal Article
Published: Wiley-Blackwell 2025
Online Access:http://hdl.handle.net/20.500.11937/97883
_version_ 1848766330741719040
author Robie, C.
Phillips, J.
Bourdage, J. S.
Christiansen, N. D.
Dunlop, Patrick
Risavy, S. D.
Speer, A.
author_facet Robie, C.
Phillips, J.
Bourdage, J. S.
Christiansen, N. D.
Dunlop, Patrick
Risavy, S. D.
Speer, A.
author_sort Robie, C.
building Curtin Institutional Repository
collection Online Access
description Large language models (LLMs), such as ChatGPT, have reshaped opportunities and challenges across various fields, including human resources (HR). Concerns have arisen about the potential for personality assessment manipulation using LLMs, posing a risk to the validity of these tools. This threat is a reality: recent research suggests that many candidates are using AI to complete pre-hire assessments. This study addresses this problem by examining whether ChatGPT can outperform humans in faking personality assessments while avoiding detection. To explore this, two experiments were conducted focusing on assessing job-relevant traits, with and without coaching, and with two methods of identifying faking, specifically using an impression management (IM) measure and an overclaiming questionnaire (OCQ). For each study, we used responses from 100 working adults recruited via the Prolific platform, which were compared to 100 replications from ChatGPT. The results revealed that while ChatGPT showed some ability to manipulate assessments, without coaching it did not consistently outperform humans. Coaching had a minimal impact on reducing IM scores for either humans or ChatGPT, but reduced OCQ bias scores for ChatGPT. These findings highlight the limitations of current faking detection measures and emphasize the need for further research to refine methods for ensuring the integrity of personality assessments in HR, particularly as artificial intelligence becomes more available to candidates.
first_indexed 2025-11-14T11:49:26Z
format Journal Article
id curtin-20.500.11937-97883
institution Curtin University Malaysia
institution_category Local University
last_indexed 2025-11-14T11:49:26Z
publishDate 2025
publisher Wiley-Blackwell
recordtype eprints
repository_type Digital Repository
spelling curtin-20.500.11937-978832025-07-22T06:34:22Z Can ChatGPT outperform humans in faking a personality assessment while avoiding detection? Robie, C. Phillips, J. Bourdage, J. S. Christiansen, N. D. Dunlop, Patrick Risavy, S. D. Speer, A. Large language models (LLMs), such as ChatGPT, have reshaped opportunities and challenges across various fields, including human resources (HR). Concerns have arisen about the potential for personality assessment manipulation using LLMs, posing a risk to the validity of these tools. This threat is a reality: recent research suggests that many candidates are using AI to complete pre-hire assessments. This study addresses this problem by examining whether ChatGPT can outperform humans in faking personality assessments while avoiding detection. To explore this, two experiments were conducted focusing on assessing job-relevant traits, with and without coaching, and with two methods of identifying faking, specifically using an impression management (IM) measure and an overclaiming questionnaire (OCQ). For each study, we used responses from 100 working adults recruited via the Prolific platform, which were compared to 100 replications from ChatGPT. The results revealed that while ChatGPT showed some ability to manipulate assessments, without coaching it did not consistently outperform humans. Coaching had a minimal impact on reducing IM scores for either humans or ChatGPT, but reduced OCQ bias scores for ChatGPT. These findings highlight the limitations of current faking detection measures and emphasize the need for further research to refine methods for ensuring the integrity of personality assessments in HR, particularly as artificial intelligence becomes more available to candidates. 2025 Journal Article http://hdl.handle.net/20.500.11937/97883 10.1111/ijsa.70015 http://creativecommons.org/licenses/by-nc/4.0/ Wiley-Blackwell fulltext
spellingShingle Robie, C.
Phillips, J.
Bourdage, J. S.
Christiansen, N. D.
Dunlop, Patrick
Risavy, S. D.
Speer, A.
Can ChatGPT outperform humans in faking a personality assessment while avoiding detection?
title Can ChatGPT outperform humans in faking a personality assessment while avoiding detection?
title_full Can ChatGPT outperform humans in faking a personality assessment while avoiding detection?
title_fullStr Can ChatGPT outperform humans in faking a personality assessment while avoiding detection?
title_full_unstemmed Can ChatGPT outperform humans in faking a personality assessment while avoiding detection?
title_short Can ChatGPT outperform humans in faking a personality assessment while avoiding detection?
title_sort can chatgpt outperform humans in faking a personality assessment while avoiding detection?
url http://hdl.handle.net/20.500.11937/97883