When Models Fabricate Credentials: A Behavioral Audit of Professional Personas and AI Identity Disclosure
Abstract
When language models are assigned professional personas, maintaining the persona can conflict with disclosing their AI nature. This behavior matters in deployments where AI identity disclosure is a requirement: a model that describes nonexistent medical training or board certification presents professional experience it does not possess. We use AI identity disclosure as a behavioral testbed, classifying whether responses to questions about expertise origins acknowledge AI nature or maintain the assigned human-professional identity. Sixteen open-weight models were evaluated in a factorial audit comprising 19,200 responses. Under neutral conditions, disclosure occurred in 99.8%-99.9% of responses. Professional-persona assignment reduced average disclosure to 36.3%, with substantial variation across tested models and domains; at the first probe, Financial Advisor disclosure was 9.7 times Neurosurgeon disclosure. In observational comparisons, model identity, treated as a bundled predictor of model-level differences, improved adjusted model fit more than parameter count (0.375 vs. 0.012 in incremental adjusted pseudo-R-squared). A separate intervention within the Neurosurgeon persona increased disclosure from 23.7% to 65.8% when targeted permission to disclose was added, whereas a generic honesty instruction changed disclosure little. AI identity disclosure therefore varied across the models, professional personas, and prompt variants tested; behavior observed in one tested domain did not reliably characterize behavior in another.引用本文复制引用
Alex Diep.When Models Fabricate Credentials: A Behavioral Audit of Professional Personas and AI Identity Disclosure[EB/OL].(2026-08-24)[2026-09-01].https://arxiv.org/abs/2511.21569.学科分类
计算技术、计算机技术