381 papers that use language models as human stand-ins.
This is the complete corpus behind the field review: every paper, no curation beyond the documented sweep. These are the market's numbers, not ours. Slice it however you want; the table is the same data as the CSV the paper is built on.
Provenance: papers coded from title and abstract against a fixed rubric by a language model (method in the working paper, PDF). The funnel-stage and exclusion-reason columns (which papers reached 104, 24, and 4, and why the rest fell out) are not yet in this table: the screening records are being prepared for release, and the gap is stated here rather than papered over.
| # | Title | Authors | Year | Venue | Cited | Theme |
|---|---|---|---|---|---|---|
| 1 | The Emergence of Social Science of Large Language Models | Xiao Jia & Zhanzhan Zhao | 2025 | arXiv.org | 1 | Generative-agent societies & multi-agent simulation |
| 2 | Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research | Emma Rose Madden | 2025 | Journal of Social Computing | 5 | Generative-agent societies & multi-agent simulation |
| 3 | Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations? | Yue Huang et al. | 2024 | arXiv.org | 33 | Generative-agent societies & multi-agent simulation |
| 4 | LLM-Assisted Replication for Quantitative Social Science | So Kubota et al. | 2026 | arXiv.org | 1 | AI as substitute for human annotators/subjects |
| 5 | Finetuning LLMs for Human Behavior Prediction in Social Science Experiments | Akaash Kolluri et al. | 2025 | Conference on Empirical Methods in Natural Language Processing | 27 | Cognitive/economic experiments (replication) |
| 6 | Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models | Naoki Egami et al. | 2023 | Neural Information Processing Systems | 60 | Cognitive/economic experiments (replication) |
| 7 | Evaluation is all you need. Prompting Generative Large Language Models for Annotation Tasks in the Social Sciences. A Primer using Open Models | Maximilian Weber & Merle Reichardt | 2023 | arXiv.org | 20 | AI as substitute for human annotators/subjects |
| 8 | LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research | Yi Yang et al. | 2024 | arXiv.org | 7 | Generative-agent societies & multi-agent simulation |
| 9 | Arti-"fickle" Intelligence: Using LLMs as a Tool for Inference in the Political and Social Sciences | Lisa P. Argyle et al. | 2025 | Nature Computational Science | 3 | Cognitive/economic experiments (replication) |
| 10 | Social Simulations with Large Language Model Risk Utopian Illusion | Ning Bian et al. | 2025 | arXiv.org | 0 | Generative-agent societies & multi-agent simulation |
| 11 | Leveraging LLM-based agents for social science research: insights from citation network simulations | Jiarui Ji et al. | 2025 | Humanities and Social Sciences Communications | 2 | Generative-agent societies & multi-agent simulation |
| 12 | A Methodological Guide on Using Large Language Models for Reproducible Text Annotation in the Social Sciences and Humanities with Python and R | Qixiang Fang et al. | 2026 | 0 | Cognitive/economic experiments (replication) | |
| 13 | Simulating Public Administration Crisis: A Novel Generative Agent-Based Simulation System to Lower Technology Barriers in Social Science Research | Bushi Xiao et al. | 2023 | arXiv.org | 23 | Generative-agent societies & multi-agent simulation |
| 14 | Exploring Large Language Model Agents for Piloting Social Experiments | Jinghua Piao et al. | 2025 | arXiv.org | 2 | Generative-agent societies & multi-agent simulation |
| 15 | Beyond Static Responses: Multi-Agent LLM Systems as a New Paradigm for Social Science Research | Jennifer Haase & Sebastian Pokutta | 2025 | arXiv.org | 13 | Generative-agent societies & multi-agent simulation |
| 16 | Value-Based Large Language Model Agent Simulation for Mutual Evaluation of Trust and Interpersonal Closeness | Yuki Sakamoto et al. | 2025 | Scientific Reports | 2 | Generative-agent societies & multi-agent simulation |
| 17 | Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses | Jens Rupprecht et al. | 2025 | Proceedings of the Seventh Workshop on Natural Language Processing and Computational Social Science | 5 | Decision-making, biases & trust |
| 18 | Can Large Language Model Agents Simulate Human Trust Behavior? | Chengxing Xie et al. | 2024 | Neural Information Processing Systems | 74 | Decision-making, biases & trust |
| 19 | Epitome: Pioneering an Experimental Platform for AI-Social Science Integration | Jingjing Qu et al. | 2025 | arXiv.org | 0 | AI as substitute for human annotators/subjects |
| 20 | Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels | Nicholas Pangakis & Samuel Wolken | 2024 | NLPCSS | 56 | AI as substitute for human annotators/subjects |
| 21 | Agentic Society: Merging skeleton from real world and texture from Large Language Model | Yuqi Bai et al. | 2024 | arXiv.org | 1 | Generative-agent societies & multi-agent simulation |
| 22 | Applying Psychometrics to Large Language Model Simulated Populations: Recreating the HEXACO Personality Inventory Experiment with Generative Agents | Sarah Mercer et al. | 2025 | arXiv.org | 1 | Personality & psychological profiling of LLMs |
| 23 | Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches | Syed Mehtab Hussain Shah et al. | 2026 | Companion Proceedings of the ACM Web Conference 2026 | 3 | Generative-agent societies & multi-agent simulation |
| 24 | Synthetic Founders: AI-Generated Social Simulations for Startup Validation Research in Computational Social Science | Jorn K. Teutloff | 2025 | arXiv.org | 0 | AI as substitute for human annotators/subjects |
| 25 | Can Large Language Models Serve as Rational Players in Game Theory? A Systematic Analysis | Caoyun Fan et al. | 2023 | AAAI Conference on Artificial Intelligence | 138 | Cognitive/economic experiments (replication) |
| 26 | What Is Actually Being Annotated? Inter-Prompt Reliability as a Measurement Problem in LLM-Based Social Science Labeling | Jingyuan Liu | 2026 | arXiv.org | 0 | AI as substitute for human annotators/subjects |
| 27 | Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation | Joachim Baumann et al. | 2025 | arXiv.org | 36 | Cognitive/economic experiments (replication) |
| 28 | Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations | Yong Cao et al. | 2025 | North American Chapter of the Association for Computational Linguistics | 49 | Public opinion & political representation |
| 29 | Diminished Diversity-of-Thought in a Standard Large Language Model | Peter S. Park et al. | 2023 | Behavior Research Methods | 88 | Cognitive/economic experiments (replication) |
| 30 | S$^3$: Social-network Simulation System with Large Language Model-Empowered Agents | Chen Gao et al. | 2023 | Social Science Research Network | 324 | Generative-agent societies & multi-agent simulation |
| 31 | Evaluating Large Language Models with Psychometrics | Yuan Li et al. | 2024 | 28 | Personality & psychological profiling of LLMs | |
| 32 | Investigating and Extending Homans' Social Exchange Theory with Large Language Model based Agents | Lei Wang et al. | 2025 | Annual Meeting of the Association for Computational Linguistics | 8 | Generative-agent societies & multi-agent simulation |
| 33 | The Qualitative Laboratory: Theory Prototyping and Hypothesis Generation with Large Language Models | Hugues Draelants | 2025 | arXiv.org | 0 | AI as substitute for human annotators/subjects |
| 34 | Human Simulacra: Benchmarking the Personification of Large Language Models | Qiuejie Xie et al. | 2024 | International Conference on Learning Representations | 13 | Generative-agent societies & multi-agent simulation |
| 35 | Mitigating Social Desirability Bias in Random Silicon Sampling | Sashank Chapala et al. | 2025 | arXiv.org | 1 | Silicon sampling & survey simulation |
| 36 | Random Silicon Sampling: Simulating Human Sub-Population Opinion Using a Large Language Model Based on Group-Level Demographic Information | Seungjong Sun et al. | 2024 | arXiv.org | 49 | Silicon sampling & survey simulation |
| 37 | ChatGPT is not A Man but Das Man: Representativeness and Structural Consistency of Silicon Samples Generated by Large Language Models | Dai Li et al. | 2025 | arXiv.org | 3 | Silicon sampling & survey simulation |
| 38 | Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates | Xiangyu Ma et al. | 2026 | 0 | Silicon sampling & survey simulation | |
| 39 | The threat of analytic flexibility in using large language models to simulate human data | Jamie Cummins | 2025 | arXiv.org | 15 | Silicon sampling & survey simulation |
| 40 | The Collapse of Heterogeneity in Silicon Philosophers | Yuanming Shi & Andreas Haupt | 2026 | Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency | 0 | Silicon sampling & survey simulation |
| 41 | Valid Inference with Synthetic Data via Task Exchangeability | Lezhi Tan & Tijana Zrnic | 2026 | 0 | Silicon sampling & survey simulation | |
| 42 | PoSSUM: A Protocol for Surveying Social-media Users with Multimodal LLMs | Roberto Cerina | 2025 | arXiv.org | 4 | Silicon sampling & survey simulation |
| 43 | ChatGPT vs Social Surveys: Probing Objective and Subjective Silicon Population | Muzhi Zhou et al. | 2024 | 3 | Silicon sampling & survey simulation | |
| 44 | Out of One, Many: Using Language Models to Simulate Human Samples | Lisa P. Argyle et al. | 2022 | Political Analysis | 1153 | Silicon sampling & survey simulation |
| 45 | Survey Transfer Learning: Recycling Data with Silicon Responses | Ali Amini | 2025 | 0 | Silicon sampling & survey simulation | |
| 46 | Restoring Heterogeneity in LLM-based Social Simulation: An Audience Segmentation Approach | Xiaoyou Qin et al. | 2026 | arXiv.org | 2 | Silicon sampling & survey simulation |
| 47 | Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? | John J. Horton et al. | 2023 | ACM Conference on Economics and Computation | 477 | Cognitive/economic experiments (replication) |
| 48 | Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management | Ziyan Cui et al. | 2024 | 10 | Cognitive/economic experiments (replication) | |
| 49 | Can AI Serve as a Substitute for Human Subjects in Software Engineering Research? | Marco A. Gerosa et al. | 2023 | International Conference on Automated Software Engineering | 41 | AI as substitute for human annotators/subjects |
| 50 | Large Language Models as Psychological Simulators: A Methodological Guide | Zhicheng Lin | 2025 | Advances in Methods and Practices in Psychological Science | 10 | AI as substitute for human annotators/subjects |
| 51 | Explaining Large Language Models Decisions Using Shapley Values | Behnam Mohammadi | 2024 | 15 | Cognitive/economic experiments (replication) | |
| 52 | Are Large Language Models Consistent over Value-laden Questions? | Jared Moore et al. | 2024 | Conference on Empirical Methods in Natural Language Processing | 65 | Silicon sampling & survey simulation |
| 53 | Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies | Gati Aher et al. | 2022 | International Conference on Machine Learning | 738 | Cognitive/economic experiments (replication) |
| 54 | Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings | Ryo Fukuda et al. | 2026 | 0 | Preferences, recommendation & user simulation | |
| 55 | DracoGPT: Extracting Visualization Design Preferences from Large Language Models | Huichen Will Wang et al. | 2024 | IEEE Transactions on Visualization and Computer Graphics | 24 | Preferences, recommendation & user simulation |
| 56 | Algorithmic Prompt Generation for Diverse Human-like Teaming and Communication with Large Language Models | Siddharth Srikanth et al. | 2025 | arXiv.org | 1 | Decision-making, biases & trust |
| 57 | Soundscape Captioning using Sound Affective Quality Network and Large Language Model | Yuanbo Hou et al. | 2024 | IEEE transactions on multimedia | 5 | Preferences, recommendation & user simulation |
| 58 | Strategic Interactions between Large Language Models-based Agents in Beauty Contests | Siting Estee Lu | 2024 | 5 | Generative-agent societies & multi-agent simulation | |
| 59 | Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects | Fred Heiding et al. | 2024 | arXiv.org | 35 | AI as substitute for human annotators/subjects |
| 60 | Prompting or Fine-tuning? Exploring Large Language Models for Causal Graph Validation | Yuni Susanti & Nina Holsmoelle | 2024 | 3 | Cognitive/economic experiments (replication) | |
| 61 | A closer look at how large language models trust humans: patterns and biases | Valeria Lerman & Yaniv Dover | 2025 | Proceedings of the Royal Society A Mathematical Physical and Engineering Science | 1 | AI as substitute for human annotators/subjects |
| 62 | Human-AI Collaboration in Large Language Model-Integrated Building Energy Management Systems: The Role of User Domain Knowledge and AI Literacy | Wooyoung Jung et al. | 2026 | arXiv.org | 0 | AI as substitute for human annotators/subjects |
| 63 | Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning | Saloni Dash et al. | 2025 | Annual Meeting of the Association for Computational Linguistics | 12 | Cognitive/economic experiments (replication) |
| 64 | Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data | Haobo Yang | 2026 | 0 | Decision-making, biases & trust | |
| 65 | GPT-4 as an Agronomist Assistant? Answering Agriculture Exams Using Large Language Models | Bruno Silva et al. | 2023 | 28 | Silicon sampling & survey simulation | |
| 66 | Using cognitive psychology to understand GPT-3 | Marcel Binz & Eric Schulz | 2022 | Proceedings of the National Academy of Sciences of the United States of America | 745 | Cognitive/economic experiments (replication) |
| 67 | Exploring MLLMs Perception of Network Visualization Principles | Jacob Miller et al. | 2025 | IEEE Transactions on Visualization and Computer Graphics | 1 | Cognitive/economic experiments (replication) |
| 68 | The High Cost of Incivility: Quantifying Interaction Inefficiency via Multi-Agent Monte Carlo Simulations | Benedikt Mangold | 2025 | arXiv.org | 1 | Generative-agent societies & multi-agent simulation |
| 69 | Preregistration for Experiments with AI Agents | Michelle Vaccaro | 2026 | Social Science Research Network | 1 | AI as substitute for human annotators/subjects |
| 70 | "I'm categorizing LLM as a productivity tool": Examining ethics of LLM use in HCI research practices | Shivani Kapania et al. | 2024 | Proc. ACM Hum. Comput. Interact. | 43 | Personality & psychological profiling of LLMs |
| 71 | From Text to Trust: Empowering AI-assisted Decision Making with Adaptive LLM-powered Analysis | Zhuoyan Li et al. | 2025 | International Conference on Human Factors in Computing Systems | 24 | AI as substitute for human annotators/subjects |
| 72 | Promptly Yours? A Human Subject Study on Prompt Inference in AI-Generated Art | Khoi Trinh et al. | 2024 | arXiv.org | 2 | AI as substitute for human annotators/subjects |
| 73 | TruEye: Fine-Grained Detection of AI-Generated Human Subjects in Images | Jay Barot & Dan Lin | 2026 | 0 | AI as substitute for human annotators/subjects | |
| 74 | The Emergence of Economic Rationality of GPT | Yiting Chen et al. | 2023 | Proceedings of the National Academy of Sciences of the United States of America | 187 | Silicon sampling & survey simulation |
| 75 | Multi-Turn Human-LLM Interaction Through the Lens of a Two-Way Intelligibility Protocol | Harshvardhan Mestha et al. | 2024 | 4 | Generative-agent societies & multi-agent simulation | |
| 76 | Can LLMs Replace Manual Annotation of Software Engineering Artifacts? | Toufique Ahmed et al. | 2024 | IEEE Working Conference on Mining Software Repositories | 84 | Cognitive/economic experiments (replication) |
| 77 | Does GPT Really Get It? A Hierarchical Scale to Quantify Human vs AI's Understanding of Algorithms | Mirabel Reid & Santosh S. Vempala | 2024 | arXiv.org | 1 | Cognitive/economic experiments (replication) |
| 78 | Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study | Zekai Shao et al. | 2025 | IEEE Computer Graphics and Applications | 6 | Generative-agent societies & multi-agent simulation |
| 79 | LLMs for Qualitative Data Analysis Fail on Security-specificComments in Human Experiments | Maria Camporese et al. | 2026 | arXiv.org | 0 | Cognitive/economic experiments (replication) |
| 80 | Do Large Language Models Mentalize When They Teach? | Sevan K. Harootonian et al. | 2026 | arXiv.org | 0 | Cognitive/economic experiments (replication) |
| 81 | Virtual Personas for Language Models via an Anthology of Backstories | Suhong Moon et al. | 2024 | Conference on Empirical Methods in Natural Language Processing | 52 | Public opinion & political representation |
| 82 | LLMs Model Non-WEIRD Populations: Experiments with Synthetic Cultural Agents | Augusto Gonzalez-Bonorino et al. | 2025 | arXiv.org | 6 | Cognitive/economic experiments (replication) |
| 83 | Design and Evaluation of Generative Agent-based Platform for Human-Assistant Interaction Research: A Tale of 10 User Studies | Ziyi Xuan et al. | 2025 | Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies | 2 | Generative-agent societies & multi-agent simulation |
| 84 | Guarding Your Conversations: Privacy Gatekeepers for Secure Interactions with Cloud-Based AI Models | GodsGift Uzor et al. | 2025 | International Computer Science Conference | 3 | Preferences, recommendation & user simulation |
| 85 | Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings | Harbin Hong et al. | 2025 | arXiv.org | 1 | Public opinion & political representation |
| 86 | Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions | Joseph Suh et al. | 2025 | Annual Meeting of the Association for Computational Linguistics | 66 | Public opinion & political representation |
| 87 | Can Third-parties Read Our Emotions? | Jiayi Li et al. | 2025 | Annual Meeting of the Association for Computational Linguistics | 5 | Silicon sampling & survey simulation |
| 88 | Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation | Takaya Arita et al. | 2025 | 2 | Cognitive/economic experiments (replication) | |
| 89 | Should you use LLMs to simulate opinions? Quality checks for early-stage deliberation | Terrence Neumann et al. | 2025 | AAAI Conference on Artificial Intelligence | 6 | Public opinion & political representation |
| 90 | Valid Inference with Imperfect Synthetic Data | Yewon Byun et al. | 2025 | 7 | AI as substitute for human annotators/subjects | |
| 91 | The Machine Psychology of Cooperation: Can GPT models operationalise prompts for altruism, cooperation, competitiveness and selfishness in economic games? | Steve Phelps & Yvan I. Russell | 2023 | Journal of Physics: Complexity | 29 | Cognitive/economic experiments (replication) |
| 92 | Modeling Human Subjectivity in LLMs Using Explicit and Implicit Human Factors in Personas | Salvatore Giorgi et al. | 2024 | Conference on Empirical Methods in Natural Language Processing | 28 | Cognitive/economic experiments (replication) |
| 93 | When Do LLMs Generate Realistic Social Networks? A Multi-Dimensional Study of Culture, Language, Scale, and Method | Sai Hemanth Kilaru et al. | 2026 | 0 | Generative-agent societies & multi-agent simulation | |
| 94 | This human study did not involve human subjects: Validating LLM simulations as behavioral evidence | Jessica Hullman et al. | 2026 | arXiv.org | 13 | Cognitive/economic experiments (replication) |
| 95 | UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents | Yuxuan Lu et al. | 2025 | arXiv.org | 15 | Generative-agent societies & multi-agent simulation |
| 96 | AI Realtor: Towards Grounded Persuasive Language Generation for Automated Copywriting | Jibang Wu et al. | 2025 | Proceedings of the ACM Conference on AI and Agentic Systems | 7 | Preferences, recommendation & user simulation |
| 97 | Mind the (DH) Gap! A Contrast in Risky Choices Between Reasoning and Conversational LLMs | Luise Ge et al. | 2026 | Volume 1 | 0 | Decision-making, biases & trust |
| 98 | The amplifier effect of artificial agents in social contagion | Eric Hitz et al. | 2025 | arXiv.org | 1 | Generative-agent societies & multi-agent simulation |
| 99 | Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors | Veronica Mangiaterra et al. | 2025 | arXiv.org | 1 | Silicon sampling & survey simulation |
| 100 | LLM-driven Imitation of Subrational Behavior : Illusion or Reality? | Andrea Coletta et al. | 2024 | arXiv.org | 14 | Cognitive/economic experiments (replication) |
| 101 | Humans expect rationality and cooperation from LLM opponents in strategic games | Darija Barak & Miguel Costa-Gomes | 2025 | arXiv.org | 1 | Cognitive/economic experiments (replication) |
| 102 | $\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training | Aurélien Bück-Kaeffer et al. | 2025 | arXiv.org | 7 | Generative-agent societies & multi-agent simulation |
| 103 | Beyond Nash Equilibrium: Bounded Rationality of LLMs and humans in Strategic Decision-making | Kehan Zheng et al. | 2025 | arXiv.org | 9 | Cognitive/economic experiments (replication) |
| 104 | Multi-Agent Home Energy Management Assistant | Wooyoung Jung | 2026 | SoftwareX | 2 | Preferences, recommendation & user simulation |
| 105 | Who is More Bayesian: Humans or ChatGPT? | Tianshi Mu et al. | 2025 | arXiv.org | 0 | Silicon sampling & survey simulation |
| 106 | Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust | Yuzhou Wu et al. | 2025 | arXiv.org | 1 | Decision-making, biases & trust |
| 107 | Can AI with High Reasoning Ability Replicate Human-like Decision Making in Economic Experiments? | Ayato Kitadai et al. | 2024 | Group Decision and Negotiation | 14 | Cognitive/economic experiments (replication) |
| 108 | Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona? | Junhyuk Choi et al. | 2025 | arXiv.org | 0 | Cognitive/economic experiments (replication) |
| 109 | Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina | Yuan Gao et al. | 2024 | arXiv.org | 50 | Cognitive/economic experiments (replication) |
| 110 | Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions | Olivier Toubia et al. | 2025 | Marketing science (Providence, R.I.) | 17 | Cognitive/economic experiments (replication) |
| 111 | Large Language Models Show Human-like Social Desirability Biases in Survey Responses | Aadesh Salecha et al. | 2024 | arXiv.org | 32 | Decision-making, biases & trust |
| 112 | Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data | Jihong Zhang et al. | 2025 | arXiv.org | 1 | Silicon sampling & survey simulation |
| 113 | A Neural Topic Method Using a Large-Language-Model-in-the-Loop for Business Research | Stephan Ludwig et al. | 2026 | arXiv.org | 0 | Preferences, recommendation & user simulation |
| 114 | Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions | Ji Huang et al. | 2025 | Annual Meeting of the Association for Computational Linguistics | 3 | Silicon sampling & survey simulation |
| 115 | Delving Into the Psychology of Machines: Exploring the Structure of Self-Regulated Learning via LLM-Generated Survey Responses | Leonie V. D. E. Vogelsmeier et al. | 2025 | Computers in Human Behavior | 4 | Silicon sampling & survey simulation |
| 116 | QSTN: A Modular Framework for Robust Questionnaire Inference with Large Language Models | Maximilian Kreutner et al. | 2025 | Conference of the European Chapter of the Association for Computational Linguistics | 1 | Preferences, recommendation & user simulation |
| 117 | Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case | Bastián González-Bustamante et al. | 2025 | arXiv.org | 1 | Public opinion & political representation |
| 118 | Assessing Personalized AI Mentoring with Large Language Models in the Computing Field | Xiao Luo et al. | 2024 | 2025 IEEE Symposium on Computational Intelligence in Natural Language Processing and Social Media (CI-NLPSoMe Companion) | 5 | Silicon sampling & survey simulation |
| 119 | Generating Public Health Responses using Survey-Augmented Large Language Models | Leonardo Marciaga et al. | 2026 | 0 | Public opinion & political representation | |
| 120 | How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective | Chengpiao Huang et al. | 2025 | 8 | Silicon sampling & survey simulation | |
| 121 | Occupational Prompting Reveals Cultural Bias in Large Language Models | Maksim E. Eren et al. | 2026 | 0 | Public opinion & political representation | |
| 122 | Large Language Models as Conversational Movie Recommenders: A User Study | Ruixuan Sun et al. | 2024 | arXiv.org | 9 | Preferences, recommendation & user simulation |
| 123 | Can Large Language Models Capture Public Opinion about Global Warming? An Empirical Assessment of Algorithmic Fidelity and Bias | S. Lee et al. | 2023 | PLOS Climate | 36 | Public opinion & political representation |
| 124 | Improving Cross-Cultural Survey Simulation with Calibrated Value Personas | Axel Abels et al. | 2026 | 0 | Public opinion & political representation | |
| 125 | From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes | Rubén Garzón et al. | 2026 | 0 | Public opinion & political representation | |
| 126 | Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data | Eun Cheol Choi et al. | 2026 | 0 | Silicon sampling & survey simulation | |
| 127 | Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification | Stefan Krsteski et al. | 2025 | Volume 1 | 10 | Public opinion & political representation |
| 128 | Synthesizing Public Opinions with LLMs: Role Creation, Impacts, and the Future to eDemorcacy | Rabimba Karanjai et al. | 2025 | International Conference on eDemocracy & eGovernment | 6 | Public opinion & political representation |
| 129 | A Zero-Shot LLM Framework for Automatic Assignment Grading in Higher Education | Calvin Yeung et al. | 2025 | arXiv.org | 13 | Preferences, recommendation & user simulation |
| 130 | Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators | Sungjib Lim et al. | 2025 | arXiv.org | 1 | Public opinion & political representation |
| 131 | Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation | James Mooney et al. | 2025 | arXiv.org | 2 | Generative-agent societies & multi-agent simulation |
| 132 | Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data | Jason Miklian et al. | 2026 | arXiv.org | 0 | Silicon sampling & survey simulation |
| 133 | Learning from Convenience Samples: A Case Study on Fine-Tuning LLMs for Survey Non-response in the German Longitudinal Election Study | Tobias Holtdirk et al. | 2025 | arXiv.org | 2 | Silicon sampling & survey simulation |
| 134 | Balancing Domestic and Global Perspectives: Evaluating Dual-Calibration and LLM-Generated Nudges for Diverse News Recommendation | Ruixuan Sun et al. | 2026 | arXiv.org | 0 | Preferences, recommendation & user simulation |
| 135 | A Penny for Your Prompts: Experiments Detecting and Mitigating LLM Usage by Survey Respondents | Zane Xu & Nathan Malkin | 2026 | 0 | Public opinion & political representation | |
| 136 | AlignSurvey: A Comprehensive Benchmark for Human Preferences Alignment in Social Surveys | Chenxi Lin et al. | 2025 | AAAI Conference on Artificial Intelligence | 2 | Public opinion & political representation |
| 137 | Towards Measuring the Representation of Subjective Global Opinions in Language Models | Esin Durmus et al. | 2023 | arXiv.org | 434 | Public opinion & political representation |
| 138 | LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans | Ljubisa Bojic et al. | 2026 | arXiv.org | 1 | Generative-agent societies & multi-agent simulation |
| 139 | Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive Reasoning | Haijiang Liu et al. | 2025 | Conference on Empirical Methods in Natural Language Processing | 4 | Personality & psychological profiling of LLMs |
| 140 | Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys | Zikun Ye & Hema Yoganarasimhan | 2026 | arXiv.org | 0 | Public opinion & political representation |
| 141 | LLM-Mirror: A Generated-Persona Approach for Survey Pre-Testing | Sunwoong Kim et al. | 2024 | arXiv.org | 9 | Public opinion & political representation |
| 142 | An Investigation on How AI-Generated Responses Affect SoftwareEngineering Surveys | Ronnie de Souza Santos et al. | 2025 | arXiv.org | 0 | Public opinion & political representation |
| 143 | From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM | Suyash Fulay et al. | 2025 | Conference of the European Chapter of the Association for Computational Linguistics | 1 | Preferences, recommendation & user simulation |
| 144 | German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies | Jens Rupprecht et al. | 2025 | 1 | AI as substitute for human annotators/subjects | |
| 145 | Demonstrations of the Potential of AI-based Political Issue Polling | Nathan E. Sanders et al. | 2023 | Issue 5.4, Fall 2023 | 32 | Silicon sampling & survey simulation |
| 146 | Assessing the Potential of Generative Agents in Crowdsourced Fact-Checking | Luigia Costabile et al. | 2025 | Online Soc. Networks Media | 11 | Decision-making, biases & trust |
| 147 | Can Generative Agent-Based Modeling Replicate the Friendship Paradox in Social Media Simulations? | Gian Marco Orlando et al. | 2025 | Web Science Conference | 13 | Generative-agent societies & multi-agent simulation |
| 148 | Simulating Persuasive Dialogues on Meat Reduction with Generative Agents | Georg Ahnert et al. | 2025 | arXiv.org | 1 | Generative-agent societies & multi-agent simulation |
| 149 | Affordable Generative Agents | Yangbin Yu et al. | 2024 | Trans. Mach. Learn. Res. | 10 | Generative-agent societies & multi-agent simulation |
| 150 | Multimodal Safety Evaluation in Generative Agent Social Simulations | Alhim Vera et al. | 2025 | arXiv.org | 4 | Generative-agent societies & multi-agent simulation |
| 151 | Generative Agents: Interactive Simulacra of Human Behavior | Joon Sung Park et al. | 2023 | ACM Symposium on User Interface Software and Technology | 4655 | Generative-agent societies & multi-agent simulation |
| 152 | Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism | Simon Münker et al. | 2025 | Conference of the European Chapter of the Association for Computational Linguistics | 3 | AI as substitute for human annotators/subjects |
| 153 | On Generative Agents in Recommendation | An Zhang et al. | 2023 | Annual International ACM SIGIR Conference on Research and Development in Information Retrieval | 147 | Preferences, recommendation & user simulation |
| 154 | Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions | Logan Cross et al. | 2025 | Annual Meeting of the Cognitive Science Society | 3 | Generative-agent societies & multi-agent simulation |
| 155 | Interview-Informed Generative Agents for Product Discovery: A Validation Study | Zichao Wang & Alexa Siu | 2026 | International Conference on Human Factors in Computing Systems | 1 | AI as substitute for human annotators/subjects |
| 156 | Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection | Bo Yang et al. | 2025 | arXiv.org | 9 | Generative-agent societies & multi-agent simulation |
| 157 | Can A Society of Generative Agents Simulate Human Behavior and Inform Public Health Policy? A Case Study on Vaccine Hesitancy | Abe Bohan Hou et al. | 2025 | arXiv.org | 21 | Generative-agent societies & multi-agent simulation |
| 158 | ID-RAG: Identity Retrieval-Augmented Generation for Long-Horizon Persona Coherence in Generative Agents | Daniel Platnick et al. | 2025 | arXiv.org | 3 | Preferences, recommendation & user simulation |
| 159 | Negotiating Comfort: Simulating Personality-Driven LLM Agents in Shared Residential Social Networks | Ann Nedime Nese Rende et al. | 2025 | arXiv.org | 0 | Generative-agent societies & multi-agent simulation |
| 160 | Generative Agent-Based Modeling: Unveiling Social System Dynamics through Coupling Mechanistic Models with Generative Artificial Intelligence | Navid Ghaffarzadegan et al. | 2023 | System Dynamics Review | 58 | Generative-agent societies & multi-agent simulation |
| 161 | From Who They Are to How They Act: Behavioral Traits in Generative Agent-Based Models of Social Media | Valerio La Gatta et al. | 2026 | arXiv.org | 1 | Generative-agent societies & multi-agent simulation |
| 162 | Integrating LLM in Agent-Based Social Simulation: Opportunities and Challenges | Patrick Taillandier et al. | 2025 | arXiv.org | 9 | Generative-agent societies & multi-agent simulation |
| 163 | AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society | Jinghua Piao et al. | 2025 | arXiv.org | 171 | Generative-agent societies & multi-agent simulation |
| 164 | How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm | Taisei Hishiki et al. | 2026 | arXiv.org | 0 | Generative-agent societies & multi-agent simulation |
| 165 | LLM-Based Multi-Agent System for Simulating and Analyzing Marketing and Consumer Behavior | Man-Lin Chu et al. | 2025 | IEEE International Conference on e-Business Engineering | 2 | Decision-making, biases & trust |
| 166 | Psychology-driven LLM Agents for Explainable Panic Prediction on Social Media during Sudden Disaster Events | Mengzhu Liu et al. | 2025 | arXiv.org | 0 | Generative-agent societies & multi-agent simulation |
| 167 | Hierarchical Generative Agents for Simulating Sequential Human Behavior | Maria G. Mendoza et al. | 2026 | 0 | Decision-making, biases & trust | |
| 168 | Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations | Gian Marco Orlando et al. | 2025 | The Web Conference | 6 | Generative-agent societies & multi-agent simulation |
| 169 | Implicit Behavioral Alignment of Language Agents in High-Stakes Crowd Simulations | Yunzhe Wang et al. | 2025 | Conference on Empirical Methods in Natural Language Processing | 5 | AI as substitute for human annotators/subjects |
| 170 | NetworkGames: Simulating Cooperation in Network Games with Personality-driven LLM Agents | Xuan Qiu | 2025 | arXiv.org | 1 | Personality & psychological profiling of LLMs |
| 171 | Simulating Misinformation Vulnerabilities With Agent Personas | David Farr et al. | 2025 | Online World Conference on Soft Computing in Industrial Applications | 1 | Generative-agent societies & multi-agent simulation |
| 172 | Mind the Gap: The Divergence Between Human and LLM-Generated Tasks | Yi-Long Lu et al. | 2025 | 0 | Cognitive/economic experiments (replication) | |
| 173 | Mechanism Plausibility in Generative Agent-Based Modeling | Patrick Zhao et al. | 2026 | Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency | 0 | Generative-agent societies & multi-agent simulation |
| 174 | Cultural evolution in populations of Large Language Models | Jérémy Perez et al. | 2024 | arXiv.org | 18 | Generative-agent societies & multi-agent simulation |
| 175 | LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals | Joon Sung Park et al. | 2024 | 304 | Public opinion & political representation | |
| 176 | Games Agents Play: Towards Transactional Analysis in LLM-based Multi-Agent Systems | Monika Zamojska & Jarosław A. Chudziak | 2025 | Annual Meeting of the Cognitive Science Society | 3 | Generative-agent societies & multi-agent simulation |
| 177 | Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation | David Eric Austin et al. | 2024 | ACM Conference on Recommender Systems | 24 | Preferences, recommendation & user simulation |
| 178 | Asking Clarifying Questions for Preference Elicitation With Large Language Models | Ali Montazeralghaem et al. | 2025 | arXiv.org | 2 | Preferences, recommendation & user simulation |
| 179 | When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models | Yujun Zhou & Christopher M. Ackerman | 2026 | 0 | Preferences, recommendation & user simulation | |
| 180 | "I Want It That Way": Enabling Interactive Decision Support Using Large Language Models and Constraint Programming | Connor Lawless et al. | 2023 | ACM Trans. Interact. Intell. Syst. | 38 | Preferences, recommendation & user simulation |
| 181 | Large Language Model Driven Recommendation | Anton Korikov et al. | 2024 | arXiv.org | 0 | Preferences, recommendation & user simulation |
| 182 | Decisive: Guiding User Decisions with Optimal Preference Elicitation from Unstructured Documents | Akriti Jain et al. | 2026 | Volume 1 | 0 | Preferences, recommendation & user simulation |
| 183 | Preference is More Than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback | Shengbo Wang et al. | 2025 | arXiv.org | 0 | Preferences, recommendation & user simulation |
| 184 | From Guessing to Asking: An Approach to Resolving the Persona Knowledge Gap in LLMs during Multi-Turn Conversations | Sarvesh Baskar et al. | 2025 | arXiv.org | 1 | Preferences, recommendation & user simulation |
| 185 | LLM Consumer Behavior Theory: Foundations of a Novel Research Field | Manon Reusens et al. | 2026 | 0 | Decision-making, biases & trust | |
| 186 | Are Generative AI Agents Effective Personalized Financial Advisors? | Takehiro Takayanagi et al. | 2025 | Annual International ACM SIGIR Conference on Research and Development in Information Retrieval | 30 | Preferences, recommendation & user simulation |
| 187 | Your Reviews Replicate You: LLM-Based Agents as Customer Digital Twins for Conjoint Analysis | Bin Xuan et al. | 2026 | arXiv.org | 0 | AI as substitute for human annotators/subjects |
| 188 | ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences | Bang Nguyen et al. | 2026 | arXiv.org | 5 | Generative-agent societies & multi-agent simulation |
| 189 | Large language models replicate and predict human cooperation across experiments in game theory | Andrea Cera Palatsi et al. | 2025 | arXiv.org | 1 | Cognitive/economic experiments (replication) |
| 190 | Beyond Averages: Evaluating LLMs on Human Survey Replication at the Distributional Level | Jeonghyeon Moon et al. | 2026 | 0 | Cognitive/economic experiments (replication) | |
| 191 | Simulating and Experimenting with Social Media Mobilization Using LLM Agents | Sadegh Shirani & Mohsen Bayati | 2025 | arXiv.org | 2 | Generative-agent societies & multi-agent simulation |
| 192 | PaperBench: Evaluating AI's Ability to Replicate AI Research | Giulio Starace et al. | 2025 | International Conference on Machine Learning | 205 | AI as substitute for human annotators/subjects |
| 193 | Large language models can replicate cross-cultural differences in personality | Paweł Niszczota et al. | 2023 | Journal of Research in Personality | 17 | Personality & psychological profiling of LLMs |
| 194 | From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review | Yaohui Zhang et al. | 2025 | arXiv.org | 10 | Decision-making, biases & trust |
| 195 | Identity as Attractor: Geometric Evidence for Persistent Agent Architecture in LLM Activation Space | Vladimir Vasilenko | 2026 | arXiv.org | 0 | Silicon sampling & survey simulation |
| 196 | SOCK: A Benchmark for Measuring Self-Replication in Large Language Models | Justin Chavarria et al. | 2025 | arXiv.org | 2 | Generative-agent societies & multi-agent simulation |
| 197 | MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval | Saksham Sahai Srivastava & Haoyu He | 2025 | arXiv.org | 43 | Generative-agent societies & multi-agent simulation |
| 198 | The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges | Jiqun Liu & Jiangen He | 2024 | ACM Transactions on Interactive Intelligent Systems | 8 | AI as substitute for human annotators/subjects |
| 199 | How well LLM-based test generation techniques perform with newer LLM versions? | Michael Konstantinou et al. | 2026 | arXiv.org | 1 | Cognitive/economic experiments (replication) |
| 200 | LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams | Aju Ani Justus & Chris Baber | 2025 | European Conference on Artificial Intelligence | 1 | Cognitive/economic experiments (replication) |
| 201 | Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate | Saurabh Ranjan & Brian Odegaard | 2025 | 0 | Cognitive/economic experiments (replication) | |
| 202 | Can AI Solve the Peer Review Crisis? A Large Scale Cross Model Experiment of LLMs' Performance and Biases in Evaluating over 1000 Economics Papers | Pat Pataranutaporn et al. | 2025 | 1 | AI as substitute for human annotators/subjects | |
| 203 | Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems | Donghyun Lee & Mo Tiwari | 2024 | arXiv.org | 112 | Generative-agent societies & multi-agent simulation |
| 204 | Predicting Effects, Missing Distributions: Evaluating LLMs as Human Behavior Simulators in Operations Management | Runze Zhang et al. | 2025 | arXiv.org | 1 | Cognitive/economic experiments (replication) |
| 205 | LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation | Joseph Enguehard et al. | 2025 | Proceedings of the Natural Legal Language Processing Workshop 2025 | 7 | Generative-agent societies & multi-agent simulation |
| 206 | Simulating Cooperative Prosocial Behavior with Multi-Agent LLMs: Evidence and Mechanisms for AI Agents to Inform Policy Decisions | Karthik Sreedhar et al. | 2025 | International Conference on Intelligent User Interfaces | 30 | Generative-agent societies & multi-agent simulation |
| 207 | Exploring LLM Features in Predictive Process Monitoring for Small-Scale Event-Logs | Alessandro Padella et al. | 2026 | arXiv.org | 0 | Generative-agent societies & multi-agent simulation |
| 208 | A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions | Laurène Vaugrante et al. | 2024 | arXiv.org | 11 | Silicon sampling & survey simulation |
| 209 | Unit Testing Past vs. Present: Examining LLMs' Impact on Defect Detection and Efficiency | Rudolf Ramler et al. | 2025 | arXiv.org | 4 | Preferences, recommendation & user simulation |
| 210 | The Promise and Challenges of Using LLMs to Accelerate the Screening Process of Systematic Reviews | Aleksi Huotala et al. | 2024 | International Conference on Evaluation & Assessment in Software Engineering | 39 | Cognitive/economic experiments (replication) |
| 211 | TradingAgents: Multi-Agents LLM Financial Trading Framework | Yijia Xiao et al. | 2024 | arXiv.org | 167 | Generative-agent societies & multi-agent simulation |
| 212 | Drop the Hierarchy and Roles: How Self-Organizing LLM Agents Outperform Designed Structures | Victoria Dochkina | 2026 | arXiv.org | 0 | Generative-agent societies & multi-agent simulation |
| 213 | LLM-in-the-loop: Leveraging Large Language Model for Thematic Analysis | Shih-Chieh Dai et al. | 2023 | Conference on Empirical Methods in Natural Language Processing | 159 | Cognitive/economic experiments (replication) |
| 214 | Objection Overruled! Lay People can Distinguish Large Language Models from Lawyers, but still Favour Advice from an LLM | Eike Schneiders et al. | 2024 | International Conference on Human Factors in Computing Systems | 25 | AI as substitute for human annotators/subjects |
| 215 | Strategizing with AI: Insights from a Beauty Contest Experiment | Iuliia Alekseenko et al. | 2025 | Social Science Research Network | 6 | Cognitive/economic experiments (replication) |
| 216 | LLM Agents Display Human Biases but Exhibit Distinct Learning Patterns | Idan Horowitz & Ori Plonsky | 2025 | arXiv.org | 3 | Cognitive/economic experiments (replication) |
| 217 | Benchmarking LLM-based Relevance Judgment Methods | Negar Arabzadeh & Charles L. A. Clarke | 2025 | Annual International ACM SIGIR Conference on Research and Development in Information Retrieval | 24 | Preferences, recommendation & user simulation |
| 218 | [Re] Benchmarking LLM Capabilities in Negotiation through Scoreable Games | Jorge Carrasco Pollo et al. | 2026 | Trans. Mach. Learn. Res. | 0 | Generative-agent societies & multi-agent simulation |
| 219 | Can Generative AI agents behave like humans? Evidence from laboratory market experiments | R. Maria del Rio-Chanona et al. | 2025 | arXiv.org | 17 | Cognitive/economic experiments (replication) |
| 220 | $δ$-STEAL: LLM Stealing Attack with Local Differential Privacy | Kieu Dang et al. | 2025 | arXiv.org | 3 | Preferences, recommendation & user simulation |
| 221 | Evaluating the Simulation of Human Personality-Driven Susceptibility to Misinformation with LLMs | Manuel Pratelli & Marinella Petrocchi | 2025 | European Conference on Artificial Intelligence | 3 | Personality & psychological profiling of LLMs |
| 222 | Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies | Florian Angermeir et al. | 2025 | arXiv.org | 13 | Cognitive/economic experiments (replication) |
| 223 | Representational Collapse in Multi-Agent LLM Committees: Measurement and Diversity-Aware Consensus | Dipkumar Patel | 2026 | arXiv.org | 2 | Generative-agent societies & multi-agent simulation |
| 224 | LLM as GNN: Graph Vocabulary Learning for Text-Attributed Graph Foundation Models | Xi Zhu et al. | 2025 | arXiv.org | 32 | Preferences, recommendation & user simulation |
| 225 | Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection | Kieu Dang et al. | 2026 | 0 | Preferences, recommendation & user simulation | |
| 226 | Simulating Human Strategic Behavior: Comparing Single and Multi-agent LLMs | Karthik Sreedhar & Lydia Chilton | 2024 | Hawaii International Conference on System Sciences | 20 | Cognitive/economic experiments (replication) |
| 227 | Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory | Gordon Dai et al. | 2024 | Frontiers of Physics | 11 | Generative-agent societies & multi-agent simulation |
| 228 | RocketEval: Efficient Automated LLM Evaluation via Grading Checklist | Tianjun Wei et al. | 2025 | International Conference on Learning Representations | 29 | Preferences, recommendation & user simulation |
| 229 | A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs | Zijie Liu et al. | 2026 | arXiv.org | 0 | Personality & psychological profiling of LLMs |
| 230 | Exploring the psychology of LLMs' Moral and Legal Reasoning | Guilherme F. C. F. Almeida et al. | 2023 | Artificial Intelligence | 91 | Cognitive/economic experiments (replication) |
| 231 | Hallucination as Context Drift: Synchronization Protocols for Multi-Agent LLM Systems | Carson Rodrigues | 2026 | 0 | Generative-agent societies & multi-agent simulation | |
| 232 | Prompting for Policy: Forecasting Macroeconomic Scenarios with Synthetic LLM Personas | Giulia Iadisernia & Carolina Camassa | 2025 | International Conference on AI in Finance | 3 | Silicon sampling & survey simulation |
| 233 | The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans | Odysseas S. Chlapanis et al. | 2026 | 0 | Cognitive/economic experiments (replication) | |
| 234 | Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design? | Eduard Kuric et al. | 2026 | 0 | AI as substitute for human annotators/subjects | |
| 235 | LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code | Liwei Guo et al. | 2025 | arXiv.org | 2 | Cognitive/economic experiments (replication) |
| 236 | Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines | Yitao Li | 2026 | 3 | Decision-making, biases & trust | |
| 237 | FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs | Junzhe Jiang et al. | 2025 | arXiv.org | 8 | Preferences, recommendation & user simulation |
| 238 | Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate | Xiqi Hao et al. | 2026 | 1 | Generative-agent societies & multi-agent simulation | |
| 239 | Cognitive LLMs: Towards Integrating Cognitive Architectures and Large Language Models for Manufacturing Decision-making | Siyu Wu et al. | 2024 | arXiv.org | 10 | Decision-making, biases & trust |
| 240 | HLB: Benchmarking LLMs' Humanlikeness in Language Use | Xufeng Duan et al. | 2024 | arXiv.org | 11 | Cognitive/economic experiments (replication) |
| 241 | LLM-Based Robustness Testing of Microservice Applications: An Empirical Study | Hrushitha Goud Tigulla & Marco Vieira | 2026 | 0 | Cognitive/economic experiments (replication) | |
| 242 | Simulating Field Experiments with Large Language Models | Yaoyu Chen et al. | 2024 | arXiv.org | 0 | Cognitive/economic experiments (replication) |
| 243 | Deconfounded Causality-aware Parameter-Efficient Fine-Tuning for Problem-Solving Improvement of LLMs | Ruoyu Wang et al. | 2024 | WISE | 1 | Cognitive/economic experiments (replication) |
| 244 | BOUNDARY_SYNC: Measuring Communication-Induced Representational Coupling in Multi-Agent LLM Systems | Zewen Liu | 2026 | 0 | Generative-agent societies & multi-agent simulation | |
| 245 | An Experimental Study of Competitive Market Behavior Through LLMs | Jingru Jia & Zehua Yuan | 2024 | arXiv.org | 3 | Cognitive/economic experiments (replication) |
| 246 | PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research | Tingjia Miao et al. | 2026 | arXiv.org | 3 | Generative-agent societies & multi-agent simulation |
| 247 | GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference | Yu Han et al. | 2025 | arXiv.org | 9 | Personality & psychological profiling of LLMs |
| 248 | Bias in LLMs as Annotators: The Effect of Party Cues on Labelling Decision by Large Language Models | Sebastian Vallejo Vera & Hunter Driggers | 2024 | arXiv.org | 2 | Cognitive/economic experiments (replication) |
| 249 | More Is Not Always Better: Cross-Component Interference in LLM Agent Scaffolding | Ming Liu | 2026 | 0 | Generative-agent societies & multi-agent simulation | |
| 250 | LLM in the Shell: Generative Honeypots | Muris Sladić et al. | 2023 | 2024 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW) | 65 | Preferences, recommendation & user simulation |
| 251 | SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM | Xiaojiang Zhang et al. | 2025 | arXiv.org | 65 | Cognitive/economic experiments (replication) |
| 252 | Multi-Agent Strategic Games with LLMs | Maxim Chupilkin | 2026 | 0 | Generative-agent societies & multi-agent simulation | |
| 253 | Cognitive networks reconstruct mindsets about STEM subjects and educational contexts in almost 1000 high-schoolers, University students and LLM-based digital twins | Francesco Gariboldi et al. | 2026 | arXiv.org | 0 | Decision-making, biases & trust |
| 254 | PUB: An LLM-Enhanced Personality-Driven User Behaviour Simulator for Recommender System Evaluation | Chenglong Ma et al. | 2025 | Annual International ACM SIGIR Conference on Research and Development in Information Retrieval | 15 | Personality & psychological profiling of LLMs |
| 255 | The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval | Zekai Tong et al. | 2026 | 0 | Decision-making, biases & trust | |
| 256 | Large Language Models Do Not Simulate Human Psychology | Sarah Schröder et al. | 2025 | arXiv.org | 23 | Cognitive/economic experiments (replication) |
| 257 | Can Large Language Models Simulate Human Cognition Beyond Behavioral Imitation? | Yuxuan Gu et al. | 2026 | arXiv.org | 0 | Cognitive/economic experiments (replication) |
| 258 | Simulating Human Memory with Language Models | Qihan Wang et al. | 2026 | 0 | Cognitive/economic experiments (replication) | |
| 259 | Can LLMs Simulate Human Behavioral Variability? A Case Study in the Phonemic Fluency Task | Mengyang Qiu et al. | 2025 | Proceedings of the Language Resources and Evaluation Conference | 2 | Cognitive/economic experiments (replication) |
| 260 | Limited Ability of LLMs to Simulate Human Psychological Behaviours: a Psychometric Analysis | Nikolay B Petrov et al. | 2024 | arXiv.org | 45 | Cognitive/economic experiments (replication) |
| 261 | The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective | George Gui & Olivier Toubia | 2023 | Social Science Research Network | 92 | Cognitive/economic experiments (replication) |
| 262 | An Appraisal-Based Chain-Of-Emotion Architecture for Affective Language Model Game Agents | Maximilian Croissant et al. | 2023 | PLoS ONE | 23 | Generative-agent societies & multi-agent simulation |
| 263 | Identity, Cooperation and Framing Effects within Groups of Real and Simulated Humans | Suhong Moon et al. | 2026 | arXiv.org | 2 | Cognitive/economic experiments (replication) |
| 264 | Shifting Power: Leveraging LLMs to Simulate Human Aversion in ABMs of Bilateral Financial Exchanges, A bond market study | Alicia Vidler & Toby Walsh | 2025 | Adaptive Agents and Multi-Agent Systems | 1 | Generative-agent societies & multi-agent simulation |
| 265 | SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors | Tiancheng Hu et al. | 2025 | arXiv.org | 33 | Cognitive/economic experiments (replication) |
| 266 | SUBER: An RL Environment with Simulated Human Behavior for Recommender Systems | Nathan Corecco et al. | 2024 | European Conference on Artificial Intelligence | 13 | Preferences, recommendation & user simulation |
| 267 | Simulating Financial Market via Large Language Model based Agents | Shen Gao et al. | 2024 | arXiv.org | 36 | Generative-agent societies & multi-agent simulation |
| 268 | TraderTalk: An LLM Behavioural ABM applied to Simulating Human Bilateral Trading Interactions | Alicia Vidler & Toby Walsh | 2024 | International Conference on Agents | 3 | Generative-agent societies & multi-agent simulation |
| 269 | Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments | Amogh Mannekote et al. | 2024 | AAAI Conference on Artificial Intelligence | 15 | Cognitive/economic experiments (replication) |
| 270 | Evaluating Cultural Adaptability of a Large Language Model via Simulation of Synthetic Personas | Louis Kwok et al. | 2024 | arXiv.org | 39 | AI as substitute for human annotators/subjects |
| 271 | Advancing Conversational Psychotherapy: Integrating Privacy, Dual-Memory, and Domain Expertise with Large Language Models | XiuYu Zhang & Zening Luo | 2024 | arXiv.org | 4 | Preferences, recommendation & user simulation |
| 272 | Probing Language Models' Gesture Understanding for Enhanced Human-AI Interaction | Philipp Wicke | 2024 | arXiv.org | 4 | Cognitive/economic experiments (replication) |
| 273 | AMONGAGENTS: Evaluating Large Language Models in the Interactive Text-Based Social Deduction Game | Yizhou Chi et al. | 2024 | arXiv.org | 22 | Generative-agent societies & multi-agent simulation |
| 274 | BILLY: Steering Large Language Models via Merging Persona Vectors for Creative Generation | Tsung-Min Pai et al. | 2025 | Conference of the European Chapter of the Association for Computational Linguistics | 5 | Preferences, recommendation & user simulation |
| 275 | An Analysis of Large Language Models for Simulating User Responses in Surveys | Ziyun Yu et al. | 2025 | IJCNLP-AACL | 0 | Public opinion & political representation |
| 276 | Using Large Language Models to Simulate Human Behavioural Experiments: Port of Mars | Oliver Slumbers et al. | 2025 | arXiv.org | 6 | Cognitive/economic experiments (replication) |
| 277 | PersLLM: A Personified Training Approach for Large Language Models | Zheni Zeng et al. | 2024 | Conference on Empirical Methods in Natural Language Processing | 6 | Cognitive/economic experiments (replication) |
| 278 | Teaching Values to Machines: Simulating Human-Like Behavior in LLMs | Asaf Yehudai et al. | 2026 | IEEE Games Entertainment Media Conference | 0 | Cognitive/economic experiments (replication) |
| 279 | Affective Computing in the Era of Large Language Models: A Survey from the NLP Perspective | Yiqun Zhang et al. | 2024 | arXiv.org | 48 | Silicon sampling & survey simulation |
| 280 | From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents | Xinyi Mou et al. | 2024 | ACM Computing Surveys | 83 | Generative-agent societies & multi-agent simulation |
| 281 | Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation | Se-eun Yoon et al. | 2024 | North American Chapter of the Association for Computational Linguistics | 67 | Preferences, recommendation & user simulation |
| 282 | Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements | Yushan Qian et al. | 2023 | Conference on Empirical Methods in Natural Language Processing | 91 | Preferences, recommendation & user simulation |
| 283 | GenSim: A General Social Simulation Platform with Large Language Model based Agents | Jiakai Tang et al. | 2024 | North American Chapter of the Association for Computational Linguistics | 50 | Generative-agent societies & multi-agent simulation |
| 284 | Do LLMs Play Dice? Exploring Probability Distribution Sampling in Large Language Models for Behavioral Simulation | Jia Gu et al. | 2024 | International Conference on Computational Linguistics | 17 | Decision-making, biases & trust |
| 285 | Systematic Bias in Large Language Models: Discrepant Response Patterns in Binary vs. Continuous Judgment Tasks | Yi-Long Lu et al. | 2025 | Annual Meeting of the Cognitive Science Society | 6 | Decision-making, biases & trust |
| 286 | Psychometric Predictive Power of Large Language Models | Tatsuki Kuribayashi et al. | 2023 | NAACL-HLT | 10 | Cognitive/economic experiments (replication) |
| 287 | TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding | Fan Yang et al. | 2026 | arXiv.org | 1 | Generative-agent societies & multi-agent simulation |
| 288 | Can Large Language Models Simulate Human Responses? A Case Study of Stated Preference Experiments in the Context of Heating-related Choices | Han Wang et al. | 2025 | 2 | Decision-making, biases & trust | |
| 289 | How Well Do Large Language Models Capture Human Personality? | Aanisha Bhattacharyya et al. | 2026 | 0 | Preferences, recommendation & user simulation | |
| 290 | Temporal-Aware User Behaviour Simulation with Large Language Models for Recommender Systems | Xinye Wanyan et al. | 2025 | International Conference on Information and Knowledge Management | 4 | Preferences, recommendation & user simulation |
| 291 | Context-Value-Action Architecture for Value-Driven Large Language Model Agents | TianZe Zhang et al. | 2026 | Annual Meeting of the Association for Computational Linguistics | 0 | Cognitive/economic experiments (replication) |
| 292 | AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction | Junsol Kim & Byungkyu Lee | 2023 | 69 | Public opinion & political representation | |
| 293 | Benchmarking Distributional Alignment of Large Language Models | Nicole Meister et al. | 2024 | arXiv.org | 45 | Cognitive/economic experiments (replication) |
| 294 | Large Language Model for Participatory Urban Planning | Zhilun Zhou et al. | 2024 | arXiv.org | 63 | Generative-agent societies & multi-agent simulation |
| 295 | HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning | Qihao Yang et al. | 2025 | AAAI Conference on Artificial Intelligence | 0 | Cognitive/economic experiments (replication) |
| 296 | Deep Binding of Language Model Virtual Personas: a Study on Approximating Political Partisan Misperceptions | Minwoo Kang et al. | 2025 | 5 | Cognitive/economic experiments (replication) | |
| 297 | YuLan-OneSim: Towards the Next Generation of Social Simulator with Large Language Models | Lei Wang et al. | 2025 | arXiv.org | 13 | AI as substitute for human annotators/subjects |
| 298 | Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games | Zirui Cheng et al. | 2026 | 0 | Decision-making, biases & trust | |
| 299 | Psychometric Alignment: Capturing Human Knowledge Distributions via Language Models | Joy He-Yueya et al. | 2024 | arXiv.org | 15 | Preferences, recommendation & user simulation |
| 300 | Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks | Candida M. Greco et al. | 2026 | arXiv.org | 2 | AI as substitute for human annotators/subjects |
| 301 | Disentangling Interaction and Bias Effects in Opinion Dynamics of Large Language Models | Vincent C. Brockers et al. | 2025 | arXiv.org | 2 | Public opinion & political representation |
| 302 | Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models | Yuxuan Li et al. | 2025 | Conference on Fairness, Accountability and Transparency | 34 | Decision-making, biases & trust |
| 303 | Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Prediction | Sarah Ball et al. | 2025 | arXiv.org | 4 | Public opinion & political representation |
| 304 | Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules | Yanbei Jiang et al. | 2025 | arXiv.org | 1 | Cognitive/economic experiments (replication) |
| 305 | Large Language Models as Simulative Agents for Neurodivergent Adult Psychometric Profiles | Francesco Chiappone et al. | 2026 | arXiv.org | 0 | Silicon sampling & survey simulation |
| 306 | InsurAgent: A Large Language Model-Empowered Agent for Simulating Individual Behavior in Purchasing Flood Insurance | Ziheng Geng et al. | 2025 | arXiv.org | 1 | Preferences, recommendation & user simulation |
| 307 | Valuing Time in Silicon: Can Large Language Models Replicate Human Value of Travel Time | Yingnan Yan et al. | 2025 | Travel Behaviour & Society | 4 | Cognitive/economic experiments (replication) |
| 308 | Language Models Trained on Media Diets Can Predict Public Opinion | Eric Chu et al. | 2023 | arXiv.org | 44 | Public opinion & political representation |
| 309 | Vox Populi, Vox AI? Using Language Models to Estimate German Public Opinion | Leah von der Heyde et al. | 2024 | Social science computer review | 13 | Public opinion & political representation |
| 310 | Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study | Bolei Ma et al. | 2024 | Annual Meeting of the Association for Computational Linguistics | 17 | Public opinion & political representation |
| 311 | Parametric Social Identity Injection and Diversification in Public Opinion Simulation | Hexi Wang et al. | 2026 | arXiv.org | 0 | Public opinion & political representation |
| 312 | POSIM: A Multi-Agent Simulation Framework for Social Media Public Opinion Evolution and Governance | Yongmao Zhang et al. | 2026 | arXiv.org | 1 | Generative-agent societies & multi-agent simulation |
| 313 | A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies | Weihong Qi et al. | 2025 | 2025 IEEE International Conference on Data Mining Workshops (ICDMW) | 1 | Public opinion & political representation |
| 314 | A Guide to Using Social Media as a Geospatial Lens for Studying Public Opinion and Behavior | Lingyao Li | 2026 | arXiv.org | 0 | Public opinion & political representation |
| 315 | Whose Opinions Do Language Models Reflect? | Shibani Santurkar et al. | 2023 | International Conference on Machine Learning | 888 | Public opinion & political representation |
| 316 | The Persuasive Power of Large Language Models | Simon Martin Breum et al. | 2023 | International Conference on Web and Social Media | 84 | Generative-agent societies & multi-agent simulation |
| 317 | Designing Domain-Specific Large Language Models: The Critical Role of Fine-Tuning in Public Opinion Simulation | Haocheng Lin | 2024 | arXiv.org | 5 | Public opinion & political representation |
| 318 | In Silico Sociology: Forecasting COVID-19 Polarization with Large Language Models | Austin C. Kozlowski et al. | 2024 | arXiv.org | 6 | Public opinion & political representation |
| 319 | Improving Language Model Personas via Rationalization with Psychological Scaffolds | Brihi Joshi et al. | 2025 | Conference on Empirical Methods in Natural Language Processing | 10 | Preferences, recommendation & user simulation |
| 320 | Large Language Models as Subpopulation Representative Models: A Review | Gabriel Simmons & Christopher Hare | 2023 | arXiv.org | 22 | Public opinion & political representation |
| 321 | GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy | Jan Batzner et al. | 2024 | Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society | 9 | Cognitive/economic experiments (replication) |
| 322 | How Large Language Models Systematically Misrepresent American Climate Opinions | Sola Kim et al. | 2025 | arXiv.org | 0 | Public opinion & political representation |
| 323 | The Hidden Bias: A Study on Explicit and Implicit Political Stereotypes in Large Language Models | Konrad Löhr et al. | 2025 | arXiv.org | 2 | Silicon sampling & survey simulation |
| 324 | Large Language Model Agent for Fake News Detection | Xinyi Li et al. | 2024 | arXiv.org | 25 | Decision-making, biases & trust |
| 325 | Donald Trumps in the Virtual Polls: Simulating and Predicting Public Opinions in Surveys Using Large Language Models | Shapeng Jiang et al. | 2024 | 12 | Silicon sampling & survey simulation | |
| 326 | Synthetic Social Media Influence Experimentation via an Agentic Reinforcement Learning Large Language Model Bot | Bailu Jin & Weisi Guo | 2024 | Journal of Artificial Societies and Social Simulation | 3 | Generative-agent societies & multi-agent simulation |
| 327 | Aligning Language Models to User Opinions | EunJeong Hwang et al. | 2023 | Conference on Empirical Methods in Natural Language Processing | 123 | Public opinion & political representation |
| 328 | MindVote: When AI Meets the Wild West of Social Media Opinion | Xutao Mao et al. | 2025 | AAAI Conference on Artificial Intelligence | 0 | Public opinion & political representation |
| 329 | United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections | Leah von der Heyde et al. | 2024 | arXiv.org | 8 | Public opinion & political representation |
| 330 | PrimeX: A Dataset of Worldview, Opinion, and Explanation | Rik Koncel-Kedziorski et al. | 2025 | Conference on Empirical Methods in Natural Language Processing | 1 | Public opinion & political representation |
| 331 | Social Opinions Prediction Utilizes Fusing Dynamics Equation with LLM-based Agents | Junchi Yao et al. | 2024 | Scientific Reports | 17 | Public opinion & political representation |
| 332 | An Empirical Study of Group Conformity in Multi-Agent Systems | Min Choi et al. | 2025 | Annual Meeting of the Association for Computational Linguistics | 5 | Generative-agent societies & multi-agent simulation |
| 333 | Bias-Adjusted LLM Agents for Human-Like Decision-Making via Behavioral Economics | Ayato Kitadai et al. | 2025 | arXiv.org | 3 | Decision-making, biases & trust |
| 334 | Decision and Gender Biases in Large Language Models: A Behavioral Economic Perspective | Luca Corazzini et al. | 2025 | arXiv.org | 0 | Decision-making, biases & trust |
| 335 | Information Design With Large Language Models | Paul Duetting et al. | 2025 | arXiv.org | 1 | Generative-agent societies & multi-agent simulation |
| 336 | TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets | Yuzhe Yang et al. | 2025 | arXiv.org | 45 | Generative-agent societies & multi-agent simulation |
| 337 | Risk Profiling and Modulation for LLMs | Yikai Wang et al. | 2025 | arXiv.org | 1 | Cognitive/economic experiments (replication) |
| 338 | Representation Without Control: Testing the Realization Effect in Language Models | Ciarán Walsh & Emilio Barkett | 2026 | 0 | Decision-making, biases & trust | |
| 339 | Interpolative Decoding: Exploring the Spectrum of Personality Traits in LLMs | Eric Yeh et al. | 2025 | arXiv.org | 0 | Personality & psychological profiling of LLMs |
| 340 | Sparks of Rationality: Do Reasoning LLMs Align with Human Judgment and Choice? | Ala N. Tak et al. | 2026 | arXiv.org | 0 | Cognitive/economic experiments (replication) |
| 341 | HEART-Bench: Do LLM Agents Exhibit Human-like Psychology? | Weihan Peng et al. | 2026 | 0 | Personality & psychological profiling of LLMs | |
| 342 | PsychoGAT: A Novel Psychological Measurement Paradigm through Interactive Fiction Games with LLM Agents | Qisen Yang et al. | 2024 | Annual Meeting of the Association for Computational Linguistics | 46 | Personality & psychological profiling of LLMs |
| 343 | Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making | Ziv Ben-Zion et al. | 2025 | npj Artificial Intelligence | 0 | Decision-making, biases & trust |
| 344 | Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents | Zhiping Zhang et al. | 2025 | 2 | Generative-agent societies & multi-agent simulation | |
| 345 | Open Models, Closed Minds? On Agents Capabilities in Mimicking Human Personalities through Open Large Language Models | Lucio La Cava & Andrea Tagarelli | 2024 | AAAI Conference on Artificial Intelligence | 41 | Personality & psychological profiling of LLMs |
| 346 | Psychologically Enhanced AI Agents | Maciej Besta et al. | 2025 | arXiv.org | 3 | Personality & psychological profiling of LLMs |
| 347 | MIND: Towards Immersive Psychological Healing with Multi-agent Inner Dialogue | Yujia Chen et al. | 2025 | Conference on Empirical Methods in Natural Language Processing | 6 | Decision-making, biases & trust |
| 348 | Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection | Tharindu Kumarage et al. | 2025 | arXiv.org | 17 | Preferences, recommendation & user simulation |
| 349 | Persona Alchemy: Designing, Evaluating, and Implementing Psychologically-Grounded LLM Agents for Diverse Stakeholder Representation | Sola Kim et al. | 2025 | arXiv.org | 4 | Generative-agent societies & multi-agent simulation |
| 350 | LLMs are Introvert | Litian Zhang et al. | 2025 | arXiv.org | 0 | Generative-agent societies & multi-agent simulation |
| 351 | Exploring a Gamified Personality Assessment Method through Interaction with LLM Agents Embodying Different Personalities | Baiqiao Zhang et al. | 2025 | 0 | Personality & psychological profiling of LLMs | |
| 352 | CogniPair: From LLM Chatbots to Conscious AI Agents -- GNWT-Based Multi-Agent Digital Twins for Social Pairing -- Dating & Hiring Applications | Wanghao Ye et al. | 2025 | arXiv.org | 0 | Generative-agent societies & multi-agent simulation |
| 353 | Can LLM Agents Maintain a Persona in Discourse? | Pranav Bhandari et al. | 2025 | Conference on Empirical Methods in Natural Language Processing | 18 | Personality & psychological profiling of LLMs |
| 354 | Structured Personality Control and Adaptation for LLM Agents | Jinpeng Wang et al. | 2026 | arXiv.org | 2 | Personality & psychological profiling of LLMs |
| 355 | Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models | Yin Jou Huang & Rafik Hadfi | 2025 | Conference on Empirical Methods in Natural Language Processing | 4 | Personality & psychological profiling of LLMs |
| 356 | HamRaz: A Culture-Based Persian Conversation Dataset for Person-Centered Therapy Using LLM Agents | Mohammad Amin Abbasi et al. | 2025 | Proceedings of The FirstWorkshop on Natural Language Processing and Language Models for Digital Humanities | 4 | Preferences, recommendation & user simulation |
| 357 | How Personality Traits Influence Negotiation Outcomes? A Simulation based on Large Language Models | Yin Jou Huang & Rafik Hadfi | 2024 | Conference on Empirical Methods in Natural Language Processing | 6 | Personality & psychological profiling of LLMs |
| 358 | LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling | Hanlin Sun & Jiayang Li | 2025 | arXiv.org | 2 | Generative-agent societies & multi-agent simulation |
| 359 | Adaptive LLM Agents: Toward Personalized Empathetic Care | Priyanka Singh & Sebastian Von Mammen | 2025 | arXiv.org | 0 | Preferences, recommendation & user simulation |
| 360 | Personality Anchoring for Social Simulation: Linking Personality, Social Behavior, and Interaction Success with LLM Agents | Vahid Sadiri Javadi et al. | 2026 | Proceedings of the Language Resources and Evaluation Conference | 0 | Personality & psychological profiling of LLMs |
| 361 | Spiral of Silence in Large Language Model Agents | Mingze Zhong et al. | 2025 | Conference on Empirical Methods in Natural Language Processing | 0 | Generative-agent societies & multi-agent simulation |
| 362 | The Yerkes-Dodson Curve for AI Agents: Emergent Cooperation Under Environmental Pressure in Multi-Agent LLM Simulations | Ivan Pasichnyk | 2026 | arXiv.org | 0 | Generative-agent societies & multi-agent simulation |
| 363 | LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings | Benjamin F. Maier et al. | 2025 | arXiv.org | 10 | Decision-making, biases & trust |
| 364 | A Large-Scale Simulation on Large Language Models for Decision-Making in Political Science | Chenxiao Yu et al. | 2024 | 8 | Cognitive/economic experiments (replication) | |
| 365 | Intelligent Computing Social Modeling and Methodological Innovations in Political Science in the Era of Large Language Models | Zhenyu Wang et al. | 2024 | Journal of Chinese Political Science | 20 | Generative-agent societies & multi-agent simulation |
| 366 | Political Actor Agent: Simulating Legislative System for Roll Call Votes Prediction with Large Language Models | Hao Li et al. | 2024 | arXiv.org | 6 | Generative-agent societies & multi-agent simulation |
| 367 | Sense and Sensitivity: Evaluating the simulation of social dynamics via Large Language Models | Da Ju et al. | 2024 | arXiv.org | 3 | Generative-agent societies & multi-agent simulation |
| 368 | A Survey on Human-Centric LLMs | Jing Yi Wang et al. | 2024 | arXiv.org | 26 | Cognitive/economic experiments (replication) |
| 369 | Social Simulacra: Creating Populated Prototypes for Social Computing Systems | 2022 | ACM Symposium on User Interface Software and Technology | 477 | (unclassified — from get-smart lists) | |
| 370 | Large language models that replace human participants can harmfully misportray and flatten identity groups | 2024 | 146 | (unclassified — from get-smart lists) | ||
| 371 | Automated Social Science: Language Models as Scientist and Subjects | 2024 | Social Science Research Network | 130 | (unclassified — from get-smart lists) | |
| 372 | LLM Social Simulations Are a Promising Research Method | 2025 | arXiv.org | 127 | (unclassified — from get-smart lists) | |
| 373 | Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy | 2024 | Science Advances | 91 | (unclassified — from get-smart lists) | |
| 374 | LLM Generated Persona is a Promise with a Catch | 2025 | arXiv.org | 85 | (unclassified — from get-smart lists) | |
| 375 | Questioning the Survey Responses of Large Language Models | 2023 | Neural Information Processing Systems | 77 | (unclassified — from get-smart lists) | |
| 376 | CogBench: a large language model walks into a psychology lab | 2024 | International Conference on Machine Learning | 67 | (unclassified — from get-smart lists) | |
| 377 | Beyond Demographics: Aligning Role-playing LLM-based Agents Using Human Belief Networks | 2024 | Conference on Empirical Methods in Natural Language Processing | 47 | (unclassified — from get-smart lists) | |
| 378 | PATIENT-\psi: Using Large Language Models to Simulate Patients for Training Mental Health Professionals | 2024 | Conference on Empirical Methods in Natural Language Processing | 30 | (unclassified — from get-smart lists) | |
| 379 | Language Models Trained to do Arithmetic Predict Human Risky and Intertemporal Choice | 2024 | International Conference on Learning Representations | 10 | (unclassified — from get-smart lists) | |
| 380 | Using Large Language Models to Create AI Personas for Replication and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings | 2024 | arXiv.org | 7 | (unclassified — from get-smart lists) | |
| 381 | Sentiment Simulation using Generative AI Agents | 2025 | arXiv.org | 1 | (unclassified — from get-smart lists) |