Investigating Persuasiveness in Large Language Models
Published in M.A. Thesis, Princeton University, 2023
Recommended citation: Promise Ekpo. "Investigating Persuasiveness in Large Language Models." M.A. Thesis, Princeton University, 2023.
M.A. thesis (advised by Prof. Jaime Fernández Fisac) quantifying the persuasiveness of GPT-3/4 using a game-theoretic framework, demonstrating a 12-50% belief shift toward false statements (e.g., prior 0.58 to posterior 0.79 for GPT-3). The work uncovered systematic social bias (e.g., misleading arguments generated for one country but refused for another on identical false claims), trained LLMs with PPO on a persuasiveness reward signal (observing emergent manipulation strategies), and proposed reward functions for persuasiveness and truthfulness as an extension of the BIG-BENCH benchmark, contributing to AI safety frameworks for responsible deployment.
