While large language models (LLMs) have demonstrated remarkable capabilities, they often exhibit significant social biases. However, conventional evaluation frameworks are largely limited to detecting a narrow set of explicit social stereotypes, faili...
While large language models (LLMs) have demonstrated remarkable capabilities, they often exhibit significant social biases. However, conventional evaluation frameworks are largely limited to detecting a narrow set of explicit social stereotypes, failing to capture the subtle and implicit biases inherent in modern models. To address this, this study redefines implicit bias as behavioral inconsistencies in model decision-making and presents a comprehensive framework for diagnosing and interpreting these implicit social biases through the lens of defeasible reasoning, a non-monotonic reasoning where inferences are subject to revision based on new information. By adopting a defeasible reasoning framework, this research examines how demographic cues disproportionately influence model reasoning.
For this purpose, this study proposes an evaluation framework comprising a dedicated dataset, curated from benign social commonsense and natural language inference corpora and adapted for five bias axes, alongside new metrics designed to quantify prediction inconsistencies. Empirical evaluation of ten instruction-tuned LLMs reveals that models exhibit systematic biases that align with real-world social hierarchies, where introducing cues to particular demographics triggers disproportionate shifts in their decisions. Notably, explicitly engaging in step-by-step reasoning further amplifies these biases, triggering the models’ latent stereotypes in intermediate reasoning steps.
In addition to empirical analysis, this study applies mechanistic interpretability techniques, specifically attribution patching and edge ablation, to localize the structural origins of these biases as a case study. Moving beyond prior interpretability research limited to explicit stereotypes, this analysis identifies that demographic-sensitive behavior difference is mediated by a remarkably sparse set of critical paths within a model. By shifting the focus from what models explicitly say to how they implicitly reason, this study provides a starting point for more systematic diagnosis and analysis of reasoning behavior and fairness in LLMs.