Effectiveness of Chain of Thought Strategies to Mitigate Sycophancy in Quantized Small Language Models
International Journal on Artificial Intelligence Tools, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Basım Tarihi: 2026
- Doi Numarası: 10.1142/s021821302650017x
- Dergi Adı: International Journal on Artificial Intelligence Tools
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, Applied Science & Technology Source, Compendex, Academic Search Ultimate (EBSCO), Business Source Ultimate (EBSCO), Engineering Source (EBSCO), Technology Collection (ProQuest)
- Anahtar Kelimeler: Chain-of-Thought (CoT), confidence calibration, explicit reasoning, implicit reasoning, low-resource languages, quantization
- Süleyman Demirel Üniversitesi Adresli: Evet
Özet
Small Language Models (SLMs) deployed in resource-constrained environments are increasingly susceptible to “sycophancy”, a behavioral failure where models prioritize user-suggested biases over factual truth. Although mitigation strategies exist for highresource settings, the effectiveness of these interventions remains largely unexplored for morphologically rich languages where subword tokenization disrupts logical coherence. To address this problem, this study investigates the efficacy of Chain-of-Thought (CoT) strategies in mitigating sycophancy within a Turkish context utilizing a 4-bit quantized Gemma-3-12B model. We evaluated performance across legal, analytical, knowledgebased, and reading comprehension domains under four conditions, namely Neutral, Sycophancy (bias injection), Implicit (silent) CoT, and Explicit (verbalized) CoT. Results reveal that while bias injection caused a catastrophic accuracy drop (from 49.8% to 17.4%), the considered silent reasoning approach failed to provide resistance against the bias. Conversely, Explicit CoT significantly mitigated sycophancy, recovering accuracy to 32.8% and correcting confidence discrepancy by reducing the reported confidence when wrong. However, this robustness incurs a six-fold increase in inference latency and a fifteen-fold increase in token consumption. We conclude that while silent reasoning is insufficient for low-resource languages, verbal reasoning serves as a necessary but costly protective constraint for high-stakes tasks.