Structural and Behavioral Analysis of Fine-Tuning SLMs for Domain Adaptation
International Journal of Software Engineering and Knowledge Engineering, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Basım Tarihi: 2026
- Doi Numarası: 10.1142/s0218194026500518
- Dergi Adı: International Journal of Software Engineering and Knowledge Engineering
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, Applied Science & Technology Source, Compendex, INSPEC, Business Source Ultimate (EBSCO), Engineering Source (EBSCO), Technology Collection (ProQuest)
- Anahtar Kelimeler: domain generalization, Gemma, Legal domain adaptation, LoRA adaptation, low-resource languages, scaling laws, Turkish NLP
- Süleyman Demirel Üniversitesi Adresli: Evet
Özet
This study examines how small language models (SLMs) behave under domain-specific fine-tuning in Turkish, a non-English and comparatively lower-resource language. Seven different models with varying sizes and architectures are evaluated across legal reasoning, reading comprehension, analytical abilities, and conceptual knowledge. Fine-tuning sets range from 100 to 10,000 supervised legal examples, enabling systematic analysis of data-scaling effects in the fine-tuning regime. The results show that performance does not improve monotonically with increased training data. Instead, models consistently exhibit a narrow efficiency zone, typically at 100–400 examples, where legal accuracy peaks before degrading due to over-specialization. Larger datasets amplify LoRA update magnitude and density without corresponding gains in accuracy, revealing a divergence between optimization progress and generalization. These effects are attributed to the fragmented nature of legal knowledge, which limits cross-topic transfer and makes models highly sensitive to overshooting the optimal fine-tuning range. Smaller end of considered models seem to lack the representational capacity for stable legal reasoning, whereas larger ones overfit rapidly without careful control. Overall, the study demonstrates that no general-purpose fine-tuning recipe can be assumed effective without comprehensive evaluation across multiple cognitive dimensions for SLMs.