SeqGPT-RL: a reinforcement-aligned hybrid intelligent system for Turkish technical-support dialogs


ARISOY A., KÜÇÜKSİLLE E. U.

Journal of Supercomputing, cilt.82, sa.12, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 82 Sayı: 12
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1007/s11227-026-08757-2
  • Dergi Adı: Journal of Supercomputing
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, Applied Science & Technology Source, Compendex, INSPEC, zbMATH, Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO), Engineering Source (EBSCO), Materials Science & Engineering Collection (ProQuest), Technology Collection (ProQuest)
  • Anahtar Kelimeler: GPT-2 adaptation, Hybrid intelligent systems, Low-resource dialog, Metric-based reinforcement learning, SeqGAN, Turkish NLP
  • Süleyman Demirel Üniversitesi Adresli: Evet

Özet

We present SeqGPT-RL, a hybrid conversational architecture for Turkish technical-support dialog generation in a low-resource institutional setting. The proposed framework combines domain-adapted GPT-2 fine-tuning, SeqGAN-based adversarial sequence regularization, and metric-based, human-reference-guided reinforcement learning to improve both semantic alignment and practical response quality. The model is trained on a curated corpus of 1759 real-world question–answer pairs collected from university help-desk interactions. Experimental results show that the full GPT-2 + SeqGAN + PPO configuration outperforms its baseline variants, achieving a BERTScore F1 of 0.8273. As an additional robustness check, we also evaluated the proposed model using the Turkish-specific BERTurk encoder, which yielded a BERTScore F1 of 0.8285. This result suggests that the observed semantic gains remain stable under a language-appropriate evaluation setting. Ablation analysis further shows that adversarial sequence learning and semantic reward alignment make complementary contributions to the overall improvement. We further compare the proposed model with alternative configurations, including DPO-based variants, DistilGPT-2-based hybrids, and a retrieval-augmented generation (RAG) baseline. Although RAG performs better on selected surface-level overlap metrics, SeqGPT-RL provides superior semantic alignment in the target domain. In addition to offline evaluation, we report deployment-oriented observations related to latency and concurrent-load behavior, highlighting the relevance of the approach for real-world information systems. The findings suggest that staged hybrid optimization can offer a practically deployable alternative for low-resource, domain-specific dialog generation under latency, privacy, and infrastructure constraints. A complementary blinded human evaluation on held-out test queries further indicated that the proposed model was generally preferred over the fine-tuned GPT-2 baseline in terms of relevance, usefulness, and technical appropriateness.