Haseeb Ullah ORCID iD Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University Pakistan
Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University, Dera Ismail Khan, Khyber Pakhtunkhwa, Pakistan.
h2seeb@gmail.com,
https://orcid.org/0009-0003-4259-7298
Husnain Saleem ORCID iD Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University Pakistan
Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University, Dera Ismail Khan, Khyber Pakhtunkhwa, Pakistan.
jilani.husnain@yahoo.com, https://orcid.org/0009-0001-7513-1086
Asia Kanwal ORCID iD University of Agriculture Pakistan
Department of Artificial Intelligence, University of Agriculture, Dera Ismail Khan, Khyber Pakhtunkhwa, Pakistan. aassia.baloch@gmail.com, https://orcid.org/0009-0004-2548-5582
Muhammad Javed ORCID iD Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University Pakistan
Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University, Dera Ismail Khan, Khyber Pakhtunkhwa, Pakistan. javed_gomal@gu.edu.pk, https://orcid.org/0000-0001-6884-6641
Hamid Masood Khan ORCID iD Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University Pakistan
Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University, Dera Ismail Khan, Khyber Pakhtunkhwa, Pakistan.
hamidmasoodkhan@gu.edu.pk, https://orcid.org/0000-0001-9235-9886
Kiran Hanif ORCID iD Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University Pakistan
Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University, Dera Ismail Khan, Khyber Pakhtunkhwa, Pakistan.
safdarkiran3@gmail.com,
https://orcid.org/0009-0000-7573-7515
User
p-ISSN: 2068 - 0473 e-ISSN: 2067 - 3957
DOI:
10.18662/brain
DOI prefix: 10.70594/brain (currently edited by EduSoft) | 10.18662/brain (when was edited by Lumen)
Frequency:
4 issues/year (occasional additional issues)
Abstracting & Indexing
Web of Science (ESCI, IF 0.6), EBSCO, Google Scholar etc.
Haseeb Ullah -
Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University (PK),
Husnain Saleem -
Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University (PK),
Asia Kanwal -
University of Agriculture (PK),
Muhammad Javed -
Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University (PK),
Hamid Masood Khan -
Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University (PK),
Kiran Hanif -
Gomal Research Institute of Computing (GRIC), Faculty of Computing, Gomal University (PK),
Abstract
Automated Urdu hate-speech detection remains challenging because annotated resources are limited, orthography varies, and harmful meaning depends on both lexical and contextual cues. This study evaluates a leakage-safe lexical-semantic framework combining 5,000-dimensional word/bigram TF-IDF features with 768-dimensional multilingual sentence embeddings, followed by binary particle swarm optimisation (BPSO) and classical ensemble learning. Experiments used the Urdu text and binary labels from MMHS11K: 8,800 balanced training records and an untouched balanced test set of 2,200 records. BPSO selected 2,849 of 5,768 hybrid dimensions, reducing dimensionality by 50.61%. On the official test set, Stacking with the complete hybrid representation achieved Macro-F1 = 0.8468 and ROC-AUC = 0.9265; the PSO-selected representation achieved Macro-F1 = 0.8391 and ROC- AUC = 0.9196. After mask selection, classifier fitting time decreased by 52.24%, excluding sentence-embedding extraction and BPSO search. Leakage-safe nested five-fold validation showed a similar trade-off: fitting time decreased by 52.64%, with Macro-F1 = 0.8444 ± 0.0066 after PSO versus 0.8492 ± 0.0110 without PSO. Soft Voting and Stacking were statistically comparable on selected features. Three-seed transformer aggregates performed better: Urdu-RoBERTa achieved the highest Macro-F1 (0.8836), and XLM-R achieved the highest ROC-AUC (0.9529). Holm-corrected exact McNemar tests confirmed that transformer aggregates significantly outperformed PSO Stacking. Overall, BPSO approximately halves representation size and downstream classical-model fitting time, with a small predictive decrease that was not statistically significant relative to complete-hybrid Stacking after correction.
Academic discipline and sub-disciplines:
Artificial Intelligence; Education; Linguistics