Arena-style human evaluation platform for LLMs across 11 Arabic dialects — 1,100+ pairwise comparisons over 67 models, ranked via Bradley–Terry (ESS = 582).
Hassan Barmandah, Fatimah Emad Eldin, Khloud Al Jallad, Omer Nacar
Presented at 1st Research Conference for Students of Makkah Region Universities (RCUSM), Jeddah.
Hassan Barmandah, et al.
Beyond Mean Absolute Error: Ancestry-Stratified Calibration and Explainability for Warfarin Dosing Models
MDPI Journal of Personalized Medicine · Under Review
Ancestry-stratified calibration and explainability analysis for machine learning-based warfarin dosing models.
Hassan Barmandah, et al. (NAMAA Community)
NAMAA Community at AraGenre 2026: From Encoder Baselines to Self-Consistent LLM Ensembling for Hierarchical Arabic Genre Classification
Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026) · [GitHub]
Zero-shot LLM classification with self-consistency sampling reached 0.7013 hierarchical macro F1, ranking 3rd of 18 teams.
Hassan Barmandah, et al. (NAMAA Community)
NAMAA Community at DialectSentEval 2026 Subtask 1: Dialect-Aware Arabic Sentiment Classification with Pretrained Transformer Models
Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026)
Dialect-aware sentiment classification across Moroccan, Egyptian, Jordanian, and Saudi Arabic — MARBERTv2 achieved the strongest result at 94.07% macro F1.
Hassan Barmandah, et al. (NAMAA Community)
NAMAA Community at HalluScoring 2026: Beyond Predictive Performance: Reliability and Generalization in Arabic Hallucination Detection
Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026) · [GitHub]
CAMeLBERT cross-encoder hallucination detector achieved 0.7596 ROC-AUC, ranking 3rd of 7 teams; temperature scaling and MC-Dropout deferral further improved calibration and reliability.
Hassan Barmandah, et al. (NAMAA Community)
NAMAA Community at StanceEval-2026: From Encoder and Decoder Model Fine-Tuning to Data-Centric Retrieval-Augmented In-Context Learning for Arabic Stance Detection
Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026)
Retrieval-augmented in-context learning ensemble for Arabic stance detection achieved Favg2 of 0.8719 (Track 1, ranked 3rd) and 0.9016 (Track 2, ranked 7th).
Hassan Barmandah, et al. (NAMAA Community)
NAMAA at DialectSentEval 2026 Subtask 2: From Reranked Sampling to Cross-Backbone Ensembling for Arabic Dialectal Sentiment Swap
Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026)
Cross-backbone ensembling for Arabic dialectal sentiment-polarity swap reached 96.78 development accuracy; best official test score was 0.7597.
Hassan Barmandah, et al.
Sarab: A Cause-Diagnostic Arabic Visual Hallucination Evaluation Benchmark for Multimodal Large Language Models
In Preparation
Cause-diagnostic Arabic visual hallucination benchmark for multimodal LLMs, evaluated across four models.
Hassan Barmandah
Sawb: Cultural Hallucination Detection in Arabic LLMs
4-model Arabic encoder ensemble (AraBERT, AraBERT-Large, ARBERTv2, MARBERTv2) for detecting cultural hallucinations — 0.9647 macro F1 on 457 validation examples, with Arabic explanations generated for all 163 detected hallucinations.
Hassan Barmandah, et al.
WarfaRisk: A Reproducible, Fairness-Aware Machine Learning Pipeline for Warfarin Dose Prediction (Prototype)
MDPI AI · Under Review
Reproducible, fairness-aware pipeline for warfarin dose prediction — leakage-free preprocessing, ablation across 9 model architectures, leave-one-ancestry-out evaluation, conformal prediction calibration, SHAP explainability, and a FHIR R4 clinical decision-support interface.