Publications

Mourad Mars, Hassan Barmandah, Abdulrhman Alassaf

Haakkim: An Arena-Style Human Preference Evaluation Platform for Arabic LLMs

IEEE Access · Under Review · Jul 2026 · [Website] [Hugging Face]

Arena-style human evaluation platform for LLMs across 11 Arabic dialects — 1,100+ pairwise comparisons over 67 models, ranked via Bradley–Terry (ESS = 582).

Hassan Barmandah, Fatimah Emad Eldin, Khloud Al Jallad, Omer Nacar

Ketaba-OCR at AR-MS NakbaNLP 2026: Efficient Adaptation of Vision-Language Models for Hand Written Recognition

Proceedings of LREC 2026 · NakbaNLP 2026 Shared Task · May 2026 · [Model]

CER = 0.0819, WER = 0.2588 on blind per-line evaluation — ranked 1st in that setting.

Hassan Barmandah, Fatimah Emad Eldin, Omer Nacar

Fine-Tashkeel at KSAA-2026: A Comprehensive Evaluation of Seq2Seq and Multimodal Approaches for Automatic Diacritization of Arabic Speech Dictation

Proceedings of LREC 2026 · KSAA-2026 Shared Task, OSACT7 Workshop · May 2026 · [GitHub]

DER = 10.56%, WER = 34.47%, SER = 79.88% — 5th of 7 teams in the KSAA-2026 shared task.

Hassan Barmandah

Saudi-Dialect-ALLaM: LoRA Fine-Tuning for Dialectal Arabic Generation

arXiv · Aug 2025 · [arXiv] [GitHub]

Presented at 1st Research Conference for Students of Makkah Region Universities (RCUSM), Jeddah.

Hassan Barmandah, et al.

Beyond Mean Absolute Error: Ancestry-Stratified Calibration and Explainability for Warfarin Dosing Models

MDPI Journal of Personalized Medicine · Under Review

Ancestry-stratified calibration and explainability analysis for machine learning-based warfarin dosing models.

Hassan Barmandah, et al. (NAMAA Community)

NAMAA Community at AraGenre 2026: From Encoder Baselines to Self-Consistent LLM Ensembling for Hierarchical Arabic Genre Classification

Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026) · [GitHub]

Zero-shot LLM classification with self-consistency sampling reached 0.7013 hierarchical macro F1, ranking 3rd of 18 teams.

Hassan Barmandah, et al. (NAMAA Community)

NAMAA Community at DialectSentEval 2026 Subtask 1: Dialect-Aware Arabic Sentiment Classification with Pretrained Transformer Models

Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026)

Dialect-aware sentiment classification across Moroccan, Egyptian, Jordanian, and Saudi Arabic — MARBERTv2 achieved the strongest result at 94.07% macro F1.

Hassan Barmandah, et al. (NAMAA Community)

NAMAA Community at HalluScoring 2026: Beyond Predictive Performance: Reliability and Generalization in Arabic Hallucination Detection

Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026) · [GitHub]

CAMeLBERT cross-encoder hallucination detector achieved 0.7596 ROC-AUC, ranking 3rd of 7 teams; temperature scaling and MC-Dropout deferral further improved calibration and reliability.

Hassan Barmandah, et al. (NAMAA Community)

NAMAA Community at StanceEval-2026: From Encoder and Decoder Model Fine-Tuning to Data-Centric Retrieval-Augmented In-Context Learning for Arabic Stance Detection

Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026)

Retrieval-augmented in-context learning ensemble for Arabic stance detection achieved Favg2 of 0.8719 (Track 1, ranked 3rd) and 0.9016 (Track 2, ranked 7th).

Hassan Barmandah, et al. (NAMAA Community)

NAMAA at DialectSentEval 2026 Subtask 2: From Reranked Sampling to Cross-Backbone Ensembling for Arabic Dialectal Sentiment Swap

Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026)

Cross-backbone ensembling for Arabic dialectal sentiment-polarity swap reached 96.78 development accuracy; best official test score was 0.7597.

Hassan Barmandah, et al.

Sarab: A Cause-Diagnostic Arabic Visual Hallucination Evaluation Benchmark for Multimodal Large Language Models

In Preparation

Cause-diagnostic Arabic visual hallucination benchmark for multimodal LLMs, evaluated across four models.

Hassan Barmandah

Sawb: Cultural Hallucination Detection in Arabic LLMs

arXiv Preprint · ICAIRE 2026 Hackathon Track 3 · [Hugging Face] [GitHub]

4-model Arabic encoder ensemble (AraBERT, AraBERT-Large, ARBERTv2, MARBERTv2) for detecting cultural hallucinations — 0.9647 macro F1 on 457 validation examples, with Arabic explanations generated for all 163 detected hallucinations.

Hassan Barmandah, et al.

WarfaRisk: A Reproducible, Fairness-Aware Machine Learning Pipeline for Warfarin Dose Prediction (Prototype)

MDPI AI · Under Review

Reproducible, fairness-aware pipeline for warfarin dose prediction — leakage-free preprocessing, ablation across 9 model architectures, leave-one-ancestry-out evaluation, conformal prediction calibration, SHAP explainability, and a FHIR R4 clinical decision-support interface.