About

I am a Senior Responsible AI Data Scientist at Humana, where I build evaluation and safety systems for foundation models deployed across healthcare and enterprise AI applications.

My work spans LLM-as-a-Judge, automated evaluation pipelines, guardrail frameworks, and AI risk assessment, helping teams measure model quality, identify failure modes, and implement safeguards that improve the reliability and trustworthiness of AI systems.

I completed my Ph.D. in Computer Science at Brown University, co-advised by Carsten Eickhoff and Ritambhara Singh. My thesis, Towards Trustworthy Clinical AI, spanned knowledge grounding, inference reliability, and behavioral control. I collaborate with Dr. Hamish Fraser and Bio-RAMP Labs on clinical decision support, and with the Masakhane community on speech and multimodal evaluation to widen healthcare access across low- and middle-income countries.

Before Brown I earned an M.Sc. at the University of Cape Town with Geoff Nitschke, and a B.Sc. at Bayero University Kano.

AI Safety LLM Evaluation Responsible AI Human-AI Interaction Multilingual NLP
Research areas

What I work on

01

LLM Evaluation & Benchmarking

Evaluation frameworks and benchmarks for the safety, factuality, and robustness of LLMs, probing adversarial vulnerabilities, the limits of self-evaluation, and failure modes in generative systems.

02

Clinical Decision Support

With Dr. Hamish Fraser and Bio-RAMP Labs, evaluating the safety, reliability, and clinical applicability of generative AI for diagnostic reasoning, with a focus on low-resource healthcare settings.

03

Global Accessibility & Multilingual AI

Through the Masakhane community, evaluating speech recognition and multimodal LLMs to improve healthcare accessibility in low- and middle-income countries.

Recent news

News

2026
Joined Humana as Senior Responsible AI Data Scientist, building LLM evaluation and responsible-AI frameworks.
May 2026
Invited talk at Google Health AI: Engineering Trustworthy Clinical AI: Knowledge Grounding, Inference Reliability, and Behavioral Control.
2026
UbuntuGuard, on local-policy approaches to AI safety, was featured by the Burnes Center for Social Change.
2025
Two papers accepted to NeurIPS 2025.
2025
AfriMed-QA won the Best Social Impact Award at ACL 2025, Vienna.
2025
Presented knowledge-graph reasoning for explainable AI-driven drug discovery at KDD 2025, Toronto.
2025
Completed a summer internship at Optum AI, UnitedHealth Group.
Selected publications

Publications

Author list abbreviated; Abdullahi, T. shown in bold. Full list on Google Scholar.

Model evaluation, safety & behavior

The Persona Paradox: Medical Personas as Behavioral Priors in Clinical Language Models
Abdullahi, T., et al.
ACL ARR 2026
Position: Benchmarking is Broken — Don't Let AI Be Its Own Judge
Cheng, Z., Wohnig, S., Gupta, R., Alam, S., Abdullahi, T., et al.
NeurIPS 2025 Position paper Paper ↗
UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages
Abdullahi, T., et al.
ACL 2026 Feature ↗
AfriVox: Probing Multilingual and Accent Robustness of Speech & Multimodal LLMs
Abdullahi, T., et al.
EACL 2026
LLM-Powered Graph Reasoning for Knowledge Discovery
Gemou, I., Abdullahi, T., Singh, R.
WiML @ NeurIPS 2025 OpenReview ↗

Foundation models, reasoning & retrieval

K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
Abdullahi, T., Gemou, I., Nayak, N., Murtaza, G., et al.
KDD 2025 arXiv ↗
Retrieval-Augmented Zero-Shot Text Classification
Abdullahi, T., Singh, R., Eickhoff, C.
ACM SIGIR ICTIR 2024 PDF ↗

Generative AI in healthcare

AfriMed-QA: Towards a Pan-African, Multi-Specialty Medical Question-Answering Benchmark Dataset
Olatunji, T., Charles, N., Owodunni, A., Yuehgoh, F., Abdullahi, T., et al.
ACL 2025 ★ Best Social Impact Award Dataset ↗
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond
Sanni, M., Abdullahi, T., et al.
NAACL 2025 arXiv ↗
Identifying and Timing Patient Outcomes in Clinician Notes Using Large Language Models
Abdullahi, T., Hamzeh, A., Sears, I., Abadi, N., Singh, R., Eickhoff, C., Abbasi, A.
Artificial Intelligence in Medicine 2026
Diagnostic Accuracy of ChatGPT-4.0 for TIA or Stroke Using Patient Symptoms and Demographic Data
Khatri, I., Zahiri, A., Abdullahi, T., et al.
Int'l Stroke Conference 2025 Paper ↗
Retrieval-Based Diagnostic Decision Support: A Mixed-Methods Study
Abdullahi, T., Mercurio, L., Singh, R., Eickhoff, C.
JMIR Medical Informatics 2024 Paper ↗
Learning to Make Rare and Complex Diagnoses with Generative AI Assistance
Abdullahi, T., Singh, R., Eickhoff, C.
JMIR Medical Education 2024 Paper ↗