I am a Senior Responsible AI Data Scientist at Humana, where I build evaluation and safety systems for foundation models deployed across healthcare and enterprise AI applications.
My work spans LLM-as-a-Judge, automated evaluation pipelines, guardrail frameworks, and AI risk assessment, helping teams measure model quality, identify failure modes, and implement safeguards that improve the reliability and trustworthiness of AI systems.
I completed my Ph.D. in Computer Science at Brown University, co-advised by Carsten Eickhoff and Ritambhara Singh. My thesis, Towards Trustworthy Clinical AI, spanned knowledge grounding, inference reliability, and behavioral control. I collaborate with Dr. Hamish Fraser and Bio-RAMP Labs on clinical decision support, and with the Masakhane community on speech and multimodal evaluation to widen healthcare access across low- and middle-income countries.
Before Brown I earned an M.Sc. at the University of Cape Town with Geoff Nitschke, and a B.Sc. at Bayero University Kano.
What I work on
LLM Evaluation & Benchmarking
Evaluation frameworks and benchmarks for the safety, factuality, and robustness of LLMs, probing adversarial vulnerabilities, the limits of self-evaluation, and failure modes in generative systems.
Clinical Decision Support
With Dr. Hamish Fraser and Bio-RAMP Labs, evaluating the safety, reliability, and clinical applicability of generative AI for diagnostic reasoning, with a focus on low-resource healthcare settings.
Global Accessibility & Multilingual AI
Through the Masakhane community, evaluating speech recognition and multimodal LLMs to improve healthcare accessibility in low- and middle-income countries.
News
Publications
Author list abbreviated; Abdullahi, T. shown in bold. Full list on Google Scholar.