Skip to Content

Michael Owusu

Michael Owusu

by Corban Swain

Morehouse College
Faculty Advisor: Prof. Marzyeh Ghassemi
Research Supervisors: Kumail Hamoud, Hara Moraitaki
Department: Electrical Engineering and Computer Science

Biography

Michael Owusu is a Computer Science major at Morehouse College from Kumasi, Ghana. He is
interested in artificial intelligence and machine learning. This summer, he is conducting research
at MIT under the mentorship of Dr. Marzyeh Ghassemi, where he studies multilingual safety in large
language models. He enjoys building software, attending research conferences, and learning about
the problems other researchers are working on. When he gets the chance, he also enjoys sharing his
own work.

Negation Understanding in Vision-Language Models: A Twi (Akan) Extension of NegBench
Michael Owusu¹, Kumail Alhamoud², Hara Moraitaki² and Marzyeh Ghassemi²

¹Department of Computer Science, Morehouse College
²Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology

Vision-language models power image search, but they ignore negation: asked for “a photo with no
cars,” they return cars. Such errors reverse meaning, which matters where absence is the message,
e.g., a scan showing “no evidence of pneumonia.” NegBench (Alhamoud et al., 2025) documented this
in English and showed that finetuning on synthetic negated captions recovers 28 points; whether the
failure or the fix holds elsewhere was untested. Within an ongoing HealthyML project extending
NegBench across languages, we built the first negation benchmark for Twi, an unevaluated Ghanaian
language: 5,914 human-verified multiple-choice questions and 500 captions translated by a native
speaker (first author). Across ten models and five languages, reading a language and understanding
negation prove separate abilities. NLLB-CLIP matches 14.0% of Twi captions to their images, against
61.0% in English, yet answers 1.2% of Twi negation questions correctly, below 25% chance. The fix
does not transfer: one finetuned model drops from 54.5% in English to 17.5% in Twi, below its
untuned baseline. Measuring image-text similarity, the object a caption names matters 8–39x more
than whether it is negated: a matching shortcut, not missing data. Negation must be addressed in
training, and without benchmarks this failure stays invisible.

« Back to profiles