Building an open evaluation suite to identify multilingual safety, control, and jailbreak failures in frontier AI systems across African languages.
Building an open evaluation suite to identify multilingual safety, control, and jailbreak failures in frontier AI systems across African languages.
Project Details
Updated 07/04/26 · Provided via application · VerifiedThis project tests whether advanced AI systems is as safe in African languages as they are in English.
Most AI safety evaluations today are built solely in English, which creates a major blind spot. A model sometimes refuse a dangerous request in English, but respond differently when the same request is written in Yoruba, Swahili, Nigerian Pidgin, Hausa, or Igbo. Safety systems may also become weaker when users switch between languages, use regional dialects, or rely on translation tools.
We plan to build an open evaluation suite that measures these failures across several African languages. The project will test areas like jailbreak resistance, harmful-request refusal, instruction following, deception, translation-based attacks, and whether models behave consistently across languages.
The goal is to give AI developers, researchers, and policymakers better tools to understand where multilingual safety protections break down before more powerful AI systems are deployed globaly. This is still an underexplored area, and we think it is important to test these systems in the languages people actually use, not just the ones most represented in current AI research.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiAdvanced AI systems are increasingly being deployed across countries, languages, and public institutions, but most safety testing is still done mainly in English. This creates a serious blind spot. A model may follow safety rules in English, but become easier to jailbreak, manipulate, or misuse when the same request is made in Yoruba, Swahili, Nigerian Pidgin, Hausa, or Igbo.
Our theory of impact is that better multilingual evaluations can reveal safety failures before more capable AI systems are deployed at scale. We will test whether models remain consistent across languages in areas such as harmful-request refusal, deception, instruction following, agent control, and translation-based attacks.
The results can help AI labs improve safeguards, help researchers design stronger evaluation methods, and help governments understand the risks of deploying models that have not been properly tested outside English. Through NKENNEAi’s partnership with NITDA, we also have a pathway to share these findings with policymakers and technical stakeholders in Nigeria, where AI adoption is growing quickly.
This project will not eliminate AI x-risk on its own. Its value is in closing an overlooked gap in global AI safety. If advanced AI systems can behave unsafely in languages that current evaluations barely cover, then developers and governments may believe those systems are more controllable than they really are. Identifying those failures earlier can support safer deployment, stronger oversight, and better international coordination around advanced AI.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.