An open benchmark red-teaming frontier LLMs for safety-guardrail failures in Bengali and other low-resource South Asian languages, with responsible disclosure to labs.
An open benchmark red-teaming frontier LLMs for safety-guardrail failures in Bengali and other low-resource South Asian languages, with responsible disclosure to labs.
Project Details
Updated 07/13/26 · Provided via application · VerifiedFrontier model safety training such as refusals, dangerous-capability guardrails, jailbreak resistance is developed and evaluated overwhelmingly in English. There is growing evidence that these safeguards degrade sharply in low-resource languages, and that translating a harmful request into a low-resource language is itself a working jailbreak vector. As frontier models are deployed to hundreds of millions of Bengali speakers (and across South Asia), this creates a real, largely unmeasured misuse surface.
I will build the first systematic safety-evaluation benchmark for Bengali, covering (1) refusal robustness on harmful requests, (2) low-resource-language jailbreak attacks, and (3) dangerous-capability elicitation compared against English baselines. I'll evaluate leading models (Claude, GPT, Gemini, Llama), quantify the English-vs-Bengali safety gap, and responsibly disclose exploitable failures to the relevant labs before any public release. Outputs: an open-source benchmark and dataset (HuggingFace), reproducible evaluation code, and a public technical report.
I'm a Bangladeshi ML researcher (Brac University) with published transformer-based Bengali NLP work and 360+ citations across deep-learning and evaluation research. As a native Bengali speaker embedded in this research community, I can construct linguistically valid adversarial data that non-speakers cannot — the core reason this gap remains unmeasured. If successful, I'll extend to Urdu, Nepali, and Sinhala as a reusable multilingual safety-eval framework.
Theory of Impact
Updated 07/21/26 · By grantmaking.aiRobust, evaluated safety behavior is a prerequisite for safely deploying increasingly capable models. If guardrails fail predictably in non-English languages, then (a) misuse of dangerous capabilities is easier via a simple translation attack, and (b) our confidence in "the model is safe" is systematically overstated, because safety is only measured in English. This project directly reduces both risks: it produces a concrete measurement of the multilingual safety gap, gives labs an actionable signal to close it (via responsible disclosure), and creates reusable infrastructure the safety community can extend to other languages. Evaluations are a recognized bottleneck in AI safety; low-resource-language coverage is an especially neglected corner of it, and neglectedness is exactly where a small grant is most counterfactual.
People
Updated 07/21/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.