Skip to main content
LLMgram · AI News · 2026-08-26

Urdu Absent from LLM Safety Evaluations Across Nine Years of WOAH

Urdu Absent from LLM Safety Evaluations Across Nine Years of WOAH

New arXiv research introduces Missed-in-Urdu scores to measure cross-script safety inconsistency, reporting that Urdu, with roughly 246 million speakers and ranked the world's tenth most spoken language, has been almost entirely absent from mainstream LLM safety evaluation and nine years of WOAH proceedings named in the paper. Five large language models are tested; the authors argue Urdu hate speech is systematically missed with concrete moderation consequences. A related benchmark finds no single model catches every harm type, from jailbreaks to implicit hate. Multilingual operators should not trust English-centric safety scores alone. The evidence packet flags the WOAH reference as possibly anomalous and notes truncated abstracts, limiting confidence in methods and full quantitative results.

Sources

Urdu Absent from LLM Safety Evaluations Across Nine Years of WOAH

Urdu Absent from LLM Safety Evaluations Across Nine Years of WOAH

'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection. Urdu, the world's tenth most spoken language with 246 million speakers, remains almost entirely absent from mainstream LLM safety evaluation and nine years of WOAH proceedings.

Key takeaway

Missed-in-Urdu scores make cross-script moderation gaps measurable, but truncated abstracts and a flagged WOAH citation need verification.

What happened

A new cs.CL paper titled 'Ghaib in Translation' introduces Missed-in-Urdu scores to quantify cross-script safety inconsistency in LLM hate speech detection, focusing on Urdu after what it describes as near-total absence from mainstream safety evaluation and nine years of WOAH proceedings.

The arXiv abstract states the study tests five large language models to investigate whether that absence has measurable consequences for content moderation reliability, while a companion paper benchmarks moderation across diverse safety scenarios and reports that no single model catches every type of harmful content.

Evidence

  • Urdu has been almost entirely absent from mainstream LLM safety evaluation and nine years of WOAH proceedings.

    arXiv cs.CL · attributed

    Urdu, the world's tenth most spoken language with 246 million speakers, remains almost entirely absent from mainstream LLM safety evaluation and nine years of WOAH proceedings.

  • The paper introduces Missed-in-Urdu scores to quantify cross-script safety inconsistency.

    arXiv cs.CL · attributed

    A new cs.CL paper introduces 'Missed-in-Urdu' scores to quantify cross-script safety inconsistency, revealing that Urdu—a top-10 global language—has been absent from LLM safety evaluation for nearly a decade.

  • Five large language models were tested for moderation reliability consequences.

    arXiv cs.CL · attributed

    To investigate whether this absence has measurable consequences for content moderation reliability, five large language model

  • Urdu-language hate speech is systematically missed by LLM safety systems.

    arXiv cs.CL · attributed

    Urdu-language hate speech is systematically missed by LLM safety systems, exposing a measurable equity gap in content moderation that affects 246 million speakers.

  • No single model catches every type of harmful content across safety scenarios.

    arXiv cs.CL · attributed

    A new arXiv benchmark evaluates LLM-based content moderation across diverse safety scenarios, revealing that no single model catches every type of harmful content, from explicit jailbreaks to implicit hate speech.

  • Moderation efficacy is highly scenario-dependent across models.

    arXiv cs.CL · attributed

    The findings suggest that moderation efficacy is highly scenario-dependent, with models excelling in some risk categories while missing others.

Why it matters

Teams deploying global LLM products may inherit undetected hate-speech exposure where low-resource languages like Urdu were never benchmarked.

Limits and uncertainties

The packet flags the WOAH reference as anomalous because WOAH is an animal health organization, suggesting possible metadata errors that warrant verification.

Both arXiv abstracts in the packet are truncated, leaving methodology, model sets, and quantitative gaps unspecified.

The companion moderation benchmark critique notes it may underrepresent low-resource languages and non-English harm vectors.

Practical implications

Builders and operators serving multilingual users should audit moderation systems for low-resource language gaps such as Urdu.

Moderation pipelines should be designed as layered or ensemble systems rather than relying on a single model across all harm scenarios.

What to watch

Whether independent verification clarifies the WOAH proceedings citation and full methodology in arXiv:2608.24191.

Release of complete quantitative Missed-in-Urdu results across the five tested large language models.

Whether scenario-level moderation benchmarks expand coverage beyond English-centric harm vectors.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: 'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection