AbjadNLP Workshop Series

Language technology for a script shared across continentsNatural Language Processing for Arabic, Perso-Arabic and Ajami writing traditions

AbjadNLP brings together researchers, language communities and technology builders working on languages written in Arabic-derived scripts. The workshop exists because script matters: it shapes tokenisation, orthography, OCR, search, generation, language resources and the ability of communities to participate fully in the digital world.

Why a workshop around writing systems?

Because language technology is not script-neutral. Orthographic variation, shaping and rendering, segmentation, transliteration, OCR, low-resource data and multilingual modelling all behave differently across writing traditions.

~1Bspeakers represented across the wider language communities
Why AbjadNLP matters

Digital inclusion starts with the languages people actually write

Arabic-derived scripts are used across Africa, the Middle East, Central Asia and South Asia by communities with very different linguistic histories. Some of these languages have mature digital resources; many others remain under-resourced, inconsistently encoded, poorly represented in training data or weakly supported by mainstream NLP systems.

That gap affects much more than benchmark scores. It affects whether people can search, translate, read, write, learn, access public information, build cultural archives and use modern AI systems in their own languages. It also affects whether linguistic heritage survives the transition from print and oral traditions into digital environments.

AbjadNLP provides a place to study these problems together while respecting the differences between Arabic, Perso-Arabic and Ajami traditions. It connects computational work with orthography, language resources, sociolinguistics and the practical needs of language communities.

The aim is not to force diverse languages into one technical mould. It is to build NLP that understands the diversity carried by a shared writing tradition.
01 · Representation

Many languages remain digitally under-resourced

Low-resource communities need corpora, lexicons, benchmarks, tools and models that reflect their actual writing practices.

02 · Writing systems

Script creates technical challenges of its own

Tokenisation, orthography, Unicode, glyph rendering, OCR, transliteration and spelling variation require dedicated attention.

03 · Modern AI

LLMs can reproduce existing resource inequalities

Strong performance in high-resource languages does not guarantee reliable behaviour for dialects, minority languages or alternative orthographies.

04 · Heritage

Language technology can help preserve written traditions

Better digital infrastructure enables communities to search, analyse, teach and preserve linguistic and cultural material.

Language coverage

A workshop defined by writing tradition, not by one language family

AbjadNLP deliberately spans languages that are typologically different but share Arabic-derived writing systems and many related computational challenges.

العربية

Arabic

Modern Standard Arabic, Classical Arabic and regional dialects, including challenges in morphology, dialect variation, code-switching, generation and evaluation.

فارسی · اردو

Perso-Arabic traditions

Persian, Urdu, Pashto, Sorani Kurdish, Sindhi, Uyghur, Azeri and other languages that use adapted Arabic-derived scripts.

عجمي

Ajami traditions

African languages including Hausa, Fula, Wolofal, Swahili, Kanuri, Mandingo and Tamazight written using Arabic-derived conventions.

What the workshop is trying to build

A sustained research community around inclusive Arabic-script NLP

AbjadNLP is designed as an ongoing workshop series rather than a one-off event, connecting research papers, shared tasks, proceedings and community-building across editions.

R

Resources

Encourage reusable corpora, dictionaries, datasets, annotation schemes, orthographic documentation and open tools.

M

Methods

Advance NLP and LLM methods that work across scripts, dialects and low-resource settings rather than only dominant languages.

E

Evaluation

Develop better benchmarks and shared tasks that expose where multilingual systems succeed, fail or behave unevenly.

C

Community

Connect researchers across regions and disciplines, including academia, industry and grassroots language communities.

Workshop series

From Abu Dhabi to Rabat to Athens

Each edition builds a permanent research record through peer-reviewed proceedings while extending the workshop's coverage of Arabic-script languages and shared research challenges.

2025 · 1st edition

AbjadNLP 2025

Abu Dhabi, UAE · 19 January 2025 · COLING 2025

The inaugural workshop established a dedicated international venue for NLP across Abjad, Perso-Arabic and Ajami writing traditions.

2026 · 2nd edition

AbjadNLP 2026

Rabat, Morocco · 28 March 2026 · EACL 2026

The second edition expanded the community and introduced four shared tasks spanning generated-text detection, authorship, style transfer and medical NLP.

2027 · 3rd edition · Current

AbjadNLP 2027

Athens, Greece · EACL 2027

The third edition continues the workshop's multilingual and community-driven agenda, with an expanded shared-task programme and wider Abjad and Ajami coverage.

A growing public research record

AbjadNLP proceedings are archived on the ACL Anthology, giving papers and shared-task work a permanent, openly accessible home and making it possible to follow the development of the field across editions.

Browse all proceedings ↗
Research themes

From foundational language resources to multilingual LLMs

The workshop welcomes work across the NLP pipeline, especially research that helps explain, document or overcome the technical consequences of writing-system and resource diversity.

Core NLP

Morphology, tokenisation, tagging, parsing, named entities, sentiment, language modelling and representation learning.

LLMs & Generative AI

Multilingual and low-resource LLMs, adaptation, evaluation, safety, generated-text detection and culturally aware generation.

Machine Translation & Speech

Translation, speech recognition, speech synthesis and cross-lingual technologies for under-represented languages.

OCR & Writing Systems

Optical character recognition, handwriting, Unicode, fonts, glyph rendering, transliteration and orthographic variation.

Language Resources

Corpora, dictionaries, benchmarks, annotation, orthography descriptions and tools for low-resource communities.

Culture, Society & Access

Code-switching, dialects, language preservation, education, cultural heritage, fairness and digital inclusion.

Workshop series

General Chair

MEDr Mo El-Haj
General Chair · AbjadNLP

Dr Mo El-Haj

Natural Language Processing researcher working on multilingual and under-resourced language technology, with a long-standing focus on Arabic NLP, language resources and cross-lingual research. AbjadNLP is organised as a continuing community workshop with collaborators and chairs changing across editions.