Building a Text Collection for Urdu Information Retrieval
Citation
RASHEED, Imran, Haider BANKA & Hamaid M. KHAN. "Building a Text Collection for Urdu Information Retrieval". Etri Journal, 43.5 (2021): 856-868.Abstract
Urdu is a widely spoken language in the Indian subcontinent with over 300 million
speakers worldwide. However, linguistic advancements in Urdu are rare compared to
those in other European and Asian languages. Therefore, by following Text Retrieval
Conference standards, we attempted to construct an extensive text collection of
85 304 documents from diverse categories covering over 52 topics with relevance
judgment sets at 100 pool depth. We also present several applications to demonstrate
the effectiveness of our collection. Although this collection is primarily intended
for text retrieval, it can also be used for named entity recognition, text summarization,
and other linguistic applications with suitable modifications. Ours is the most
extensive existing collection for the Urdu language, and it will be freely available for
future research and academic education.