
Sindhi Open Lexicon Dataset (223K+ Entries) for AI, NLP & Computational Linguistics. This project is a large-scale structured lexical dataset for the Sindhi language containing over 223,000 entries including definitions, linguistic metadata, and normalized forms. Sindhi is a historically rich but low-resource language in AI. This dataset aims to support NLP, AI systems, and computational linguistics. Objectives - Provide AI-ready Sindhi dataset - Support NLP research - Enable search engines, chatbots, OCR, and language tools - Preserve linguistic heritage digitally Dataset Features - 223,000+ entries - Definitions in Sindhi - Variants with/without diacritics - Normalized text - Domain classification - Formats: CSV, JSONL, SQLite
sindhi, low-resource-language, nlp, ai, linguistics, lexicon
sindhi, low-resource-language, nlp, ai, linguistics, lexicon
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
