A Part-of-Speech (POS) tagger for the Hindi language built with NLTK. It trains a sequence of n-gram taggers with back-off on the Universal Dependencies Hindi-HDTB treebank and evaluates tagging accuracy on a held-out test split.
POS tagging assigns a grammatical category (noun, verb, adjective, …) to each word in a sentence. It is a foundational step for downstream NLP tasks such as parsing, machine translation, and named-entity recognition.
This project implements POS tagging two ways:
| Script | Approach | Data |
|---|---|---|
NLP_Project.py |
DefaultTagger → Unigram → Bigram → Trigram taggers chained with back-off | UD Hindi-HDTB (.conllu) |
importnltk.py |
TnT (Trigrams'n'Tags) statistical tagger | NLTK indian corpus (hindi.pos) |
The primary, up-to-date implementation is NLP_Project.py.
- Parse the CoNLL-U dataset files, extracting
(word, UPOS)pairs per sentence. - Build a back-off chain so rarer n-grams fall back to simpler models:
TrigramTagger → BigramTagger → UnigramTagger → DefaultTagger('NN'). - Evaluate the trigram tagger's accuracy on the test split.
- Tokenize a sample Hindi sentence with
indic-nlp-libraryand print its tags.
The tagger reads the UD Hindi-HDTB treebank splits in CoNLL-U format (not committed to this repo; see Data for the download):
hi_hdtb-ud-train.conllu: training sentenceshi_hdtb-ud-test.conllu: test sentences
Each token line is tab-separated; column 2 is the word form and column 4 is the universal POS tag.
# Python 3.10+ (nltk 3.10 requires it)
pip install -r requirements.txtThe script downloads the required NLTK data packages (punkt, punkt_tab,
indian, nonbreaking_prefixes) automatically on first run.
python NLP_Project.pyExample output:
Model Accuracy: 88.xx%
Tagged Sentence:
दो/NUM आदमी/NOUN आए/VERB ।/PUNCT
(The exact accuracy depends on the installed NLTK version and data.)
To try the alternative TnT-based tagger:
python importnltk.py.
├── NLP_Project.py # Main n-gram back-off tagger (UD Hindi-HDTB)
├── importnltk.py # Alternative TnT tagger (NLTK indian corpus)
├── hi_hdtb-ud-train.conllu # Training data (download, not committed)
├── hi_hdtb-ud-test.conllu # Test data (download, not committed)
├── Hindipos.pdf # Project report
└── requirements.txt
- Python 3.8+
- NLTK 3.8.1
- indic-nlp-library
The Hindi UD treebank files (hi_hdtb-ud-train.conllu, hi_hdtb-ud-test.conllu, ~52 MB) are not committed. Download the UD Hindi-HDTB treebank from https://universaldependencies.org/ (or https://github.com/UniversalDependencies/UD_Hindi-HDTB) and place the .conllu files in the repo root to run the tagger.