Skip to content

About

N-gram back-off POS tagger for Hindi (NLTK) trained on the UD Hindi-HDTB treebank, ~88% accuracy.

Topics

Resources

Stars

8 stars

Watchers

0 watching

Forks

Latest commit

 

History

17 Commits

Folders and files

Repository files navigation

POS Tagger for Hindi (NLTK)

A Part-of-Speech (POS) tagger for the Hindi language built with NLTK. It trains a sequence of n-gram taggers with back-off on the Universal Dependencies Hindi-HDTB treebank and evaluates tagging accuracy on a held-out test split.

Overview

POS tagging assigns a grammatical category (noun, verb, adjective, …) to each word in a sentence. It is a foundational step for downstream NLP tasks such as parsing, machine translation, and named-entity recognition.

This project implements POS tagging two ways:

Script Approach Data
NLP_Project.py DefaultTagger → Unigram → Bigram → Trigram taggers chained with back-off UD Hindi-HDTB (.conllu)
importnltk.py TnT (Trigrams'n'Tags) statistical tagger NLTK indian corpus (hindi.pos)

The primary, up-to-date implementation is NLP_Project.py.

How it works (NLP_Project.py)

  1. Parse the CoNLL-U dataset files, extracting (word, UPOS) pairs per sentence.
  2. Build a back-off chain so rarer n-grams fall back to simpler models: TrigramTagger → BigramTagger → UnigramTagger → DefaultTagger('NN').
  3. Evaluate the trigram tagger's accuracy on the test split.
  4. Tokenize a sample Hindi sentence with indic-nlp-library and print its tags.

Dataset

The tagger reads the UD Hindi-HDTB treebank splits in CoNLL-U format (not committed to this repo; see Data for the download):

  • hi_hdtb-ud-train.conllu: training sentences
  • hi_hdtb-ud-test.conllu: test sentences

Each token line is tab-separated; column 2 is the word form and column 4 is the universal POS tag.

Setup

# Python 3.10+ (nltk 3.10 requires it)
pip install -r requirements.txt

The script downloads the required NLTK data packages (punkt, punkt_tab, indian, nonbreaking_prefixes) automatically on first run.

Usage

python NLP_Project.py

Example output:

Model Accuracy: 88.xx%

Tagged Sentence:
दो/NUM आदमी/NOUN आए/VERB ।/PUNCT

(The exact accuracy depends on the installed NLTK version and data.)

To try the alternative TnT-based tagger:

python importnltk.py

Project structure

.
├── NLP_Project.py            # Main n-gram back-off tagger (UD Hindi-HDTB)
├── importnltk.py             # Alternative TnT tagger (NLTK indian corpus)
├── hi_hdtb-ud-train.conllu   # Training data (download, not committed)
├── hi_hdtb-ud-test.conllu    # Test data (download, not committed)
├── Hindipos.pdf              # Project report
└── requirements.txt

Tech stack

  • Python 3.8+
  • NLTK 3.8.1
  • indic-nlp-library

Data

The Hindi UD treebank files (hi_hdtb-ud-train.conllu, hi_hdtb-ud-test.conllu, ~52 MB) are not committed. Download the UD Hindi-HDTB treebank from https://universaldependencies.org/ (or https://github.com/UniversalDependencies/UD_Hindi-HDTB) and place the .conllu files in the repo root to run the tagger.

About

N-gram back-off POS tagger for Hindi (NLTK) trained on the UD Hindi-HDTB treebank, ~88% accuracy.

Topics

Resources

Stars

8 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages