เฅฅ เคถเฅเคฐเฅ เคเคฃเฅเคถเคพเคฏ เคจเคฎเค เฅฅ
Processing a 1000-page tantric manuscript with crash-safe resume capability
|
Ancient Sanskrit and Hindi manuscriptsโtantras, stotras, and sacred textsโare being lost to time. Existing OCR tools:
|
OCR-Devnagari combines local OCR speed with Gemini AI accuracy: |
| Feature | Description | |
|---|---|---|
| ๐ | Multi-Engine Support | 5 OCR backends to choose from |
| ๐ง | Smart Hybrid Mode | EasyOCR + Gemini for optimal results |
| ๐๏ธ | Mantra Detection | Auto-detect and preserve sacred text |
| โก | High Performance | Async concurrent workers |
| ๐พ | Crash-Safe | Resume from any interruption |
| ๐ | Live Progress | Real-time tracking with ETA |
| ๐ก๏ธ | Graceful Shutdown | Ctrl+C saves all work |
| ๐งน | Memory Efficient | Handles 1000+ page PDFs |
| โ | Response Validation | Rejects invalid OCR results |
# Clone the repository
git clone https://github.com/rajeshkanaka/OCR-Devnagari.git
cd OCR-Devnagari
# Install with UV (recommended)
uv sync && uv pip install easyocr
# Or with pip
pip install -r requirements.txt && pip install easyocr# Option A: Vertex AI (Recommended for production)
export GOOGLE_CLOUD_PROJECT="your-project"
export GOOGLE_CLOUD_LOCATION="global"
export GOOGLE_GENAI_USE_VERTEXAI=1
# Option B: API Key (Quick setup)
export GEMINI_API_KEY="your-key"# ๐ฅ Hybrid mode โ 90% savings, maximum accuracy
python -m ocr_hindi ocr manuscript.pdf --pages "all"
# ๐ 100% FREE local processing
python -m ocr_hindi ocr manuscript.pdf -e easyocr
# ๐ Premium Gemini mode for critical documents
python -m ocr_hindi ocr manuscript.pdf -e gemini
|
|
| Engine | Cost | Accuracy | Speed | Best For |
|---|---|---|---|---|
| ๐ hybrid | ~$0.30/1K | โญโญโญโญโญ | โกโกโก | Recommended |
| ๐ easyocr | FREE | โญโญโญโญ | โกโก | Budget-conscious |
| ๐ marker | FREE | โญโญโญโญโญ | โกโกโก | Structured PDFs |
| ๐ tesseract | FREE | โญโญโญ | โกโกโกโก | Simple documents |
| ๐ gemini | ~$2/1K | โญโญโญโญโญ | โกโกโกโก | Critical accuracy |
"Write once, crash anywhere, resume everywhere"
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ ๐ PDF Input โ
โโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ ๐ INTELLIGENT ROUTING โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ โ
โ โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโ โ
โ โ hybrid โ โ easyocr โ โ marker โ โtesseract โ โ gemini โ โ
โ โ DEFAULT โ โ FREE โ โ FREE โ โ FREE โ โ PREMIUM โ โ
โ โโโโโโฌโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโ โ
โ โ โ
โ โผ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ ๐ง HYBRID DECISION ENGINE โ โ
โ โ โ โ
โ โ โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโ โ โ
โ โ โ EasyOCR โ โโโโโโโถ โ Confidence Checkโ โโโโโโโถ โ Mantra โ โ โ
โ โ โ FREE โ โ < 85% ? โ โ Detected? โ โ โ
โ โ โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโ โโโโโโโโฌโโโโโโโ โ โ
โ โ โ โ โ โ
โ โ โผ โผ โ โ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ โ
โ โ โ ๐ Gemini 2.0 Flash โ โ โ
โ โ โ โข thinking_level: "low" โ โ โ
โ โ โ โข media_resolution: "high" โ โ โ
โ โ โ โข Token tracking for cost โ โ โ
โ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ ๐ก๏ธ CRASH-SAFE PIPELINE โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ โ
โ โโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โ
โ โ OCR โโโโโโถโ Cache โโโโโโถโ Progress โโโโโโถโ Release โ โ
โ โ Process โ โ Atomic Write โ โ Update โ โ Memory โ โ
โ โโโโโโโโโโโโ โ page_NNN.txt โ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโ โ
โ โโโโโโโโโโโโโโโโ โ
โ โ
โ On interrupt (Ctrl+C) or crash: โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ โ All cached pages preserved โ Resume skips completed pages โ โ
โ โ โ No duplicate API charges โ Output merged from cache โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ ๐ Markdown Output + ๐ฐ Cost Report โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Intelligent detection of sacred text patterns ensures mantras are always verified with maximum accuracy
|
เคฌเฅเค เคฎเคจเฅเคคเฅเคฐ |
เคฎเคจเฅเคคเฅเคฐ เคธเคฎเคพเคชเฅเคคเคฟ |
เคถเฅเคฒเฅเค เคเคฟเคนเฅเคจ |
เคตเคฟเคญเคพเค เคธเฅเคเค |
# Process entire manuscript with intelligent routing
python -m ocr_hindi ocr sacred_text.pdf --pages "all"
# Adjust confidence threshold (higher = more Gemini verification)
python -m ocr_hindi ocr sacred_text.pdf --confidence 0.90
# Disable mantra verification for faster processing
python -m ocr_hindi ocr sacred_text.pdf --no-verify-mantras
# Process specific page ranges
python -m ocr_hindi ocr sacred_text.pdf --pages "1-100,200-250"
# Use more workers for faster processing
python -m ocr_hindi ocr sacred_text.pdf --workers 10# EasyOCR โ Good Hindi/Devanagari support, no API needed
python -m ocr_hindi ocr book.pdf -e easyocr
# Marker โ Best for structured books and PDFs
python -m ocr_hindi ocr book.pdf -e marker
# Tesseract โ Fast, requires system installation
python -m ocr_hindi ocr book.pdf -e tesseract# Maximum accuracy for critical manuscripts
python -m ocr_hindi ocr rare_manuscript.pdf -e gemini
# With high concurrency
python -m ocr_hindi ocr rare_manuscript.pdf -e gemini --workers 15# List all available engines with details
python -m ocr_hindi engines
# Validate your setup (dependencies + authentication)
python -m ocr_hindi validate
# View PDF information
python -m ocr_hindi info manuscript.pdf
# Dry run โ see what would be processed
python -m ocr_hindi ocr manuscript.pdf --dry-run
# Resume interrupted processing
python -m ocr_hindi ocr manuscript.pdf --resume| Option | Description | Default |
|---|---|---|
-e, --engine |
OCR engine (hybrid, easyocr, marker, tesseract, gemini) |
hybrid |
-p, --pages |
Page range (all, 1-50, 1,5,10-20) |
interactive |
-w, --workers |
Concurrent workers (1-20) | 5 |
-c, --confidence |
Hybrid threshold (0.0-1.0) | 0.85 |
--verify-mantras |
Verify mantra pages with Gemini | true |
-r, --resume |
Resume from previous progress | false |
-n, --dry-run |
Preview without processing | false |
--dpi |
PDF rendering quality | 200 |
your_manuscript/
โโโ ๐ manuscript.pdf # Original file
โโโ ๐ manuscript_unicode.md # โจ Final output (Devanagari text)
โโโ ๐ ocr_manuscript_20240120_143022.log # Processing log
โโโ ๐ .ocr_progress_manuscript.json # Resume state
โโโ ๐ .ocr_cache_manuscript/ # ๐ก๏ธ Crash-safe cache
โโโ page_0001.txt # Individual page cache
โโโ page_0001.meta.json # Page metadata
โโโ page_0002.txt
โโโ ...
| Mode | 1000 Pages | Throughput | Cost | Notes |
|---|---|---|---|---|
| ๐ Hybrid | ~90 min | ~11 ppm | ~$1 | Best value |
| ๐ EasyOCR | ~120 min | ~8 ppm | $0 | 100% free |
| ๐ Marker | ~60 min | ~16 ppm | $0 | Structured PDFs |
| ๐ Gemini | ~45 min | ~22 ppm | ~$10 | Max accuracy |
ppm = pages per minute โข Tested on M1 MacBook Pro with 10 workers
โ "poppler not found"
# macOS
brew install poppler
# Ubuntu/Debian
sudo apt-get install poppler-utils
# Windows - Download from:
# https://github.com/oschwartz10612/poppler-windows/releasesโ "EasyOCR not installed"
uv pip install easyocr
# or
pip install easyocrโ "Tesseract not installed"
# macOS
brew install tesseract tesseract-lang
# Ubuntu/Debian
sudo apt install tesseract-ocr tesseract-ocr-hin tesseract-ocr-san
# Windows - Download installer from:
# https://github.com/UB-Mannheim/tesseract/wikiโ Authentication errors
# Verify Vertex AI setup
gcloud auth application-default login
gcloud config set project YOUR_PROJECT_ID
# Or use API key instead
export GEMINI_API_KEY="your-api-key-here"
# Test authentication
python -m ocr_hindi validateโ Rate limiting (429 errors)
# Reduce concurrent workers
python -m ocr_hindi ocr book.pdf --workers 3
# The system will automatically retry with exponential backoffโ High memory usage
# Reduce workers (each worker holds images in memory)
python -m ocr_hindi ocr book.pdf --workers 2
# Or process in smaller batches
python -m ocr_hindi ocr book.pdf --pages "1-100"
python -m ocr_hindi ocr book.pdf --pages "101-200" --resumeContributions are what make the open source community amazing!
|
๐ Bug Reports |
๐ก Feature Ideas |
๐ง Pull Requests |
๐ Documentation |
# Fork, clone, and create a branch
git clone https://github.com/YOUR_USERNAME/OCR-Devnagari.git
cd OCR-Devnagari
git checkout -b feature/amazing-feature
# Make your changes, then
git commit -m "Add amazing feature"
git push origin feature/amazing-feature
# Open a Pull Request ๐MIT License โ Free for personal and commercial use
See LICENSE for details