Skip to content

Latest commit

 

History

7 Commits

Folders and files

Repository files navigation

AI-Transcribe-Pro banner

🎙️ AI-Transcribe-Pro

Turn any audio or video file into accurate text, subtitles, or structured data — fully offline, after a one-time model download.

Python License Offline faster-whisper PRs Welcome

Features • Requirements • Setup • Usage • Models • Deployment • FAQ • Contributing • License


📖 Overview

AI-Transcribe-Pro converts audio and video files into plain text, subtitles (SRT/VTT), or structured JSON, powered by faster-whisper — a highly optimized reimplementation of OpenAI's Whisper built on CTranslate2 and Silero VAD.

It's designed to handle real-world files, not just short clips:

  • 🕐 Works on anything from a few seconds to 10+ hours of audio.
  • ⚡ Long recordings are auto-chunked and transcribed in parallel across your CPU cores — not processed sequentially.
  • 💾 Crash-safe checkpointing — if the process is interrupted, resume instead of starting over.
  • 📴 Model weights are downloaded once; every run after that works completely offline, with no telemetry or internet dependency.

Whether you're transcribing interviews, lectures, podcasts, meetings, or generating subtitles for video content, AI-Transcribe-Pro gives you a self-hosted, private alternative to cloud transcription APIs.


✨ Features

Category Details
📄 Output formats Plain TXT, subtitle formats SRT and VTT, and structured JSON (with timestamps and segment metadata)
🧩 Long-file handling Automatic chunking of long recordings with true parallel processing across CPU cores (not sequential)
🔁 Crash-safe resume Checkpointing means a crashed or interrupted job resumes from where it left off — just re-upload the same file
📥 One-time setup Whisper model weights download once to a local cache; every subsequent run is fully offline
🌐 Multi-language ~28 languages available via dropdown out of the box, but any of Whisper's ~99 supported languages works by editing LANGUAGE_OPTIONS
🇮🇳 Hindi/Urdu disambiguation Hindi and Urdu are acoustically near-identical to the model; the app applies an explainable heuristic that biases detection toward Hindi
🛡️ Anti-hallucination safeguards VAD max-segment cap + repetition penalty tuned to stop Whisper's known "repeated phrase" looping on long, unbroken speech
🖥️ CPU-first, GPU-ready Runs on CPU using int8 quantization by default; automatically uses CUDA/GPU if available for faster processing
🔒 Privacy by design No audio, video, or transcript ever leaves your machine — ideal for sensitive or confidential recordings

🧰 Requirements

  • Python 3.9 – 3.12 (3.13 also works — see the note in Setup)
  • No GPU required — runs on CPU with int8 quantization out of the box
  • Automatically uses a CUDA-capable GPU if one is available, for significantly faster transcription
  • ~1–3 GB free disk space for model weights, depending on the model size chosen
  • Works on Windows, macOS, and Linux

🚀 Setup

# 1. Clone the repository
git clone https://github.com/jastfan/AI-Transcribe-Pro.git
cd AI-Transcribe-Pro

# 2. Create and activate a virtual environment
python -m venv venv

# Windows
venv\Scripts\activate

# macOS / Linux
source venv/bin/activate

# 3. Install dependencies
pip install -r requirements.txt

Python 3.13 users: the app runs fine, but some dependency wheels may need to build from source the first time. If you hit install errors, prefer Python 3.9–3.12 in your virtual environment.


▶️ Usage

Start the app:

python text_converter.py

Then open your browser at:

http://localhost:8084

From there:

  1. Upload an audio or video file (most common formats are supported via ffmpeg).
  2. Choose your target language (or leave on auto-detect).
  3. Pick a model size (see Choosing a model).
  4. Select your desired output format — TXT, SRT, VTT, or JSON.
  5. Run the transcription and download your result once complete.

Networking behavior:

Run Internet required?
First run of any given model size ✅ Yes — downloads model weights once to ./model_cache/
Every run after that ❌ No — fully offline, no network calls at all

🧠 Choosing a model

Whisper model size is a trade-off between speed and accuracy. Pick based on your use case:

Model Speed Accuracy Best for
tiny ⚡ Fastest 🔻 Lowest Quick drafts, short clips, rapid prototyping
base ⚖️ Balanced ✅ Good Default choice — solid for most everyday use-cases
small 🐢 Slower 🎯 Best Important, long, or noisy recordings where accuracy matters most

Larger models generally use more RAM/VRAM and take longer per file, but produce fewer transcription errors — especially with accents, background noise, or technical vocabulary.


🏗️ Deployment notes

Transcription is CPU/GPU-bound work — the machine running the model needs real compute. The web UI/client itself stays lightweight regardless of scale.

Use-case Suggested setup
Personal / local use Your own machine, tiny or base model
Small public web app 4–8 core VPS (e.g. Hetzner, DigitalOcean)
Larger public app / high volume GPU cloud instance (e.g. RunPod, Vast.ai — T4/A10 GPUs)
Offline desktop app (.exe) Package with PyInstaller; runs entirely on the end-user's own CPU

For production deployments:

  • Run behind a proper WSGI server (e.g. gunicorn or waitress) instead of Flask's built-in development server.
  • Put a reverse proxy (e.g. nginx) in front for HTTPS termination and to support larger file uploads.
  • Consider a job queue (e.g. Celery/RQ) if you expect concurrent long-running transcription jobs from multiple users.
  • Monitor disk usage in ./model_cache/ and any temp/chunk directories, especially for very long files.

❓ FAQ

Does this send my audio to any external API? No. Once the model weights are cached locally, transcription runs entirely on your machine with zero network calls.

Can I add more languages to the dropdown? Yes — edit the LANGUAGE_OPTIONS list in text_converter.py. Whisper itself supports ~99 languages; the dropdown only exposes a curated subset (~28) by default.

Why does it sometimes confuse Hindi and Urdu? The two languages are acoustically almost indistinguishable to Whisper's underlying model. AI-Transcribe-Pro applies an explainable heuristic that biases detection toward Hindi when the model is uncertain — you can override the detected language manually if needed.

My long recording has repeated/looping phrases in the output — what's happening? This is a known Whisper failure mode on long, unbroken speech without pauses. The app already applies a VAD max-segment cap and a repetition penalty to mitigate this, but extremely quiet or noisy audio may still trigger it occasionally — try a larger model or manually split the file.

Do I need a GPU? No — it runs fine on CPU with int8 quantization. A CUDA-capable GPU is used automatically if detected, which speeds things up considerably for larger models or longer files.


🤝 Contributing

Contributions, issues, and feature requests are welcome!

  1. Fork the repository
  2. Create a feature branch: git checkout -b feature/your-feature
  3. Commit your changes: git commit -m "Add your feature"
  4. Push to the branch: git push origin feature/your-feature
  5. Open a Pull Request

Please open an issue first for major changes so we can discuss what you'd like to do.


🗺️ Roadmap ideas

  • Speaker diarization (who-said-what)
  • Batch/queue mode for multiple files at once
  • Docker image for one-command deployment
  • REST API mode alongside the web UI

(Have an idea? Open an issue — contributions welcome.)


📜 License

Released under the MIT License — see LICENSE for the full text.


🙏 Acknowledgements

  • faster-whisper — the CTranslate2-based inference engine powering transcription
  • OpenAI Whisper — the original speech recognition model
  • Silero VAD — voice activity detection used for chunking and anti-hallucination

Made with ❤️ for anyone who needs private, offline, accurate transcription.
⭐ If this project helps you, consider starring the repo!

Releases

Packages

Contributors

Languages