Turn any audio or video file into accurate text, subtitles, or structured data — fully offline, after a one-time model download.
Features • Requirements • Setup • Usage • Models • Deployment • FAQ • Contributing • License
AI-Transcribe-Pro converts audio and video files into plain text, subtitles (SRT/VTT), or structured JSON, powered by faster-whisper — a highly optimized reimplementation of OpenAI's Whisper built on CTranslate2 and Silero VAD.
It's designed to handle real-world files, not just short clips:
- 🕐 Works on anything from a few seconds to 10+ hours of audio.
- ⚡ Long recordings are auto-chunked and transcribed in parallel across your CPU cores — not processed sequentially.
- 💾 Crash-safe checkpointing — if the process is interrupted, resume instead of starting over.
- 📴 Model weights are downloaded once; every run after that works completely offline, with no telemetry or internet dependency.
Whether you're transcribing interviews, lectures, podcasts, meetings, or generating subtitles for video content, AI-Transcribe-Pro gives you a self-hosted, private alternative to cloud transcription APIs.
| Category | Details |
|---|---|
| 📄 Output formats | Plain TXT, subtitle formats SRT and VTT, and structured JSON (with timestamps and segment metadata) |
| 🧩 Long-file handling | Automatic chunking of long recordings with true parallel processing across CPU cores (not sequential) |
| 🔁 Crash-safe resume | Checkpointing means a crashed or interrupted job resumes from where it left off — just re-upload the same file |
| 📥 One-time setup | Whisper model weights download once to a local cache; every subsequent run is fully offline |
| 🌐 Multi-language | ~28 languages available via dropdown out of the box, but any of Whisper's ~99 supported languages works by editing LANGUAGE_OPTIONS |
| 🇮🇳 Hindi/Urdu disambiguation | Hindi and Urdu are acoustically near-identical to the model; the app applies an explainable heuristic that biases detection toward Hindi |
| 🛡️ Anti-hallucination safeguards | VAD max-segment cap + repetition penalty tuned to stop Whisper's known "repeated phrase" looping on long, unbroken speech |
| 🖥️ CPU-first, GPU-ready | Runs on CPU using int8 quantization by default; automatically uses CUDA/GPU if available for faster processing |
| 🔒 Privacy by design | No audio, video, or transcript ever leaves your machine — ideal for sensitive or confidential recordings |
- Python 3.9 – 3.12 (3.13 also works — see the note in Setup)
- No GPU required — runs on CPU with int8 quantization out of the box
- Automatically uses a CUDA-capable GPU if one is available, for significantly faster transcription
- ~1–3 GB free disk space for model weights, depending on the model size chosen
- Works on Windows, macOS, and Linux
# 1. Clone the repository
git clone https://github.com/jastfan/AI-Transcribe-Pro.git
cd AI-Transcribe-Pro
# 2. Create and activate a virtual environment
python -m venv venv
# Windows
venv\Scripts\activate
# macOS / Linux
source venv/bin/activate
# 3. Install dependencies
pip install -r requirements.txtPython 3.13 users: the app runs fine, but some dependency wheels may need to build from source the first time. If you hit install errors, prefer Python 3.9–3.12 in your virtual environment.
Start the app:
python text_converter.pyThen open your browser at:
http://localhost:8084
From there:
- Upload an audio or video file (most common formats are supported via
ffmpeg). - Choose your target language (or leave on auto-detect).
- Pick a model size (see Choosing a model).
- Select your desired output format — TXT, SRT, VTT, or JSON.
- Run the transcription and download your result once complete.
Networking behavior:
| Run | Internet required? |
|---|---|
| First run of any given model size | ✅ Yes — downloads model weights once to ./model_cache/ |
| Every run after that | ❌ No — fully offline, no network calls at all |
Whisper model size is a trade-off between speed and accuracy. Pick based on your use case:
| Model | Speed | Accuracy | Best for |
|---|---|---|---|
tiny |
⚡ Fastest | 🔻 Lowest | Quick drafts, short clips, rapid prototyping |
base |
⚖️ Balanced | ✅ Good | Default choice — solid for most everyday use-cases |
small |
🐢 Slower | 🎯 Best | Important, long, or noisy recordings where accuracy matters most |
Larger models generally use more RAM/VRAM and take longer per file, but produce fewer transcription errors — especially with accents, background noise, or technical vocabulary.
Transcription is CPU/GPU-bound work — the machine running the model needs real compute. The web UI/client itself stays lightweight regardless of scale.
| Use-case | Suggested setup |
|---|---|
| Personal / local use | Your own machine, tiny or base model |
| Small public web app | 4–8 core VPS (e.g. Hetzner, DigitalOcean) |
| Larger public app / high volume | GPU cloud instance (e.g. RunPod, Vast.ai — T4/A10 GPUs) |
| Offline desktop app (.exe) | Package with PyInstaller; runs entirely on the end-user's own CPU |
For production deployments:
- Run behind a proper WSGI server (e.g.
gunicornorwaitress) instead of Flask's built-in development server. - Put a reverse proxy (e.g. nginx) in front for HTTPS termination and to support larger file uploads.
- Consider a job queue (e.g. Celery/RQ) if you expect concurrent long-running transcription jobs from multiple users.
- Monitor disk usage in
./model_cache/and any temp/chunk directories, especially for very long files.
Does this send my audio to any external API? No. Once the model weights are cached locally, transcription runs entirely on your machine with zero network calls.
Can I add more languages to the dropdown?
Yes — edit the LANGUAGE_OPTIONS list in text_converter.py. Whisper itself supports ~99 languages; the dropdown only exposes a curated subset (~28) by default.
Why does it sometimes confuse Hindi and Urdu? The two languages are acoustically almost indistinguishable to Whisper's underlying model. AI-Transcribe-Pro applies an explainable heuristic that biases detection toward Hindi when the model is uncertain — you can override the detected language manually if needed.
My long recording has repeated/looping phrases in the output — what's happening? This is a known Whisper failure mode on long, unbroken speech without pauses. The app already applies a VAD max-segment cap and a repetition penalty to mitigate this, but extremely quiet or noisy audio may still trigger it occasionally — try a larger model or manually split the file.
Do I need a GPU? No — it runs fine on CPU with int8 quantization. A CUDA-capable GPU is used automatically if detected, which speeds things up considerably for larger models or longer files.
Contributions, issues, and feature requests are welcome!
- Fork the repository
- Create a feature branch:
git checkout -b feature/your-feature - Commit your changes:
git commit -m "Add your feature" - Push to the branch:
git push origin feature/your-feature - Open a Pull Request
Please open an issue first for major changes so we can discuss what you'd like to do.
- Speaker diarization (who-said-what)
- Batch/queue mode for multiple files at once
- Docker image for one-command deployment
- REST API mode alongside the web UI
(Have an idea? Open an issue — contributions welcome.)
Released under the MIT License — see LICENSE for the full text.
- faster-whisper — the CTranslate2-based inference engine powering transcription
- OpenAI Whisper — the original speech recognition model
- Silero VAD — voice activity detection used for chunking and anti-hallucination
Made with ❤️ for anyone who needs private, offline, accurate transcription.
⭐ If this project helps you, consider starring the repo!
