Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
-
Updated
Jun 3, 2026
Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channel integration, agent state management, and observability.
A New End-to-end Framework for Evaluating Voice Agents
🇺🇦 Open Source Ukrainian Text-to-Speech datasets
A Docker-based OpenAI-compatible Text-to-Speech API server powered by Kyutai's TTS models with GPU acceleration support.
Just a simple multimodal avatar interaction platform
A voice-based AI chat interface built with Next.js and ElevenLabs. Start and stop real-time conversations with an animated UI that reflects agent status. Fully responsive and deployable via Vercel with environment-based agent configuration.
Open-source real-time Voice AI infrastructure in Go. Stream audio via WebRTC or WebSocket, connect STT → LLM → TTS pipelines, and build scalable voice agents and conversational AI applications.
A unified benchmarking framework for evaluating Voice AI agents across conversational quality, audio realism, latency metrics, and safety guardrails with scalable multi-language stress testing.
MLX Porting Toolkit — an agent-guided, evidence-gated pipeline (scaffold → convert → parity → benchmark) plus a portable skill for porting PyTorch/Hugging Face models to Apple MLX.
🇺🇦 Ukrainian RAD-TTS++ models (decoder + models with 3 voices) and HiFiGAN model
A voice-based AI chat interface built with Next.js and ElevenLabs. Start and stop real-time conversations with an animated UI that reflects agent status. Fully responsive and deployable via Vercel with environment-based agent configuration.
Legacy Speech AI examples with migration links to the current Brainiall TTS and transcription services.
Code-switching ASR adaptation for strong multilingual speech recognition models. Synthetic CSW data generation, Whisper adaptation, Bayesian LoRA (BLoRA), and robust multilingual ASR evaluation.
MCP Server for Brainiall Speech AI - pronunciation assessment, speech-to-text, and text-to-speech
Open source AI voice calling agent for Twilio phone calls, built with FastAPI, Google ADK, and Gemini Live API
Enterprise-Grade Secure ASR Diarization Pipeline - HIPAA-compliant speech processing service combining automatic speech recognition with speaker diarization. Features modular architecture, comprehensive security, and production-ready deployment.
A source-linked directory of free and trial LLM APIs, multimodal models, embeddings, speech, translation, safety, and other inference endpoints. Companion catalog for freellmapi.io.
Interruptible voice-agent runtime for structured interview prototypes, with VAD-based interruption handling and modular speech backends.
Add a description, image, and links to the speech-ai topic page so that developers can more easily learn about it.
To associate your repository with the speech-ai topic, visit your repo's landing page and select "manage topics."