Conversational Voice Agent powered by CSM (Conversational Speech Model)
A low-latency voice agent pipeline with acoustic conversation context: VAD → STT → LLM → CSM (TTS). Built on top of huggingface/speech-to-speech, this fork replaces the standard TTS with Sesame AI's CSM-1B — a conversational speech model that generates voice conditioned on the acoustic history of the conversation.
graph LR
MIC["🎤 Mic"] --> VAD["VAD<br/>Silero v5"]
VAD -->|"audio segments"| STT["STT<br/>Faster-Whisper"]
STT -->|"transcription"| LLM["LLM<br/>DeepSeek V4 Pro"]
LLM -->|"response text"| CSM["TTS<br/>CSM-1B"]
CSM -->|"24kHz audio"| SPK["🔊 Speaker"]
CSM -.->|"acoustic context<br/>(user audio + transcription)"| CSM
Most voice assistants treat each utterance in isolation — same voice, same intonation, zero memory of how you just spoke. CSM-1B changes that: it conditions each response on the acoustic history of the conversation. If you whispered — it whispers back. If you laughed — its prosody shifts.
This is a portfolio project demonstrating:
- Multi-component real-time AI pipeline design
- GPU/CPU resource partitioning (CPU for VAD+STT, GPU for TTS, external API for LLM)
- Integration of cutting-edge open models (CSM-1B, Faster-Whisper)
- Modular handler architecture — swap any component without touching the rest
graph TD
subgraph "CPU + RAM (62 GB)"
MIC2["Microphone Input<br/>16kHz mono PCM"]
VAD2["Voice Activity Detection<br/>Silero VAD v5<br/>~50 MB RAM"]
STT2["Speech-to-Text<br/>Faster-Whisper large-v3<br/>~3 GB RAM (int8)"]
HIST["Conversation History<br/>SQLite / JSON<br/>~negligible"]
end
subgraph "GPU (RTX 4090 24 GB)"
CSM2["Text-to-Speech<br/>CSM-1B<br/>~6 GB VRAM (bfloat16)"]
end
subgraph "External"
LLM2["Language Model<br/>DeepSeek V4 Pro<br/>via local proxy :8088"]
end
MIC2 --> VAD2
VAD2 -->|"speech segments"| STT2
STT2 -->|"text transcription"| HIST
HIST -->|"conversation context"| LLM2
LLM2 -->|"response text"| CSM2
CSM2 -->|"24kHz → 16kHz resampled"| SPK2["Speaker Output"]
VAD2 -.->|"audio (for acoustic context)"| CSM2
STT2 -.->|"text (for acoustic context)"| CSM2
| Resource | Component | Why Here? |
|---|---|---|
| CPU + RAM | VAD | Silero VAD is lightweight (~50 MB). CPU inference is real-time. |
| CPU + RAM | STT | Faster-Whisper large-v3 on int8 runs comfortably on CPU. |
| CPU + RAM | Conversation History | SQLite or JSON — negligible memory footprint. |
| GPU (RTX 4090) | CSM-1B | Autoregressive transformer with KV-cache. CPU would be 10-30× slower than real-time. GPU required. |
| External | DeepSeek LLM | 131K context, state-of-the-art instruction following. Runs via local proxy on 127.0.0.1:8088. |
# Clone the fork
git clone https://github.com/c0gni/vox-runtime-sts.git
cd sts-pc-assistant
# Activate the virtual environment
uv sync
source .venv/bin/activateexport DEEPSEEK_API_KEY="your-key" # for LLM via local proxy# Terminal 1: Start the voice agent server
s2s-server
# Terminal 2: Start the microphone client
s2s-clientOr manually:
speech-to-speech \
--mode realtime \
--stt faster-whisper \
--faster_whisper_stt_model_name /lang_models/whisper-large-v3 \
--faster_whisper_stt_device cuda \
--faster_whisper_stt_gen_language pl \
--llm_backend chat-completions \
--model_name deepseek-v4-pro \
--responses_api_base_url http://127.0.0.1:8088/v1 \
--responses_api_api_key "$DEEPSEEK_API_KEY" \
--tts qwen3handler threads (one per component, connected by queues):
recv_audio ──▶ VAD ──▶ STT ──▶ TranscriptionNotifier
│
▼
LLM ──▶ LMOutputProcessor
│
▼
CSM ──▶ send_audio
text_output_queue (side channel for client events):
SpeechStarted → PartialTranscription → TranscriptionCompleted
→ AssistantText → TokenUsage → SpeechStopped
stateDiagram-v2
[*] --> Listening
Listening --> DetectingSpeech : VAD triggered
DetectingSpeech --> Transcribing : speech ended
Transcribing --> Generating : transcription done
Generating --> Speaking : CSM streaming
Speaking --> Listening : response done
DetectingSpeech --> Listening : speech too short
Generating --> DetectingSpeech : user interrupted
Speaking --> DetectingSpeech : user interrupted
- Core pipeline (VAD → STT → LLM → TTS with Qwen3)
- CSM-1B TTS handler — replace Qwen3 with conversational speech model
- Acoustic context — pipe user audio through to CSM for prosody conditioning
- Conversation history — persistent, local, resumable
- Wake word — "Hey Computer" activation via Porcupine/OpenWakeWord
- Agent backend — connect to Claude Code / custom agent for tool use during conversation
This is a fork of huggingface/speech-to-speech — a production-grade voice agent pipeline used in thousands of Reachy Mini robots.
Upstream changes preserved: This fork tracks upstream/main. All existing features (all STT/LLM/TTS backends, all run modes, OpenAI Realtime API compatibility) remain functional.
Apache 2.0 — same as upstream.
- Sesame AI Labs for CSM-1B
- Hugging Face for the speech-to-speech pipeline
- SYSTRAN for Faster-Whisper
- DeepSeek for the LLM API