Skip to content

About

Build local voice agents with open-source models

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Repository files navigation

VoxRuntime STS

Conversational Voice Agent powered by CSM (Conversational Speech Model)

License Python

A low-latency voice agent pipeline with acoustic conversation context: VAD → STT → LLM → CSM (TTS). Built on top of huggingface/speech-to-speech, this fork replaces the standard TTS with Sesame AI's CSM-1B — a conversational speech model that generates voice conditioned on the acoustic history of the conversation.

graph LR
    MIC["🎤 Mic"] --> VAD["VAD<br/>Silero v5"]
    VAD -->|"audio segments"| STT["STT<br/>Faster-Whisper"]
    STT -->|"transcription"| LLM["LLM<br/>DeepSeek V4 Pro"]
    LLM -->|"response text"| CSM["TTS<br/>CSM-1B"]
    CSM -->|"24kHz audio"| SPK["🔊 Speaker"]

    CSM -.->|"acoustic context<br/>(user audio + transcription)"| CSM
Loading

Why This Project?

Most voice assistants treat each utterance in isolation — same voice, same intonation, zero memory of how you just spoke. CSM-1B changes that: it conditions each response on the acoustic history of the conversation. If you whispered — it whispers back. If you laughed — its prosody shifts.

This is a portfolio project demonstrating:

  • Multi-component real-time AI pipeline design
  • GPU/CPU resource partitioning (CPU for VAD+STT, GPU for TTS, external API for LLM)
  • Integration of cutting-edge open models (CSM-1B, Faster-Whisper)
  • Modular handler architecture — swap any component without touching the rest

Pipeline

graph TD
    subgraph "CPU + RAM (62 GB)"
        MIC2["Microphone Input<br/>16kHz mono PCM"]
        VAD2["Voice Activity Detection<br/>Silero VAD v5<br/>~50 MB RAM"]
        STT2["Speech-to-Text<br/>Faster-Whisper large-v3<br/>~3 GB RAM (int8)"]
        HIST["Conversation History<br/>SQLite / JSON<br/>~negligible"]
    end

    subgraph "GPU (RTX 4090 24 GB)"
        CSM2["Text-to-Speech<br/>CSM-1B<br/>~6 GB VRAM (bfloat16)"]
    end

    subgraph "External"
        LLM2["Language Model<br/>DeepSeek V4 Pro<br/>via local proxy :8088"]
    end

    MIC2 --> VAD2
    VAD2 -->|"speech segments"| STT2
    STT2 -->|"text transcription"| HIST
    HIST -->|"conversation context"| LLM2
    LLM2 -->|"response text"| CSM2
    CSM2 -->|"24kHz → 16kHz resampled"| SPK2["Speaker Output"]

    VAD2 -.->|"audio (for acoustic context)"| CSM2
    STT2 -.->|"text (for acoustic context)"| CSM2
Loading

Hardware Allocation Strategy

Resource Component Why Here?
CPU + RAM VAD Silero VAD is lightweight (~50 MB). CPU inference is real-time.
CPU + RAM STT Faster-Whisper large-v3 on int8 runs comfortably on CPU.
CPU + RAM Conversation History SQLite or JSON — negligible memory footprint.
GPU (RTX 4090) CSM-1B Autoregressive transformer with KV-cache. CPU would be 10-30× slower than real-time. GPU required.
External DeepSeek LLM 131K context, state-of-the-art instruction following. Runs via local proxy on 127.0.0.1:8088.

Getting Started

Prerequisites

# Clone the fork
git clone https://github.com/c0gni/vox-runtime-sts.git
cd sts-pc-assistant

# Activate the virtual environment
uv sync
source .venv/bin/activate

Environment

export DEEPSEEK_API_KEY="your-key"  # for LLM via local proxy

Quick Launch

# Terminal 1: Start the voice agent server
s2s-server

# Terminal 2: Start the microphone client
s2s-client

Or manually:

speech-to-speech \
    --mode realtime \
    --stt faster-whisper \
    --faster_whisper_stt_model_name /lang_models/whisper-large-v3 \
    --faster_whisper_stt_device cuda \
    --faster_whisper_stt_gen_language pl \
    --llm_backend chat-completions \
    --model_name deepseek-v4-pro \
    --responses_api_base_url http://127.0.0.1:8088/v1 \
    --responses_api_api_key "$DEEPSEEK_API_KEY" \
    --tts qwen3

Architecture

handler threads (one per component, connected by queues):

  recv_audio ──▶ VAD ──▶ STT ──▶ TranscriptionNotifier
                                     │
                                     ▼
                                  LLM ──▶ LMOutputProcessor
                                               │
                                               ▼
                                            CSM ──▶ send_audio

  text_output_queue (side channel for client events):
    SpeechStarted → PartialTranscription → TranscriptionCompleted
    → AssistantText → TokenUsage → SpeechStopped
stateDiagram-v2
    [*] --> Listening
    Listening --> DetectingSpeech : VAD triggered
    DetectingSpeech --> Transcribing : speech ended
    Transcribing --> Generating : transcription done
    Generating --> Speaking : CSM streaming
    Speaking --> Listening : response done
    DetectingSpeech --> Listening : speech too short
    Generating --> DetectingSpeech : user interrupted
    Speaking --> DetectingSpeech : user interrupted
Loading

Roadmap

  • Core pipeline (VAD → STT → LLM → TTS with Qwen3)
  • CSM-1B TTS handler — replace Qwen3 with conversational speech model
  • Acoustic context — pipe user audio through to CSM for prosody conditioning
  • Conversation history — persistent, local, resumable
  • Wake word — "Hey Computer" activation via Porcupine/OpenWakeWord
  • Agent backend — connect to Claude Code / custom agent for tool use during conversation

Fork Baseline

This is a fork of huggingface/speech-to-speech — a production-grade voice agent pipeline used in thousands of Reachy Mini robots.

Upstream changes preserved: This fork tracks upstream/main. All existing features (all STT/LLM/TTS backends, all run modes, OpenAI Realtime API compatibility) remain functional.

License

Apache 2.0 — same as upstream.

Acknowledgments

About

Build local voice agents with open-source models

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages