@goodandready/dsh-voice

Voice input for DeepSeek Harness: dictation chunked by pauses and voice messages, each with its own provider fallback chain (Deepgram, Groq, HuggingFace, local whisper.cpp, plus any OpenAI-compatible API of your own).


Keywords
dsh, dsh-plugin, deepseek-harness, voice, stt, dictation, whisper, deepgram, groq, openrouter, openai-compatible
License
MIT
Install
npm install @goodandready/dsh-voice@0.8.18

Documentation

๐Ÿ“ฆ @goodandready/dsh-voice

Zero-Latency Streaming Dictation & Multi-Provider Voice Input for DeepSeek Harness

npm version license DSH Plugin Node version

All Author Projects

๐Ÿ‡ฌ๐Ÿ‡ง English โ€ข ๐Ÿ‡ท๐Ÿ‡บ ะ ัƒััะบะธะน โ€ข ๐Ÿ‡จ๐Ÿ‡ณ ไธญๆ–‡่ฏดๆ˜Ž


โšก Overview

dsh-voice brings voice superpowers to the DeepSeek Harness Web UI. Whether you need hands-free real-time streaming dictation segmented on natural breath pauses or crisp voice notes with keyboard/mouse Push-to-Talk gestures, dsh-voice ensures your audio is never lost thanks to automatic multi-provider fallback chains.

graph LR
    subgraph Client [Browser Web UI]
        Mic[๐ŸŽ™๏ธ Dictation Mic] -->|VAD Cut on Pause| Stream[Audio Chunks]
        Wave[๐ŸŒŠ Voice Message] -->|Hold / Release| PTT[Push-to-Talk]
    end

    subgraph Host [DSH Host Backend]
        Stream --> FFMPEG[ffmpeg 16kHz Transcoder]
        PTT --> FFMPEG
        FFMPEG --> Chain{Fallback Chain}
        
        Chain -->|1st Priority| P1[Deepgram / Nova-2]
        Chain -.->|On Rate Limit / 429| P2[Groq / Whisper Turbo]
        Chain -.->|On Failure| P3[Local whisper.cpp / Offline]
    end

    subgraph Output [Target]
        P1 --> Composer[๐Ÿ’ฌ Web Composer / Chat]
        P2 --> Composer
        P3 --> Composer
    end

    style Client fill:#1e1e2e,stroke:#89b4fa,stroke-width:2px,color:#cdd6f4
    style Host fill:#181825,stroke:#cba6f7,stroke-width:2px,color:#cdd6f4
    style Output fill:#11111b,stroke:#a6e3a1,stroke-width:2px,color:#cdd6f4
Loading

โœจ Key Features

  • ๐ŸŽ™๏ธ Streaming Dictation with VAD: Speech is automatically sliced at natural pauses (vadSilenceMs, default 700ms) and typed into the composer in real time.
  • ๐ŸŒŠ Voice Notes with Cancel Window: Record your thought and have it automatically dispatched to the agent after a safety countdown (autoSendMs, default 4000ms).
  • ๐ŸŽฎ Tactile Push-to-Talk:
    • Mouse: Hold the wave button โ€” releasing sends the message; dragging pointer away discards.
    • Keyboard: Hold Ctrl (or custom hotkey) for hands-free speaking; press Esc to cancel.
  • โšก Zero-Latency In-Browser Captions (browser): Chrome Web Speech API recognition runs 100% locally with live floating captions as you speak.
  • ๐Ÿ›ก๏ธ Ironclad Multi-Provider Fallbacks: If your primary cloud provider runs out of credits or hits a 429 rate limit, requests seamlessly fail over down the chain.
  • ๐Ÿง  Context Glossary Injection: Automatically extracts code variables and identifiers from your composer draft to steer STT model accuracy on technical jargon.
  • ๐ŸŽต Embedded Audio Player: Preview, scrubber, and playback of your recorded voice message directly in chat and the composer dock.
  • ๐Ÿ”‡ Hardware Noise Suppression Toggle: Configurable in settings to toggle browser-level noise suppression, echo cancellation, and auto gain control.
  • ๐Ÿ“Š Provider Latency & Health Dashboard: Live visual telemetry of provider latency (ms), success rates, and errors directly within the settings UI.
  • ๐Ÿ”’ Zero API Key Leakage: Keys are resolved on the host via ctx.credentials (credentialRef) and never transmitted to browser clients.
  • ๐Ÿ–ฅ๏ธ Offline Local Whisper Server: Automatically boots and manages whisper.cpp (whisper-server) with on-the-fly ffmpeg transcode.
  • โšก SenseVoice-ONNX / Sherpa-ONNX (0.8.11): Ultra-fast (~50โ€“100ms) non-autoregressive local STT engine with automatic emotion/event tag stripping. Supports both Sherpa-ONNX HTTP and OpenAI-compatible endpoints.
  • ๐ŸŒ Realtime Audio Streaming (0.8.11): Low-latency WebSocket bridge (/dsh-voice/realtime) for OpenAI Realtime API or local Sherpa-ONNX streaming. API keys stay securely on the host.
  • ๐ŸŒŠ Liquid Wave & Dynamic Orb Visualizer (0.8.12): Smooth animated audio visualization in the recording pill with real-time mic volume reactivity. Switch between organic multi-layer liquid waves, pulsating radiant orb, classic bars, or off.

๐ŸŽฎ Four Ways to Speak

Mode Gesture / Trigger Behavior
Dictation Click ๐ŸŽ™๏ธ Mic Speech is sliced on pauses (vadSilenceMs) and typed live into composer
Voice Message Click ๐ŸŒŠ Wave Records until stopped, then sends after cancel window (autoSendMs)
Mouse PTT Hold ๐ŸŒŠ Wave Records while held; release sends message, drag off button to discard
Keyboard PTT Hold Ctrl Hands-free recording; release sends message, press Esc to discard

Tip

You can customize the keyboard modifier in settings (hotkey: Control, Alt, Shift, or any KeyboardEvent.code).


๐Ÿ› ๏ธ Supported Providers Matrix

Provider Key Service Backend Default Model Credential Ref Features & Notes
browser Web Speech API Native Browser None Zero latency, floating live captions in Chrome
deepgram Deepgram API nova-2 DEEPGRAM_API_KEY Ultra-fast cloud transcription
groq Groq Whisper whisper-large-v3-turbo GROQ_API_KEY Near-instant inference speed
hf HuggingFace Inference openai/whisper-large-v3 HF_TOKEN High-accuracy open Whisper
local-whisper Local whisper.cpp Server defined None 100% private, offline, no internet needed
sensevoice SenseVoice-ONNX / Sherpa-ONNX SenseVoiceSmall None Ultra-fast (~50ms) local non-autoregressive STT

๐Ÿš€ Ready-Made Presets (Plug & Play)

Just specify the name in your fallback chain and add the corresponding API key:

  • openai (whisper-1) โ†’ OPENAI_API_KEY
  • siliconflow (SenseVoiceSmall) โ†’ SILICONFLOW_API_KEY
  • mistral (voxtral-mini-latest) โ†’ MISTRAL_API_KEY
  • openrouter (google/gemini-2.5-flash) โ†’ OPENROUTER_API_KEY
  • deepinfra (whisper-large-v3-turbo) โ†’ DEEPINFRA_API_KEY
  • fireworks (whisper-v3-turbo) โ†’ FIREWORKS_API_KEY

๐Ÿ“ฆ Quick Installation

dsh plugin --profile web add @goodandready/dsh-voice

Important

Restart DSH Web UI after installation (systemctl --user restart dsh-web) and refresh your browser tab.


โš™๏ธ Configuration

Open Settings โ†’ Plugins โ†’ Plugin settings โ†’ Voice in the Web UI:

- id: dsh-voice
  config:
    dictation:
      language: ru
      vadSilenceMs: 700
      chain:
        - provider: deepgram
        - provider: groq
        - provider: local-whisper
    message:
      language: ru
      autoSendMs: 4000
      chain:
        - provider: openai
        - provider: local-whisper
    hotkey: Control
    autoStart: true
    whisperModel: /models/ggml-medium-q8_0.bin

๐Ÿค– Agent Tool & HTTP API

Agent Tool (transcribe_audio)

Registers transcribe_audio(file_path, language?) in ctx.tools, allowing agents to analyze audio files, interview recordings, and voice notes directly from disk.

Internal HTTP Endpoints

  • POST /dsh-voice/transcribe โ€” { dataBase64, mimeType, mode } โ†’ { ok, text, provider, tookMs }
  • POST /dsh-voice/polish โ€” { text } โ†’ { ok, text }
  • GET /dsh-voice/status โ€” Returns daemon status, active chains, SenseVoice and realtime config.
  • GET /dsh-voice/realtime โ€” WebSocket upgrade for low-latency audio streaming (OpenAI Realtime API / Sherpa-ONNX). Accepts binary audio chunks, returns JSON text deltas.

๐Ÿ“„ License

MIT ยฉ GooDAnDReaDY