The cascade pipeline — VAD → STT → LLM → TTS — bottoms out around 1.3–1.6s voice-to-voice: every stage adds serial latency, and end-of-turn has to be reconstructed from transcripts with heuristics. A speech-to-speech model collapses the cascade — audio in, one model, audio out — with turn-taking decided by the model listening. Production S2S APIs run 500–800ms.
Three packages layer on top of each other:
| Package | Adds |
|---|---|
glove-voice-s2s | Run a built Glove agent on a realtime S2S model |
glove-voice-avatar | A lip-synced face over the agent's audio |
glove-voice-livekit | LiveKit as the room transport, plus LiveKit-native avatars |
The whole progression is preserved as runnable examples in the repo: examples/layered-voice → server-voice → s2s-rooms → avatar-rooms → livekit-rooms.
pnpm add glove-voice-s2s| Piece | What it is |
|---|---|
S2SAdapter | The provider contract: one live session — audio in/out, tool calls as events, a text side-channel |
OpenAIRealtimeAdapter | device mode (WebRTC, browser-only): owns the mic and plays the reply itself |
OpenAIRealtimeSocketAdapter | transport mode (WebSocket, Node + browser): 24 kHz PCM both ways |
OpenAILiveAdapter | Server transport for GPT-Live: continuous 16/24 kHz PCM with Responses delegation to Glove tools |
GeminiLiveAdapter | transport mode: 16 kHz in, 24 kHz out — the mode a server-hosted room needs |
RealtimeAgent | Runs a built Glove on an S2S model: its prompt and tools configure the session |
createS2SAdapter | Provider/model/credential factory — args first, S2S_* env second |
s2sDrivenModel | The Glove model slot for S2S-driven agents, optionally carrying the full realtime config |
runConformance | The behavioural suite every adapter must pass |
Every adapter declares mode: "device" | "transport" so a host can refuse a mismatch loudly at startup instead of discovering silence on the first call. Device opens the microphone and plays the reply itself — least code, browser only. Transport moves PCM and nothing else: the only mode a server room or phone bridge can use, because there is no microphone in the process.
Author the agent exactly as you always do — tools, prompt, store — and hand it to RealtimeAgent. One definition, two runtimes: the same tools serve text turns through the normal loop and voice turns through the provider's.
import { RealtimeAgent, s2sDrivenModel } from "glove-voice-s2s";
// The cleanest form: the model slot carries the realtime config, and
// RealtimeAgent derives the provider session from the agent itself.
const agent = new Glove({
model: s2sDrivenModel({
label: "s2s-front",
provider: "openai", // or S2S_PROVIDER
voice: "marin", // or S2S_VOICE
turnDetection: { type: "semantic_vad", eagerness: "low" }, // typed knobs
}),
systemPrompt, store, displayManager, compaction_config,
}).fold(myTool).build();
const rt = new RealtimeAgent({ agent });
await rt.start();Or pass an explicit adapter — it always wins over the model-slot config:
const rt = new RealtimeAgent({
agent, // a built Glove (IGloveRunnable)
adapter: createS2SAdapter({ provider: "gemini" }),
instructions: SPOKEN_PERSONA, // re-voice the text prompt for speech
excludeTools: ["render_chart"], // withhold tools that don't belong in a call
});
rt.on("user_said", (t) => log("caller:", t));
rt.on("agent_said", (t) => log("agent:", t));
await rt.start();
// transport mode: wire audio yourself
micStream.on("pcm", (pcm) => rt.sendAudio(pcm));
rt.adapter.on("audio", (pcm, format) => speaker.play(pcm, format.sampleRate));
// push an async result into the live call — the model relays it out loud
rt.inject("the lookup finished: covered until 2031", { respond: true });What the voice path deliberately does not do. The provider owns the loop, so the Glove Executor never runs: requiresPermission is not enforced (put gated tools in excludeTools); display.pushAndWait tools get no handOver and will throw (exclude them — voice-first tools should return descriptive data instead); and tool calls and transcripts are not persisted to the store or fired as subscriber events (use RealtimeAgent's own user_said / agent_said / tool_started / tool_finished events to log).
What is shared with the text path: tool definitions and JSON schemas, Zod input validation before run, the system prompt, and the renderData-stays-client-side contract — the bridge strips renderData and summary before anything reaches the provider, exactly like the model adapters do.
| Env | Meaning |
|---|---|
S2S_PROVIDER | openai (WS transport) | openai-webrtc (browser device) | openai-live (server WS) | gemini. Unset: whichever key exists, OpenAI first |
S2S_MODEL | Model id; unset uses the provider default |
OPENAI_API_KEY / GEMINI_API_KEY | The credential when no getToken/apiKey is passed — server-side only |
S2S_TURN_DETECTION | OpenAI Realtime: semantic_vad (default) | server_vad (snappier barge-in) |
S2S_BACKEND_MODEL | GPT-Live Responses backend; default gpt-5.6-luna |
A missing credential fails at construction with the env var name, not at connect() with a 401. Realtime and Gemini expose their full turn-taking surface as typed config, so a mistyped field fails at compile time instead of being silently ignored:
// OpenAI — the model judges WHETHER you were done
createS2SAdapter({ provider: "openai", turnDetection: {
type: "semantic_vad",
eagerness: "low", // low | medium | high | auto
}});
// …or threshold-driven
createS2SAdapter({ provider: "openai", turnDetection: {
type: "server_vad",
threshold: 0.6, // how loud counts as speech
silence_duration_ms: 700, // trailing silence before end-of-turn
prefix_padding_ms: 300,
idle_timeout_ms: 10_000,
}});
// turnDetection: null → manual / push-to-talk
// Gemini
createS2SAdapter({ provider: "gemini", realtimeInput: {
automaticActivityDetection: {
startOfSpeechSensitivity: "START_SENSITIVITY_LOW",
endOfSpeechSensitivity: "END_SENSITIVITY_LOW",
silenceDurationMs: 700,
},
activityHandling: "NO_INTERRUPTION", // default is barge-in
}});Available in glove-voice-s2s 0.4.0. Select openai-live through the same factory or agent model slot. The voice model delegates reasoning and tool selection to a Responses backend; RealtimeAgent executes the agent's existing tools. Selecting gpt-live-1 on the openaiadapter uses the wrong protocol.
import { RealtimeAgent, s2sDrivenModel } from "glove-voice-s2s";
// Use this model in your existing Glove agent definition.
const model = s2sDrivenModel({
provider: "openai-live",
model: "gpt-live-1", // default voice model
backendModel: "gpt-5.6-luna", // independent Responses model
voice: "marin",
sampleRate: 24000, // or 16000, in both directions
instructions: "Speak briefly. Delegate lookups before answering.",
});
// Build agent with this model, your tools, and your system prompt.
const rt = new RealtimeAgent({ agent });
rt.on("transcript", ({ role, delta, startMs, endMs }) => {
// Preserve timed fragments and allow overlapping speakers.
});
rt.on("usage", ({ seconds, final }) => {
// Cumulative duration: replace the previous snapshot, do not sum it.
});
await rt.start();
// Host supplies continuous paced PCM, including silence, and plays audio events.
// On teardown, await rt.stop() before releasing the host resources.OPENAI_API_KEY authenticates the server WebSocket. Browser apps need a server relay or LiveKit; the Realtime token helper below does not create a Live session. instructions sets the voice persona; backendInstructions can override the backend's agent prompt. Parallel tool calls are off by default; when enabled, all tool results are collected before the backend continues once.
transcriptfragments. Live emits no final user_said or agent_said turns. capabilities.transcriptsis continuous.capabilities.speechLifecycleis host. Feed actual playback state into rt.adapter.notifyPlaybackState?.(speaking). Avatar hosts must supply utterance boundaries; missing packets and backend response completion do not establish a voice turn ending. Existing room examples need host changes. attachRealtime relays PCM but does not fill input silence gaps or forward captions.interrupt() flushes and mutes output until the host calls resumeOutput?.(). Tool work continues. Do not add Realtime VAD cancellation, audio commits, or response.cancel.respond: false injects thinking;respond: true injects commentary that Live may paraphrase. Updates are split into bounded appends and are non-atomic. Instruction refreshes append; voice changes require a new session.rt.stop() for final usage. Timeout or transport loss rejects with finalization unconfirmed. Voice duration and the Responses backend incur separate charges.Live client delegation is not implemented. The existing voice permission/display limitations still apply, and this session does not run Foundry's durable lifecycle. Expose a tool backed by the Foundry client for work that needs an instance, conversation, and observable run. Protocol and local transport tests do not establish live account access or audible behavior. See the package guide for the full contract.
API keys never reach the browser. Mint an ephemeral token server-side:
import { createOpenAIRealtimeToken } from "glove-voice-s2s/server";
const { token } = await createOpenAIRealtimeToken({
apiKey: process.env.OPENAI_API_KEY!,
instructions: SPOKEN_PERSONA,
voice: "marin",
tools: [{ name: "delegate_to_worker", description: "…", parameters: {} }],
});import { OpenAIRealtimeAdapter } from "glove-voice-s2s";
const s2s = new OpenAIRealtimeAdapter({
getToken: () => fetchToken("/api/voice/s2s-token"),
});
s2s.on("tool_call", async ({ callId, name, arguments: args }) => {
const result = await runWorker(JSON.parse(args).request); // your heavy agent
s2s.sendToolResult(callId, result); // relayed out loud
});
await s2s.connect();A realtime avatar provider is a lip-sync renderer over an audio stream: PCM in, a talking face out on a WebRTC surface. That is exactly the shape of the audio events a transport-mode S2SAdapter already emits — so the avatar is a rendering layer, not a replacement for any of the stack. The mic path, tools and delegation are untouched.
mic ──▶ S2S model (brain + voice) ──▶ agent PCM ──▶ AvatarAdapter ──▶ the face
│ tool calls unchanged (provider WebRTC surface)
▼
worker over the meshpnpm add glove-voice-avatarimport { RealtimeAgent } from "glove-voice-s2s";
import { TavusEchoAdapter, attachAvatar } from "glove-voice-avatar";
const rt = new RealtimeAgent({ agent }); // the voice stack, exactly as before
await rt.start();
const avatar = new TavusEchoAdapter({
apiKey: process.env.TAVUS_API_KEY!, // server-side only
faceId: process.env.TAVUS_FACE_ID!,
// palId omitted → ensureEchoPal() reuses-or-creates a MINIMAL echo PAL
// (no greeting, no TTS layer) so the ONLY voice is ever the agent's.
sendInteraction: (event) => duct.send({ t: "avatar_interaction", event }),
});
const detach = await attachAvatar(rt, avatar);
avatar.view; // { kind: "webrtc-room", url: "https://…" } — hand to the clientattachAvatar is the whole bridge: audio → sendAudio, agent_speech_stopped → endUtterance, interrupted → interrupt. Barge-in therefore follows the voice automatically. AvatarView is a tagged union — a WebRTC room URL (Tavus/Daily) or an SDK session token (Anam) — so a client knows how to attach without knowing the provider.
| Adapter | Mode |
|---|---|
TavusEchoAdapter | Tavus pipeline_mode: "echo" — our PCM as base64 24 kHz events; the caller joins the conversation's Daily room |
AnamPassthroughAdapter | Anam audio-passthrough (Anam's own LLM/TTS stay out of the loop). The server mints the token, the browser owns the SDK session, so the adapter needs a sendCommand courier |
Writing your own? Implement AvatarAdapter — connect(), sendAudio(), endUtterance(), interrupt() (always safe, conformance-enforced) — and run runAvatarConformance against a fake transport.
glove-voice-livekit is two halves sharing one room connection. LiveKitTransport is the room leg every LiveKit-backed voice host otherwise hand-rolls: join, publish the agent's voice as a paced WebRTC track, feed remote mic tracks back out as PCM events, carry JSON on the data channel. Barge-in is server-authoritative — clear() flushes the outbound AudioSource queue, so there is no client playback buffer to chase.
pnpm add glove-voice-livekitimport { LiveKitTransport, attachRealtime, mintParticipantToken } from "glove-voice-livekit";
const transport = new LiveKitTransport({
url: process.env.LIVEKIT_URL!,
token: await mintParticipantToken(
{ apiKey: process.env.LIVEKIT_API_KEY!, apiSecret: process.env.LIVEKIT_API_SECRET! },
{ roomName: "call-42", identity: "agent" },
),
});
await transport.connect();
attachRealtime(rt, transport); // mics → model, model → track, interrupt → flush
await rt.start();TavusLiveKitAvatar and AnamLiveKitAvatar implement the same AvatarAdapter contract (and pass its conformance suite), so a face over LiveKit is interchangeable with the Daily-based one. Under the hood they speak LiveKit's published avatar protocol: the provider's worker joins your room as a second participant and publishes synchronized voice and face itself. A Glove agent is indistinguishable from a LiveKit Agents worker as far as the avatar can tell.
import { TavusLiveKitAvatar, mintAvatarToken } from "glove-voice-livekit";
import { attachAvatar } from "glove-voice-avatar";
// The avatar publishes the voice on the agent's behalf — don't double it.
const transport = new LiveKitTransport({ url, token, publishAgentAudio: false });
await transport.connect();
attachRealtime(rt, transport, { agentAudio: false });
const avatar = new TavusLiveKitAvatar({
apiKey: process.env.TAVUS_API_KEY!,
faceId: process.env.TAVUS_FACE_ID!, // minimal echo PAL ensured automatically
livekitUrl: url,
avatarToken: await mintAvatarToken(creds, { roomName: "call-42" }),
});
await attachAvatar(rt, avatar);