1. Create a session
Create a session with anIdempotency-Key and the options you want. The response returns a
sessionId, a MyVocal transcriptionId and a streamUrl.
POST /sound_clone/api/v1/stt/realtime/sessions
Set the input encoding, the commit mode and the requested capabilities.
2. Send your audio with the right encoding
inputEncoding must match the bytes you send. pcm_s16le_16000 is signed 16-bit little-endian mono
at 16 kHz; mulaw_8000 is one μ-law byte per sample. Always declare the microphone’s real sample
rate, or a second of audio is interpreted as the wrong duration.
Send audio as it is captured, in frames of about 100 ms. Any frame size is accepted (there is no
minimum frame length or frame rate), but each message has a fixed overhead, so merge very small
pieces such as a browser’s 128-sample AudioWorklet quanta. When you stream a file, pace frames by
sample time (one second of audio per second) and keep the socket’s local send buffer bounded instead
of writing the whole file at once. capturedSamples counts everything captured so far, including
audio not sent yet; sampleOffset counts the samples already sent in the current epoch.
3. Connect the socket
A server client sends theaccessKey header. A browser must not put its long-lived key in a URL: it
uses a short-lived, single-use ticket and connects with ?ticket=. In a customer-facing app, your
backend keeps the API key and creates the session and ticket; send only the ticket and session
information to the browser. The browser then connects straight to MyVocal’s socket, so your backend
needs no CORS opening and no general proxy: only fixed routes to create a session, issue a ticket,
read a session and finish it, for sessions it created. The
browser example is a runnable version of that pattern.
POST /sound_clone/api/v1/stt/realtime/sessions/{sessionId}/tickets
Issues a single-use browser ticket bound to the owner and the session.
session.ready, transcript.delta, transcript.final, transcript.alignment, transcript.revision,
transcript.entities, usage.updated, session.notice, session.completed and session.error,
each with an eventId, sessionId and a monotonic sequence.
4. Pause, resume or commit
segment.commitsubmits the audio so far and keeps the session recording.session.pause/session.resumeclose and reopen a recording segment (a new epoch).session.finishflushes the tail, settles once and completes the task.session.completedarrives once the last final line and the supplements you requested are in, or at the cutoff of about 8 seconds after the last final line; a supplement still missing then is reported with aSUPPLEMENT_NOT_RECEIVEDnotice rather than waited for. See Finish a session.
POST /sound_clone/api/v1/stt/realtime/sessions/{sessionId}/finish
The terminal call; a repeated finish returns the same task and never charges twice.
5. Events in detail
Client to server, each as{ "eventType", "epoch", "payload" }:
Server to client, each as
{ "eventId", "sessionId", "sequence", "eventType", "payload" }:
All 64-bit quantities are decimal strings;
accepting, rotate and epoch stay boolean/number.
Each line’s languageCode (two-letter, for example en) and languageProbability are that line’s
detection result when the model returned one; languageHint from the session is a request, not a
detection. A frame that is not valid JSON, an unknown eventType, invalid base64 or a malformed
capturedSamples is answered with session.error INPUT_INVALID and does not change the session.
6. Rotation, recovery and errors
- A
usage.updatedwithrotate: trueandaccepting: falsemeans the server wants a new segment: callsession.pause, thensession.resume(a newepoch), and continue sending from the new epoch. The frame that receivedrotate: truewas not stored: re-send every frame after the last acknowledgedsentSamples, starting again atsampleOffset0 in the new epoch. This applies to the last frame too: beforesession.finish, wait untilsentSamplescovers every captured sample and no rotation is pending.session.readyreports the live epoch and cursors so a reconnect knows where to resume. - A server
session.errorcarries a MyVocal code (INPUT_INVALID,PROVIDER_UNAVAILABLE,INSUFFICIENT_CHARACTERS,TEMPORARILY_UNAVAILABLE, …). A plain disconnect is not a successful finish. Read the session over REST, then call finish to settle the audio the server can confirm. The result may bePARTIAL; unsent or unconfirmed audio is not restored by reconnecting. Full pause/resume across a service replacement is not yet verified. - A repeated frame at the same
sampleOffset, a repeatedsegment.commitand a repeatedsession.finishare idempotent; they neither double-count nor create a second task.
heartbeatIntervalMs does not keep the
client-to-MyVocal connection alive on its own.
Complete examples
- Node.js client using a one-use ticket, paced by sample time:
examples/stt/stt_realtime_minimal.mjs - Python server socket, paced by sample time:
examples/stt/stt_realtime_minimal.py - Browser microphone: a same-origin Node.js backend that keeps the key plus the page, ~100 ms
AudioWorkletframes, bounded stop/finish and close-out:examples/stt/browser/(node server.mjs, see the examples README)
