Skip to main content
The real-time API streams audio over a WebSocket and returns partial and final transcripts as they are produced. Like Batch, it uses the existing MyVocal API key, the same Characters balance and the same History. Read availability and known limitations, especially the connection recovery and processing-storage notes.

1. Create a session

Create a session with an Idempotency-Key and the options you want. The response returns a sessionId, a MyVocal transcriptionId and a streamUrl.

POST /sound_clone/api/v1/stt/realtime/sessions

Set the input encoding, the commit mode and the requested capabilities.

2. Send your audio with the right encoding

inputEncoding must match the bytes you send. pcm_s16le_16000 is signed 16-bit little-endian mono at 16 kHz; mulaw_8000 is one μ-law byte per sample. Always declare the microphone’s real sample rate, or a second of audio is interpreted as the wrong duration. Send audio as it is captured, in frames of about 100 ms. Any frame size is accepted (there is no minimum frame length or frame rate), but each message has a fixed overhead, so merge very small pieces such as a browser’s 128-sample AudioWorklet quanta. When you stream a file, pace frames by sample time (one second of audio per second) and keep the socket’s local send buffer bounded instead of writing the whole file at once. capturedSamples counts everything captured so far, including audio not sent yet; sampleOffset counts the samples already sent in the current epoch.

3. Connect the socket

A server client sends the accessKey header. A browser must not put its long-lived key in a URL: it uses a short-lived, single-use ticket and connects with ?ticket=. In a customer-facing app, your backend keeps the API key and creates the session and ticket; send only the ticket and session information to the browser. The browser then connects straight to MyVocal’s socket, so your backend needs no CORS opening and no general proxy: only fixed routes to create a session, issue a ticket, read a session and finish it, for sessions it created. The browser example is a runnable version of that pattern.

POST /sound_clone/api/v1/stt/realtime/sessions/{sessionId}/tickets

Issues a single-use browser ticket bound to the owner and the session.
Send audio with a MyVocal event envelope and read the results:
The server answers with session.ready, transcript.delta, transcript.final, transcript.alignment, transcript.revision, transcript.entities, usage.updated, session.notice, session.completed and session.error, each with an eventId, sessionId and a monotonic sequence.

4. Pause, resume or commit

  • segment.commit submits the audio so far and keeps the session recording.
  • session.pause / session.resume close and reopen a recording segment (a new epoch).
  • session.finish flushes the tail, settles once and completes the task. session.completed arrives once the last final line and the supplements you requested are in, or at the cutoff of about 8 seconds after the last final line; a supplement still missing then is reported with a SUPPLEMENT_NOT_RECEIVED notice rather than waited for. See Finish a session.

POST /sound_clone/api/v1/stt/realtime/sessions/{sessionId}/finish

The terminal call; a repeated finish returns the same task and never charges twice.

5. Events in detail

Client to server, each as { "eventType", "epoch", "payload" }: Server to client, each as { "eventId", "sessionId", "sequence", "eventType", "payload" }: All 64-bit quantities are decimal strings; accepting, rotate and epoch stay boolean/number. Each line’s languageCode (two-letter, for example en) and languageProbability are that line’s detection result when the model returned one; languageHint from the session is a request, not a detection. A frame that is not valid JSON, an unknown eventType, invalid base64 or a malformed capturedSamples is answered with session.error INPUT_INVALID and does not change the session.

6. Rotation, recovery and errors

  • A usage.updated with rotate: true and accepting: false means the server wants a new segment: call session.pause, then session.resume (a new epoch), and continue sending from the new epoch. The frame that received rotate: true was not stored: re-send every frame after the last acknowledged sentSamples, starting again at sampleOffset 0 in the new epoch. This applies to the last frame too: before session.finish, wait until sentSamples covers every captured sample and no rotation is pending. session.ready reports the live epoch and cursors so a reconnect knows where to resume.
  • A server session.error carries a MyVocal code (INPUT_INVALID, PROVIDER_UNAVAILABLE, INSUFFICIENT_CHARACTERS, TEMPORARILY_UNAVAILABLE, …). A plain disconnect is not a successful finish. Read the session over REST, then call finish to settle the audio the server can confirm. The result may be PARTIAL; unsent or unconfirmed audio is not restored by reconnecting. Full pause/resume across a service replacement is not yet verified.
  • A repeated frame at the same sampleOffset, a repeated segment.commit and a repeated session.finish are idempotent; they neither double-count nor create a second task.
The production socket can disconnect after about 60 seconds without traffic. Server clients can send a WebSocket ping every 20 seconds while idle; browser clients should handle disconnects and read the session before deciding what to do next. heartbeatIntervalMs does not keep the client-to-MyVocal connection alive on its own.

Complete examples

Read Characters and quotes for the billing rule.