> ## Documentation Index
> Fetch the complete documentation index at: https://wiz-myvocal.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Speech to Text — Real-time quickstart

> Stream microphone or line audio to MyVocal and receive partial and final transcripts.

The real-time API streams audio over a WebSocket and returns partial and final transcripts as they are
produced. Like Batch, it uses the existing MyVocal API key, the same Characters balance and the same
History.

Read [availability and known limitations](/guides/stt-availability), especially the connection
recovery and processing-storage notes.

## 1. Create a session

Create a session with an `Idempotency-Key` and the options you want. The response returns a
`sessionId`, a MyVocal `transcriptionId` and a `streamUrl`.

<Card title="POST /sound_clone/api/v1/stt/realtime/sessions" icon="plus" href="/api-reference/stt/createSession">
  Set the input encoding, the commit mode and the requested capabilities.
</Card>

```bash theme={null}
# Generate once for this session; reuse it and the same body for a retry.
IDEMPOTENCY_KEY=$(python -c 'import uuid; print(uuid.uuid4())')
SESSION=$(curl -s https://api.myvocal.ai/sound_clone/api/v1/stt/realtime/sessions \
  -H "accessKey: $MYVOCAL_ACCESS_KEY" -H "Content-Type: application/json" \
  -H "Idempotency-Key: $IDEMPOTENCY_KEY" \
  -d '{"languageHint":"en","options":{"inputEncoding":"pcm_s16le_16000"}}')
```

## 2. Send your audio with the right encoding

`inputEncoding` must match the bytes you send. `pcm_s16le_16000` is signed 16-bit little-endian mono
at 16 kHz; `mulaw_8000` is one μ-law byte per sample. Always declare the microphone's real sample
rate, or a second of audio is interpreted as the wrong duration.

Send audio as it is captured, in frames of about **100 ms**. Any frame size is accepted (there is no
minimum frame length or frame rate), but each message has a fixed overhead, so merge very small
pieces such as a browser's 128-sample `AudioWorklet` quanta. When you stream a file, pace frames by
sample time (one second of audio per second) and keep the socket's local send buffer bounded instead
of writing the whole file at once. `capturedSamples` counts everything captured so far, including
audio not sent yet; `sampleOffset` counts the samples already sent in the current epoch.

## 3. Connect the socket

A server client sends the `accessKey` header. A browser must not put its long-lived key in a URL: it
uses a short-lived, single-use ticket and connects with `?ticket=`. In a customer-facing app, your
backend keeps the API key and creates the session and ticket; send only the ticket and session
information to the browser. The browser then connects straight to MyVocal's socket, so your backend
needs no CORS opening and no general proxy: only fixed routes to create a session, issue a ticket,
read a session and finish it, for sessions it created. The
[browser example](#complete-examples) is a runnable version of that pattern.

<Card title="POST /sound_clone/api/v1/stt/realtime/sessions/{sessionId}/tickets" icon="ticket" href="/api-reference/stt/createTicket">
  Issues a single-use browser ticket bound to the owner and the session.
</Card>

Send audio with a MyVocal event envelope and read the results:

```json theme={null}
{ "eventType": "audio.append", "epoch": 1,
  "payload": { "audioBase64": "...", "sampleOffset": 0, "capturedSamples": 16000 } }
```

The server answers with `session.ready`, `transcript.delta`, `transcript.final`, `transcript.alignment`, `transcript.revision`,
`transcript.entities`, `usage.updated`, `session.notice`, `session.completed` and `session.error`,
each with an `eventId`, `sessionId` and a monotonic `sequence`.

## 4. Pause, resume or commit

* `segment.commit` submits the audio so far and **keeps** the session recording.
* `session.pause` / `session.resume` close and reopen a recording segment (a new epoch).
* `session.finish` flushes the tail, settles once and completes the task. `session.completed`
  arrives once the last final line and the supplements you requested are in, or at the cutoff of
  about 8 seconds after the last final line; a supplement still missing then is reported with a
  `SUPPLEMENT_NOT_RECEIVED` notice rather than waited for. See
  [Finish a session](/api-reference/stt/finishSession).

<Card title="POST /sound_clone/api/v1/stt/realtime/sessions/{sessionId}/finish" icon="flag-checkered" href="/api-reference/stt/finishSession">
  The terminal call; a repeated finish returns the same task and never charges twice.
</Card>

## 5. Events in detail

Client to server, each as `{ "eventType", "epoch", "payload" }`:

| `eventType` | `payload` |
| - | - |
| `audio.append` | `audioBase64` (the frame in the declared encoding), `sampleOffset` (samples already sent), `capturedSamples` (total captured), optional `precedingText` |
| `segment.commit` | `capturedSamples` |
| `session.pause` / `session.resume` | `capturedSamples` |
| `session.finish` | `capturedSamples` |

Server to client, each as `{ "eventId", "sessionId", "sequence", "eventType", "payload" }`:

| `eventType` | `payload` |
| - | - |
| `session.ready` | the current session view (epoch, cursors, `accepting`) |
| `transcript.delta` | `text` (temporary; replaces the previous delta) |
| `transcript.final` | `id`, `text`, `revisedText`, `languageCode`, `languageProbability`, `entities`, `words`, `characterAlignments` |
| `transcript.revision` | `id`, `text`, `revisedText` (the edited transcript arrived for that line) |
| `transcript.entities` | `id`, `entities` (detected entities for that line; an empty list when its entities result arrived with none) |
| `transcript.alignment` | `id`, `languageCode`, `languageProbability`, `words`, `characterAlignments` (late language/alignment) |
| `usage.updated` | `estimatedCharacters`, `billableCharacters`, `elapsedMs`, `sentSamples`, `receivedSamples`, `capturedSamples` (all decimal strings), `accepting`, `epoch`, and `rotate`/`errorCode` when present |
| `session.notice` | `code`, `message`; for example `PROCESSING_STORAGE_NOT_APPLIED`, or `SUPPLEMENT_NOT_RECEIVED` with `supplements` (`alignment`, `entities`, `revision`, `language`) and `lineIds` |
| `session.completed` | the terminal session view |
| `session.error` | `errorCode` (a MyVocal code) |

All 64-bit quantities are decimal strings; `accepting`, `rotate` and `epoch` stay boolean/number.
Each line's `languageCode` (two-letter, for example `en`) and `languageProbability` are that line's
detection result when the model returned one; `languageHint` from the session is a request, not a
detection. A frame that is not valid JSON, an unknown `eventType`, invalid base64 or a malformed
`capturedSamples` is answered with `session.error` `INPUT_INVALID` and does not change the session.

## 6. Rotation, recovery and errors

* A `usage.updated` with `rotate: true` and `accepting: false` means the server wants a new segment:
  call `session.pause`, then `session.resume` (a new `epoch`), and continue sending from the new
  epoch. The frame that received `rotate: true` was **not** stored: re-send every frame after the
  last acknowledged `sentSamples`, starting again at `sampleOffset` 0 in the new epoch. This applies
  to the last frame too: before `session.finish`, wait until `sentSamples` covers every captured
  sample and no rotation is pending.
  `session.ready` reports the live epoch and cursors so a reconnect knows where to resume.
* A server `session.error` carries a MyVocal code (`INPUT_INVALID`, `PROVIDER_UNAVAILABLE`,
  `INSUFFICIENT_CHARACTERS`, `TEMPORARILY_UNAVAILABLE`, …). A plain disconnect is not a successful
  finish. Read the session over REST, then call [finish](/api-reference/stt/finishSession) to settle
  the audio the server can confirm. The result may be `PARTIAL`; unsent or unconfirmed audio is not
  restored by reconnecting. Full pause/resume across a service replacement is not yet verified.
* A repeated frame at the same `sampleOffset`, a repeated `segment.commit` and a repeated
  `session.finish` are idempotent; they neither double-count nor create a second task.

The production socket can disconnect after about 60 seconds without traffic. Server clients can
send a WebSocket ping every 20 seconds while idle; browser clients should handle disconnects and
read the session before deciding what to do next. `heartbeatIntervalMs` does not keep the
client-to-MyVocal connection alive on its own.

## Complete examples

* Node.js client using a one-use ticket, paced by sample time: [`examples/stt/stt_realtime_minimal.mjs`](https://github.com/MyVocal-AI/API/blob/main/myvocal-api-docs/examples/stt/stt_realtime_minimal.mjs)
* Python server socket, paced by sample time: [`examples/stt/stt_realtime_minimal.py`](https://github.com/MyVocal-AI/API/blob/main/myvocal-api-docs/examples/stt/stt_realtime_minimal.py)
* Browser microphone: a same-origin Node.js backend that keeps the key plus the page, \~100 ms
  `AudioWorklet` frames, bounded stop/finish and close-out: [`examples/stt/browser/`](https://github.com/MyVocal-AI/API/tree/main/myvocal-api-docs/examples/stt/browser)
  (`node server.mjs`, see the [examples README](https://github.com/MyVocal-AI/API/blob/main/myvocal-api-docs/examples/stt/README.md))

Read [Characters and quotes](/guides/characters-and-quotes) for the billing rule.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.