Skip to main content
Speech to Text turns an audio file into a transcript. The Batch API is asynchronous: you upload the file, submit a transcription, poll the task and download the result when it is COMPLETED. Batch and real-time results share the same History and the same Characters balance. Everything below uses the existing MyVocal API key. Start by reading Authentication. See availability and known limitations before integrating optional notifications, speaker matching or processing-storage controls.

1. Read capabilities

GET /sound_clone/api/v1/stt/capabilities

Confirms the account state, the rate for this account and the accepted enum values.

2. Upload the audio

Create an upload, send the parts, then complete it. The response gives you an uploadId for the next step.

POST /sound_clone/api/v1/stt/uploads

Presigned multi-part upload; the server probes the file and builds the canonical audio.
The request field is sizeBytes (the response reports the stored size separately). For each returned part, ask for a presigned URL and upload the bytes:
Split the file using the returned partSizeBytes; PUT each part to its signed url and collect the object-storage response ETag. The snippets above show the request sequence; the complete Python example implements the byte splitting and PUT requests. See Create an upload, Sign a part and Complete an upload for the response fields. Submit only after the completed upload reports READY.

3. Submit the transcription

Send the uploadId (or a public mediaUrl instead) with the options you want. Use an Idempotency-Key so a retry never pays twice: the same key and the same body return the same task.

POST /sound_clone/api/v1/stt/transcriptions

Quotes, reserves Characters and starts the task in one call.

4. Poll, then download

GET /sound_clone/api/v1/stt/transcriptions/{transcriptionId}

Read the task, transcript and billing summary. Stop polling at COMPLETED, PARTIAL or FAILED; inspect the result or error instead of polling a failed task forever.

GET /sound_clone/api/v1/stt/transcriptions/{transcriptionId}/download

txt (the default when format is omitted), json, srt, segmented_json, html, docx or pdf.
Download is offered for COMPLETED tasks only; any other status returns RESULT_NOT_READY. A PARTIAL task (a real-time session that ended with some audio unconfirmed) is terminal and never becomes COMPLETED: read its confirmed text from the task view instead of downloading.

Reading the result

  • Language. languageHint echoes what you asked for (auto when omitted); it is not a detection result. detectedLanguage and transcript.languageCode are filled only when the model returned a language, and use the model’s three-letter code (for example eng); they are null otherwise. Real-time lines use two-letter codes (for example en). Compare languages with that difference in mind.
  • Edited transcript. With rewriteInstruction, transcript.revisionStatus is APPLIED (the edited text is in transcript.revisedText), EMPTY (the model returned no edited text), FAILED (with a MyVocal revisionErrorCode) or NOT_REQUESTED. The original text, the segments and the entity offsets always refer to the original transcript.
  • Plain text. txt puts a header line above each segment with its time range (when the segment is timed) and its speaker label, then the text, with a blank line between segments. Turn either part off with exportOptions (includeTimestamps: false, includeSpeakers: false).
  • Subtitles. SRT needs timing: segments without timestamps (for example with alignmentLevel: none) produce no cues, so an untimed result gives an empty SRT. Line wrapping with maxCharactersPerLine breaks at word boundaries; text without spaces, or a single word longer than the limit, is split between whole characters, so no text is lost and no character is cut. maxSegmentChars counts the separators MyVocal inserts when it rebuilds a segment.

Options and export controls

options controls Batch processing. A representative request:
Read the Batch option reference and the live constraints in capabilities. exportFormats only pre-generates those formats when the task completes; every format in downloadFormats stays downloadable either way. Each exportOptions entry applies its controls to one format, both to the pre-generated file and to later downloads of that format, and does not change another format. An unsupported combination is refused at submit with INPUT_INVALID rather than silently ignored. An entity category the model rejects ends the task as FAILED with INPUT_INVALID (not MEDIA_UNREADABLE); MyVocal does not keep its own category allowlist.

Public media URLs

With mediaUrl, MyVocal downloads the file itself, identifying as MyVocal-STT-MediaFetcher/1.0 (+https://docs.myvocal.ai), and re-checks the public address of the real connection and of every redirect (at most 5). A failure carries details.field = "mediaUrl", a details.reason and, when the host answered, details.sourceStatus: MEDIA_UNREADABLE means the file was downloaded but is not readable audio.

Completion notifications

Enable a signed completion callback inside the same options block. Notifications concern COMPLETED tasks; keep polling to detect failures and other terminal outcomes.
Completion delivery is available, but production verification of raw-body signatures, retry delivery and receiver de-duplication is not yet complete. Use polling as your authoritative completion path and verify your receiver before relying on notifications.
  • A notification is sent when notifyUrl is present, unless notifyOnCompletion is false. notifyOnCompletion: true without notifyUrl is refused with INPUT_INVALID; it is never silently accepted. notifyOnCompletion: false with a notifyUrl sends nothing.
  • notifyUrl must be a public http/https URL; private, loopback, link-local and cloud-metadata addresses are refused on the real connection.
  • notificationSigningSecret is write-only: it is stored encrypted and never returned or logged. Changing the URL or the secret under the same Idempotency-Key is a different request.
  • MyVocal attempts delivery with bounded exponential backoff and a stable eventId. Delivery can fail after retries are exhausted; a receiver can also receive duplicates. Each signed delivery carries three MyVocal headers: Read the timestamp from X-MyVocal-Timestamp, then compute the HMAC over the concatenation of the timestamp, a . and the raw request body bytes:
    Verify against the raw bytes exactly as received; do not re-serialize the JSON, and do not sign the body without the timestamp prefix. Then de-duplicate by X-MyVocal-Event-Id.
  • Polling is the authoritative result. A missed or failed callback never changes the task; do not retry the transcription because a callback was missed.

Complete example

The runnable Batch client implements upload, submit, poll and download, then deletes the test transcription. It stops polling at COMPLETED, PARTIAL or FAILED; a PARTIAL task prints its readable text and status and stays in History. Running it again after completion creates a new billable task: examples/stt/stt_batch_minimal.py. Read Characters and quotes for how the charge and the shared balance work, and Asynchronous jobs and retries for retry rules.