Skip to main content
POST
Create a transcription
Submits one transcription from an uploadId or a public mediaUrl. The call quotes the work, reserves Characters and starts the asynchronous task in a single request.
string
required
API key for authentication.
string
required
Binds this request. The same key and the same body return the same task; the same key with a different body is refused as IDEMPOTENCY_CONFLICT.

Body

string
An upload from POST /uploads. Use exactly one of uploadId or mediaUrl.
string
A public http/https URL. The server fetches it over a real connection and re-checks every redirect; private, loopback, link-local and cloud-metadata addresses are refused. There is no domain or port allowlist. A failure reports details.reason (DESTINATION_NOT_ALLOWED, SOURCE_REFUSED, SOURCE_UNAVAILABLE, SOURCE_UNREACHABLE, TOO_MANY_REDIRECTS or UNSUPPORTED_TYPE) and, when the host answered, details.sourceStatus; see public media URLs.
number
Which audio track to transcribe. Required when the source contains more than one audio track; use the index from the upload’s media.audioTracks.
string
The spoken language, or omit for automatic detection. The hint is echoed as languageHint; it is never reported as the detected language.
string
A title for the History row.
string
Optional model identifier; defaults to myvocal_stt_v1.
object
Processing options: separateSpeakers, maxSpeakerCount, alignmentLevel, entityCategories, redactCategories, redactionStyle, rewriteInstruction, removeDisfluencies, identifyConversationRoles, vocabularyHints, separateChannels, channelResultMode, exportFormats, inputEncoding, processingContentStorage and more. Option combinations that the model does not support are refused at submit with INPUT_INVALID rather than silently dropped. Completion notifications are options too: options.notifyOnCompletion (true to enable), options.notifyUrl (a public URL MyVocal posts the signed event to), options.clientMetadata (opaque caller data echoed back) and options.notificationSigningSecret (the write-only HMAC key). Polling stays the authoritative result. See capabilities for the live list.
boolean
When true, this call waits for a terminal state inside the same request.
number
The wait budget in milliseconds. If the task is still running at the budget, the request returns the task with its identity; it is not cancelled.

Batch options

All fields below belong inside options. Omit fields you do not need.

Option combinations

  • rewriteInstruction cannot be combined with entity detection, redaction or separateChannels.
  • Multi-channel processing does not separate speakers. Do not combine separateChannels with maxSpeakerCount, speakerMergeSensitivity, identifyConversationRoles or pcm_s16le_16 input. The selected audio track can have at most five channels for transcription.
  • Multi-channel combined results require alignment and cannot be combined with entity detection or redaction.
  • matchKnownSpeakers is unavailable. Do not request it.
Each exportOptions entry has a required format and optional controls: Only relevant formats use each control; json is the public transcription view. An entry applies to the pre-generated file and to later downloads of that format. Without timestamps (for example alignmentLevel: none) SRT has no timed cues and is empty. These options do not change recognition or billing.
An entity category the model rejects is detected when the task runs: the task ends FAILED with INPUT_INVALID. See parameter errors found by the model.

Response

Successful REST calls return { code: 1, message, data }. The fields below are inside data; JSON downloads are the documented exception and return the transcript view directly.
number
1 for success.
string
Result message.
object