Skip to main content
Interpretation takes an uploaded audio or video file and produces a translated, dubbed version of it, asynchronously. It is not live/interpreting in real time and it is not a microphone or WebSocket service: you upload a file, then poll until the result is ready. No cloned voice, workspace session or callback URL is required — only the accessKey header.
Runnable clients for this guide: examples/interpretation/quickstart.py and examples/interpretation/quickstart.mjs. Each one implements the whole flow below, including the multipart upload and the polling.

Endpoint summary

What input is supported

Read the accepted formats and limits from the service instead of hard-coding them:
  • audioFormats — accepted audio containers/extensions.
  • videoFormats — accepted video containers/extensions.
  • exportFormats — mp3, wav, flac, mp4. mp4 is the composed video and only applies to a video source.
  • maxSourceBytes — the maximum declared upload size (a JSON string).
  • languages[] — the language entries you may request, each shaped { languageKey, displayName, sourceSupport }. Send languageKey (for example es), not displayName.
  • sourceSupport — source-transcription evidence for that key, one of CONFIRMED_ALIAS, CONFIRMED_NATIVE, REJECTED or UNVERIFIED. It is separate from whether the key is a valid target language.
  • languageCatalogState — CONFIGURED when an authoritative list is configured.
Share the same production limits note: the concrete production values are the ones the service returns at request time. Audio and video sources are both supported; a successful audio-source run does not by itself prove that every video container has been exercised end to end.

Step 1: Read capabilities

Look at account.accessState, planRates, account.charactersPerMinutePerLanguage, quoteTtlSeconds and the language/format lists. If languageCatalogState is EMPTY, quoting is not available for any language yet. accessState is ENABLED only when generation is allowed; FREE_LOCKED, ENTERPRISE_UNRESOLVED, SYNCING and UNRESOLVED each mean generation is not currently available, for different reasons.

Step 2: Create the project

sourceLanguage may be omitted, sent as null, as an empty string or as AUTO to auto-detect. settingsVersion is the concurrency anchor: draft and quote requests can carry the version they were based on, and a mismatch is rejected instead of silently overwriting someone else’s change.

Step 3: Set the draft (optional)

For PATCH, “field omitted or null” means keep the current value, while a provided array — including an empty one — replaces the whole list. cloningStrength is only defaulted to 7 when the stored draft has no value yet; it is not re-defaulted on every PATCH.

Step 4: Create the upload session

  • filename and size are required. size is the exact byte length of the file you will upload, and it must be greater than 0 and not greater than maxSourceBytes (read it from capabilities).
  • contentType is optional; when omitted or blank the service uses application/octet-stream.
  • partSizeBytes and totalParts come from the service. Use both as returned; do not assume a fixed 16 MiB chunk size.

Step 5: Sign each part

Ask for one presigned URL per part number, from 1 to totalParts:

Step 6: Upload the bytes

Send the bytes for that part to the returned url, using the returned method, and include every header in requiredHeaders.
The presigned url points at object storage, not at MyVocal. Do not send your accessKey (or any credential) to that URL — the signature in the URL is the authorization. Keep the ETag response header of every part: you need all of them in the next step.A part URL is short-lived (15 minutes by default). If you run out of time, request a fresh signed URL for the same part number — the session itself lasts 24 hours.

Step 7: Complete the upload

Send all parts in part-number order with the exact ETags returned by storage. A missing, extra or mismatched part list is rejected with code = 47126. After the object is assembled, the service probes the media and reads the real container, streams and duration. The response state may already be READY, or still be COMPLETING/PROBING; poll the next endpoint until it is READY.

Step 8: Poll the upload until the source is ready

Poll while state is CREATED, UPLOADING, COMPLETING or PROBING. When it is READY, media describes the detected input (inputFormat, containerName, audioStreamCount, video, dimensions) and sourceVersion is set.
COMPLETING and PROBING do not mean you can generate yet. Only READY means the source is usable. A FAILED state carries an errorCode; 47125 means the file exceeded the size limit or had no content, 47127 means storage was unavailable.
After a successful upload, GET /projects/{projectId} shows source.state = FROZEN with the detected durationMs and media. That frozen source version is what generation uses.

Step 9: Quote the target languages

The cost formula is ceil(durationMs × rate / 60000) per new target language. The rate comes from your plan (planRates / charactersPerMinutePerLanguage).
All Characters values (perTargetCharacters, totalCharacters, availableCharacters, estimatedRemainingCharacters, and every targets[].characters) are JSON strings. Convert them to integers before comparing or adding. Do not compare them lexicographically.
Two important states:
  • state = "ALL_TARGETS_EXIST" means every requested language already exists for this source. Then quoteId is null and totalCharacters is "0". Read existingTargets and do not call generation with a null quoteId.
  • Otherwise state = "ACTIVE" and quoteId is set. Quotes expire after quoteTtlSeconds.
Adding a language later only charges for the new target languages.

Step 10: Generate

Idempotency here has two distinct behaviours, and neither is a byte-for-byte replay:
  • Same key, same body: the existing acceptance is returned; its target/job state may already have advanced. The stable anchors are acceptanceId and each targetId/jobId.
  • Same quote, different key: if that quote has already been accepted, the existing acceptance is returned again without charging twice, before any price re-check. This is intentional and differs from the Music rule.
  • Same key, different body (code = 47111): a real conflict. Do not silently switch to a new key; re-read the project to find out what already happened.

Step 11: Poll each target

Poll GET /projects/{projectId} until every targets[].state is READY, or until a target reaches FAILED_RELEASED (which is the only retryable state). Per-target progress is QUEUED → PREPARING → TRANSLATING → EXPORTING → READY.
  • summaryState rolls the project up: DRAFT (no targets), PROCESSING, PARTIAL_READY, READY (all targets ready) or FAILED.
  • jobs[] and GET /jobs/{jobId} give the per-job view. RECONCILING means the outcome is still being reconciled — it is not a failure and not a refund. Only a real ledger RELEASED state may be described as released.
  • GET /jobs/{jobId} returns 47117 when the job does not exist or does not belong to your account. The API never distinguishes the two.
  • nextPollAfterMs is currently null (the service does not publish a suggested interval yet). Use your own bounded interval, for example 2–5 seconds with jitter.

Retrying a failed target

Only a target with action = "RETRY" (state FAILED_RELEASED) may be retried:
The plan returns the server’s characters, rateVersion and nextReservationCycle. Confirm them with POST /targets/{targetId}/retry; the amount fields are a confirmation of the server’s plan, not client-side pricing. OPERATION_RETRY is replay-safe.

Step 12: Play the result

GET /projects/{projectId} gives each ready target an outputAssetId:
The response contains state, mimeType, a temporary url, expiresAt and renewalAttemptLimit (currently 1: one renewal is allowed, then the authorization error must be surfaced). Request it when you need it; when it has expired, request a new one instead of caching it.

Step 13: Export a file

format must be one of exportFormats (mp3, wav, flac, mp4); mp4 requires a video source. The response returns exportId and state. Export is asynchronous: state starts as PROCESSING, and RETRY is a transient retry of the same export.

Step 14: Download the export

Poll this same exportId while state is PROCESSING or RETRY; neither is a failure, and neither has a download URL yet. Do not create a new generation to work around it. When state is READY, url is the temporary product URL — download the bytes and, if the URL has expired, call this endpoint again for a fresh one. A non-empty url alone is not success: check that the returned state is READY and that the transfer actually produced media bytes.

Managing projects

  • GET /sound_clone/api/v1/interpretation/projects?page=1&pageSize=20 lists your projects (pageSize is 10, 20 or 50).
  • PATCH /sound_clone/api/v1/interpretation/projects/{projectId}/name renames a project.
  • DELETE /sound_clone/api/v1/interpretation/projects/{projectId} deletes it.
  • DELETE /sound_clone/api/v1/interpretation/projects/{projectId}/source removes the uploaded source and returns the project to a draft response.

Billing in one sentence

Quoting is free and charges nothing; generation reserves Characters per language, and each language settles or is released independently according to its real outcome. See Characters, quotes and billing.