# Generate the key once; reuse it and the same body if the response is lost.
IDEMPOTENCY_KEY=$(python -c 'import uuid; print(uuid.uuid4())')
curl --request POST \
--url https://api.myvocal.ai/sound_clone/api/v1/stt/realtime/sessions \
--header "accessKey: $MYVOCAL_ACCESS_KEY" \
--header "Idempotency-Key: $IDEMPOTENCY_KEY" \
--header 'Content-Type: application/json' \
--data '{"languageHint":"en","options":{"inputEncoding":"pcm_s16le_16000","segmentCommitMode":"manual"}}'
{
"code": 1,
"message": "success",
"data": {
"sessionId": "<sessionId>",
"transcriptionId": "<transcriptionId>",
"status": "RECORDING",
"modelId": "myvocal_stt_realtime_v1",
"streamUrl": "<returned socket path>",
"effectiveOptions": {"inputEncoding":"pcm_s16le_16000","sampleRateHz":16000}
}
}
Speech to Text
Create a real-time session
Start a streaming transcription session and get its socket URL.
POST
/
sound_clone
/
api
/
v1
/
stt
/
realtime
/
sessions
# Generate the key once; reuse it and the same body if the response is lost.
IDEMPOTENCY_KEY=$(python -c 'import uuid; print(uuid.uuid4())')
curl --request POST \
--url https://api.myvocal.ai/sound_clone/api/v1/stt/realtime/sessions \
--header "accessKey: $MYVOCAL_ACCESS_KEY" \
--header "Idempotency-Key: $IDEMPOTENCY_KEY" \
--header 'Content-Type: application/json' \
--data '{"languageHint":"en","options":{"inputEncoding":"pcm_s16le_16000","segmentCommitMode":"manual"}}'
{
"code": 1,
"message": "success",
"data": {
"sessionId": "<sessionId>",
"transcriptionId": "<transcriptionId>",
"status": "RECORDING",
"modelId": "myvocal_stt_realtime_v1",
"streamUrl": "<returned socket path>",
"effectiveOptions": {"inputEncoding":"pcm_s16le_16000","sampleRateHz":16000}
}
}
Creates a real-time session. Audio is then streamed over the returned
streamUrl. The session shares
History and the Characters balance with Batch.
# Generate the key once; reuse it and the same body if the response is lost.
IDEMPOTENCY_KEY=$(python -c 'import uuid; print(uuid.uuid4())')
curl --request POST \
--url https://api.myvocal.ai/sound_clone/api/v1/stt/realtime/sessions \
--header "accessKey: $MYVOCAL_ACCESS_KEY" \
--header "Idempotency-Key: $IDEMPOTENCY_KEY" \
--header 'Content-Type: application/json' \
--data '{"languageHint":"en","options":{"inputEncoding":"pcm_s16le_16000","segmentCommitMode":"manual"}}'
{
"code": 1,
"message": "success",
"data": {
"sessionId": "<sessionId>",
"transcriptionId": "<transcriptionId>",
"status": "RECORDING",
"modelId": "myvocal_stt_realtime_v1",
"streamUrl": "<returned socket path>",
"effectiveOptions": {"inputEncoding":"pcm_s16le_16000","sampleRateHz":16000}
}
}
Header
string
required
API key for authentication.
string
required
Binds this request. The same key and the same body (including
title and options) return the same
session; a changed title or option under the same key is IDEMPOTENCY_CONFLICT.Body
string
The spoken language, or omit for automatic detection.
string
A title for the History row, up to 200 characters. Defaults to a generic recording title.
object
Optional processing controls.
Hide properties
Hide properties
string
default:"pcm_s16le_16000"
The byte format you will send:
pcm_s16le_8000, pcm_s16le_16000, pcm_s16le_22050,
pcm_s16le_24000, pcm_s16le_44100, pcm_s16le_48000 or mulaw_8000. PCM is signed 16-bit
little-endian mono; μ-law is one byte per sample. The chosen rate drives the whole session timeline.string
default:"manual"
manual (you commit segments) or vad (the model decides segment boundaries).string[]
Candidate language hints, up to 8. This is not the detected language.
number
Speech threshold for
vad, between 0 and 1.number
Silence before a
vad commit, from 0 to 60000 milliseconds.number
Minimum speech length, an integer from 0 to 60000 milliseconds.
number
Minimum silence length, an integer from 0 to 60000 milliseconds.
boolean
Word and character alignment. Mutually exclusive with
options.suppressBackgroundSpeech; when
background filtering is on and this is unset, alignment defaults to off.boolean
Return the detected language.
string[]
Key terms to bias recognition, up to 50 entries.
boolean
Remove filler words.
string[]
Detect entities of these categories. Mutually exclusive with
options.rewriteInstruction.string
Instruction for an edited transcript, up to 2000 characters. Mutually exclusive with
options.entityCategories.boolean
Filter background speech. Mutually exclusive with timestamps.
number
Processing connection keepalive interval between
500 and 10000 milliseconds. This is not
a client-to-MyVocal socket heartbeat; see connection guidance.boolean
Request processing-side content storage. Account eligibility is not verified; if the model returns
a warning that it was not applied, the session reports a
PROCESSING_STORAGE_NOT_APPLIED notice. Neither acceptance nor absence of a notice guarantees
zero retention; see current limitations.Response
Successful REST calls return{ code: 1, message, data }. The fields below are inside data;
JSON downloads are the documented exception and return the transcript view directly.
number
1 for success.string
Result message.
object
Hide properties
Hide properties
string
The session id (
rts_...).string
The History id (
stt_...), the same entity Batch results use.string
The session state:
RECORDING, PAUSED or CLOSED, or the task status once closed.string
myvocal_stt_realtime_v1.string
The MyVocal socket path. Connect it as a WebSocket; identity comes from the
accessKey header or a
ticket, never from the URL string alone.object
The options in effect, including the resolved
inputEncoding and sampleRateHz.