Audio transcription
Multipart uploads, JSON and accounting shared with chat.
Availability
POST /v1/integrations/inference/audio/transcriptions uses INTEGRATION_TRANSCRIPTION_MODEL, configured on the Pura server after EU verification and plan activation. Check services.audio_transcription with capabilities(). subscription_transcription_not_configured (503) means audio is inactive; do not fall back to another provider or the wallet.
File and fields
Send multipart/form-data with one nonempty file, at most 25,000,000 bytes. Accepted extensions: mp3, mp4, mpeg, mpga, m4a, wav, webm, ogg. The provider must support the actual format: the extension does not validate content.
Optional fields: language (two lowercase letters), prompt (at most 4,000 characters), temperature (finite, from 0 to 1), response_format (json or verbose_json). Do not send model: selection is server-side and the field is rejected with 422. Duplicate or unknown fields and multiple files are rejected. Let FormData/HTTPX generate the boundary.
TypeScript example
The Blob must remain immutable during the request. The SDK can replay all bytes after token refresh; invalid files are rejected before networking.
const transcript = await pura.transcribe({
file: audioBlob, filename: 'recording.wav',
language: 'it',
response_format: 'verbose_json',
});
console.log(transcript.text);Python example
Use bytes and an explicit filename. The response contains text and preserves segments when returned by the provider. This is not a Realtime endpoint and does not provide audio streaming.
transcript = pura.transcribe(
audio_bytes, filename="recording.wav", language="it",
response_format="verbose_json")
print(transcript["text"])Usage and data
Audio uses the same account, grant, pending slots and limits as chat. Pura does not store the file in request records or the ledger. Provider processing remains subject to service terms and residency/retention documentation; do not infer a zero-retention guarantee from this endpoint.