Delegated inference
Services and server configuration
Use GET /v1/integrations/inference/capabilities to check available services. Chat and streaming use INTEGRATION_CHAT_MODEL; audio uses INTEGRATION_TRANSCRIPTION_MODEL. These variables belong exclusively to the Pura server. The host app does not receive a model catalog or select a model. All services share the OAuth grant and account limits. Unconfigured audio returns subscription_transcription_not_configured.
Request boundaries
Send messages and supported generation options, or a file for transcription. Do not send model: the backend rejects it with 422 even in direct HTTP calls. Routing overrides, provider credentials and metadata are rejected. Pura inserts the model and billing correlation on the server.
Errors and retry
HTTP 402: weekly_limit_exceeded, monthly_limit_exceeded or subscription_inactive. Read usage reset times before retrying a cap error. HTTP 429 subscription_requests_pending means two requests await billing; avoid an immediate retry loop. Revoked or expired credentials require OAuth handling. An uncertain timeout may still incur provider cost: do not blindly replay it.
curl https://ai.puradigital.it/v1/integrations/inference/chat/completions \
-H "Authorization: Bearer YOUR_OAUTH_ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"Ciao"}]}'