Chat and streaming
OpenAI-compatible contract for delegated endpoints.
Chat request
POST /v1/integrations/inference/chat/completions requires Authorization: Bearer ACCESS_TOKEN, the inference scope and messages. Do not send model: the backend uses INTEGRATION_CHAT_MODEL, configured exclusively on the Pura server. Chat and streaming share this mapping. model and routing, credential or metadata overrides are rejected with 422 even in direct HTTP calls.
curl https://ai.puradigital.it/v1/integrations/inference/chat/completions \
-H "Authorization: Bearer YOUR_OAUTH_ACCESS_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"messages":[{"role":"user","content":"Hello"}]}'Per-account catalog
GET /v1/integrations/inference/capabilities returns services.chat and services.audio_transcription, computed from server configuration and paid-period entitlements. It does not return a model catalog or offer a selector. Chat/streaming responses identify their model publicly as pura.
SSE and cancellation
Set stream=true to receive text/event-stream. Preserve events and the [DONE] terminator; do not transform the stream into ordinary JSON. Client disconnection does not cancel provider cost already incurred and accounting can finish after the connection closes.
Admission and in-flight requests
Pura limits pending requests to two per account, shared across all apps and chat/audio. Requests are checked before forwarding; final usage arrives after completion or through callback recovery. A small overrun from already in-flight requests is possible.
Timeouts and uncertain outcomes
If the connection drops after forwarding, the provider may already have executed the request. Do not automatically repeat inference as though it were idempotent: a second call may create a new charge. Check usage and let the user decide whether to retry.