How the Gateway Works

Understand the request flow from your application to the LLM provider and back.

Architecture

Your application talks to Pura LLM using the OpenAI API format. We authenticate your API key, enforce quotas, route to the configured provider, and return the response — without storing your content.

Your App → Pura LLM Gateway → LLM Provider (EU) ↑ ↓ ↓ OpenAI SDK Auth + Billing Inference Metadata only

Request flow

  1. Your app sends a request to /v1/chat/completions with your Pura LLM API key
  2. The gateway validates your key, checks wallet balance and rate limits
  3. The request is routed to the configured LLM provider in the EU
  4. The provider response is returned to your app in OpenAI format
  5. Token usage metadata is recorded for billing — content is never stored

Supported endpoints

All standard OpenAI-compatible endpoints are available:

EndpointDescription
POST /v1/chat/completionsChat completions with messages
POST /v1/completionsLegacy text completions
POST /v1/embeddingsText embeddings
GET /v1/modelsList available models

OpenAI compatibility

Any library, framework, or tool that supports the OpenAI API works with Pura LLM. Change base_url and api_key — that's it.

Inference traceability: compliance_metadata v1

Contract for the compliance-enabled gateway. Each deployment must pass the conformity gate before activation; the current catalog is not yet certified.

The v2 contract guarantees provider, actual model, EU data_zone, ingress_region, original upstream_request_id, gateway_request_id and UTC timestamp. The ingress region does not identify the execution region: Azure Data Zone Standard and Bedrock geographic profiles can distribute requests across certified zone regions. The region field remains reserved for regional v1 deployments.

provider_raw is best effort: its presence, structure and completeness are not guaranteed. It contains only metadata permitted by provider-specific allowlists. Do not use it as a contractual dependency.

JSON extension — illustrative values
{
  "extensions": {
    "compliance_metadata": {
      "schema_version": "2",
      "provider": "azure",
      "model": "verified-deployment-model",
      "data_zone": "EU",
      "ingress_region": "Italy North",
      "upstream_request_id": "provider-request-id",
      "gateway_request_id": "2c5b008e-3931-4ce8-a7d5-f5e7c38215ea",
      "timestamp": "2026-09-29T12:00:00.000Z",
      "provider_raw": {
        "headers": {
          "apim-request-id": "provider-request-id"
        }
      }
    }
  }
}

The core is available in HTTP headers before the first content event. SSE repeats it in a final chunk with empty choices, before [DONE]. Handle chunks without choices; an interrupted stream cannot guarantee the final chunk.

http
X-Pura-Compliance-Schema-Version: 2
X-Pura-Compliance-Provider: azure
X-Pura-Compliance-Model: verified-deployment-model
X-Pura-Compliance-Data-Zone: EU
X-Pura-Compliance-Ingress-Region: Italy North
X-Pura-Compliance-Upstream-Request-Id: provider-request-id
X-Pura-Compliance-Gateway-Request-Id: 2c5b008e-3931-4ce8-a7d5-f5e7c38215ea
X-Pura-Compliance-Timestamp: 2026-09-29T12:00:00.000Z

The extension is nested in extensions.compliance_metadata. The default mode and X-Pura-Compliance: 2 return the route contract. X-Pura-Compliance: headers preserves the standard OpenAI body and returns metadata in headers without a final metadata chunk. X-Pura-Compliance: 1 is reserved for regional v1 routes and rejects geographic zone routes.

Missing mandatory fields or ambiguous provider attribution produce HTTP 502 before a successful response is opened. Uncertified routes cannot be activated.

Images retain their base64 strings, signed URLs and embedded markers, such as C2PA, without transcoding. We do not watermark text: the gateway does not control generation logits.

For completed requests we record stable traceability fields without prompts, responses or provider_raw. Details are available in Logs after activating the integration. Historical logs and failed deliveries may lack this metadata.