Skip to main content

Availability

The GenAI metric names, tracing, and Server-Timing are available from v2.2.0.

Overview

Observability in AI Studio is the operational data about the gateway data path: metrics, traces, request timing, health, and logs. Use it to connect AI Studio to your monitoring stack, such as Prometheus, Grafana, or an OpenTelemetry Collector. The embedded gateway in AI Studio and the Edge Gateway use the same proxy code. For this reason, the metrics, traces, and timing on this page work the same on both. Other features also give operational data. They have their own pages:

Metrics

AI Studio and the Edge Gateway expose Prometheus metrics on a /metrics endpoint. The endpoint is secure by default. To turn it on, and for the list of governance metrics, refer to Prometheus and OpenTelemetry Metrics.
If you do not set METRICS_AUTH_TOKEN or METRICS_ALLOW_UNAUTHENTICATED=true, AI Studio and the Edge Gateway do not register the metrics endpoint. A scrape of the Edge Gateway then returns 404. A scrape of AI Studio returns 200 with the HTML of the admin console, because AI Studio serves the console for unknown paths. Prometheus reports this as a parse error.

GenAI Metrics

From v2.2.0, AI Studio also exports metrics that follow the OpenTelemetry semantic conventions for generative AI. Dashboards that you build for other AI gateways can use the same names. The GenAI metrics have these labels:
  • gen_ai_operation_name: chat for LLM requests, and execute_tool for tool calls.
  • gen_ai_provider_name: the LLM vendor.
  • gen_ai_request_model: the model of the request.
  • gen_ai_token_type: on gen_ai_client_token_usage only.
  • error_type: the HTTP status code. AI Studio adds this label only when the request fails.
AI Studio reports some vendors with the names from the conventions: The token type is input for prompt tokens and output for completion tokens. The conventions have no value for prompt cache tokens. AI Studio reports them as cache_read and cache_write.
The GenAI conventions are not stable yet. The metric names can change in a later release of the conventions.

Legacy Metric Names

Three earlier metrics now have GenAI equivalents: By default, AI Studio exports both names, so your dashboards continue to work after an upgrade. The token metric also changed its type, from a counter to a histogram. Where you used aistudio_llm_tokens_total, use gen_ai_client_token_usage_sum. When your dashboards use the GenAI names, set METRICS_LEGACY_NAMES=false to stop the three earlier metrics. The other aistudio_* metrics, such as cost and policy blocks, have no GenAI equivalent. AI Studio always exports them.

Example Queries

Tracing

AI Studio and the Edge Gateway can export OpenTelemetry traces to a collector over OTLP gRPC. Export is off by default. The endpoint format sets the transport security:
  • host:port: AI Studio connects without TLS.
  • https://host:port: AI Studio connects with TLS.
  • http://host:port: AI Studio connects without TLS.
Spans can include the model names of your requests. If the collector is outside your network, use an https:// endpoint. If you turn on tracing without an endpoint, AI Studio logs an error and continues without tracing. The service name is tyk-ai-studio for AI Studio and tyk-microgateway for the Edge Gateway.

Spans

Each LLM request through the gateway creates one span. The span name has the form {operation} {model}, for example chat gpt-4o. The span has the same gen_ai.* attributes as the metrics. When the request fails, the span status is error, and the span records http.response.status_code.

Trace Propagation

The gateway joins the trace of the caller. It reads the W3C traceparent header of the incoming request, and sends the trace context to the LLM vendor. Thus, one request stays on one trace from your ingress, through AI Studio, to the model server. Propagation works also when tracing is off. The ENABLE_TRACING setting controls only whether AI Studio records spans. If you trace the services in front of and behind the gateway, the trace stays connected without tracing in AI Studio.

Gateway Timing Headers

To see how much time the gateway adds to one request, set GATEWAY_SERVER_TIMING=true. Set it on the Edge Gateway, or on AI Studio for its embedded gateway. LLM responses then include a standard Server-Timing header. Browser developer tools and most HTTP clients show this header.
All durations are in milliseconds: The header contains only the values that are known when the gateway writes the headers. On a chunked response, such as a streamed response, the gateway also writes a Server-Timing trailer after the body. The trailer has the complete values, including total, gw, and gw-ttfb. For the OpenAI-compatible endpoints (/ai and the unified /v1 endpoint), the gateway calls the vendor through an internal second request. On these endpoints, upstream is the time of the vendor call, and gw-pre includes the internal request.
The header shows internal timing to every client. Turn it on for tests and diagnosis, not for permanent use in production.

Health Endpoints

AI Studio and the Edge Gateway have health endpoints that need no authentication. Use them for load balancer checks and Kubernetes probes. The /health/detailed endpoint shows details about the plugins of the Edge Gateway, without authentication. Do not make it available outside your network. For example, block the path at your load balancer or ingress. The admin console shows the connection of each Edge Gateway to AI Studio. Go to Edge Gateways to see the connection, the configuration sync status, and the last heartbeat of each edge. Edge Gateways page with one edge that is Connected and Synced, with its version and last heartbeat Open an edge to see its session and the checksum of its loaded configuration. The page also shows the time of its last sync acknowledgment. For the meaning of each value, refer to Edge Gateway Management. Edge Gateway detail page with the connection status, version, build hash, and last heartbeat

Logs

AI Studio and the Edge Gateway write logs to standard output. Use LOG_FORMAT=json when a log collector reads the Edge Gateway logs. The Edge Gateway does not start if LOG_LEVEL or LOG_FORMAT has a value that is not in the list.

Kubernetes

The Edge Gateway Helm chart (microgateway) sets the metrics and tracing variables from its values: The default metrics.allowUnauthenticated: true is for scrapes from inside the cluster. If the metrics endpoint is available from outside the cluster, set metrics.allowUnauthenticated to false and use metrics.bearerTokenSecret. The chart does not render a ServiceMonitor unless metrics.allowUnauthenticated is true or metrics.bearerTokenSecret.name is set. This prevents a ServiceMonitor that scrapes an endpoint that does not exist. For the full deployment, refer to Deploy on Kubernetes.