Availability
The GenAI metric names, tracing, and Server-Timing are available from v2.2.0.
Overview
Observability in AI Studio is the operational data about the gateway data path: metrics, traces, request timing, health, and logs. Use it to connect AI Studio to your monitoring stack, such as Prometheus, Grafana, or an OpenTelemetry Collector. The embedded gateway in AI Studio and the Edge Gateway use the same proxy code. For this reason, the metrics, traces, and timing on this page work the same on both. Other features also give operational data. They have their own pages:Metrics
AI Studio and the Edge Gateway expose Prometheus metrics on a/metrics endpoint. The endpoint is secure by default. To turn it on, and for the list of governance metrics, refer to Prometheus and OpenTelemetry Metrics.
If you do not set
METRICS_AUTH_TOKEN or METRICS_ALLOW_UNAUTHENTICATED=true, AI Studio and the Edge Gateway do not register the metrics endpoint. A scrape of the Edge Gateway then returns 404. A scrape of AI Studio returns 200 with the HTML of the admin console, because AI Studio serves the console for unknown paths. Prometheus reports this as a parse error.GenAI Metrics
From v2.2.0, AI Studio also exports metrics that follow the OpenTelemetry semantic conventions for generative AI. Dashboards that you build for other AI gateways can use the same names.
The GenAI metrics have these labels:
gen_ai_operation_name:chatfor LLM requests, andexecute_toolfor tool calls.gen_ai_provider_name: the LLM vendor.gen_ai_request_model: the model of the request.gen_ai_token_type: ongen_ai_client_token_usageonly.error_type: the HTTP status code. AI Studio adds this label only when the request fails.
The token type is
input for prompt tokens and output for completion tokens. The conventions have no value for prompt cache tokens. AI Studio reports them as cache_read and cache_write.
The GenAI conventions are not stable yet. The metric names can change in a later release of the conventions.
Legacy Metric Names
Three earlier metrics now have GenAI equivalents:
By default, AI Studio exports both names, so your dashboards continue to work after an upgrade. The token metric also changed its type, from a counter to a histogram. Where you used
aistudio_llm_tokens_total, use gen_ai_client_token_usage_sum.
When your dashboards use the GenAI names, set METRICS_LEGACY_NAMES=false to stop the three earlier metrics. The other aistudio_* metrics, such as cost and policy blocks, have no GenAI equivalent. AI Studio always exports them.
Example Queries
Tracing
AI Studio and the Edge Gateway can export OpenTelemetry traces to a collector over OTLP gRPC. Export is off by default.
The endpoint format sets the transport security:
host:port: AI Studio connects without TLS.https://host:port: AI Studio connects with TLS.http://host:port: AI Studio connects without TLS.
https:// endpoint.
If you turn on tracing without an endpoint, AI Studio logs an error and continues without tracing.
The service name is tyk-ai-studio for AI Studio and tyk-microgateway for the Edge Gateway.
Spans
Each LLM request through the gateway creates one span. The span name has the form{operation} {model}, for example chat gpt-4o. The span has the same gen_ai.* attributes as the metrics. When the request fails, the span status is error, and the span records http.response.status_code.
Trace Propagation
The gateway joins the trace of the caller. It reads the W3Ctraceparent header of the incoming request, and sends the trace context to the LLM vendor. Thus, one request stays on one trace from your ingress, through AI Studio, to the model server.
Propagation works also when tracing is off. The ENABLE_TRACING setting controls only whether AI Studio records spans. If you trace the services in front of and behind the gateway, the trace stays connected without tracing in AI Studio.
Gateway Timing Headers
To see how much time the gateway adds to one request, setGATEWAY_SERVER_TIMING=true. Set it on the Edge Gateway, or on AI Studio for its embedded gateway. LLM responses then include a standard Server-Timing header. Browser developer tools and most HTTP clients show this header.
The header contains only the values that are known when the gateway writes the headers. On a chunked response, such as a streamed response, the gateway also writes a
Server-Timing trailer after the body. The trailer has the complete values, including total, gw, and gw-ttfb.
For the OpenAI-compatible endpoints (/ai and the unified /v1 endpoint), the gateway calls the vendor through an internal second request. On these endpoints, upstream is the time of the vendor call, and gw-pre includes the internal request.
Health Endpoints
AI Studio and the Edge Gateway have health endpoints that need no authentication. Use them for load balancer checks and Kubernetes probes.
The
/health/detailed endpoint shows details about the plugins of the Edge Gateway, without authentication. Do not make it available outside your network. For example, block the path at your load balancer or ingress.
The admin console shows the connection of each Edge Gateway to AI Studio. Go to Edge Gateways to see the connection, the configuration sync status, and the last heartbeat of each edge.


Logs
AI Studio and the Edge Gateway write logs to standard output.
Use
LOG_FORMAT=json when a log collector reads the Edge Gateway logs. The Edge Gateway does not start if LOG_LEVEL or LOG_FORMAT has a value that is not in the list.
Kubernetes
The Edge Gateway Helm chart (microgateway) sets the metrics and tracing variables from its values:
The default
metrics.allowUnauthenticated: true is for scrapes from inside the cluster. If the metrics endpoint is available from outside the cluster, set metrics.allowUnauthenticated to false and use metrics.bearerTokenSecret.
The chart does not render a ServiceMonitor unless metrics.allowUnauthenticated is true or metrics.bearerTokenSecret.name is set. This prevents a ServiceMonitor that scrapes an endpoint that does not exist.
For the full deployment, refer to Deploy on Kubernetes.