Tyk AI Studio 2.2 is GA (and it’s mostly about the boring bits)

Tyk AI Studio 2.2 is generally available today. It makes the gateway one governed front door for every model, tool and MCP server your organisation uses, and it ships the unglamorous controls (ownership, approvals, audit, budgets) that decide whether AI survives first contact with your security team.

The office fridge problem

Most companies’ AI estate looks like the office fridge. Everyone has something in there, nothing is labelled, and there’s a shared OpenAI key taped to the door. Something at the back has been quietly growing since the last reorg, and nobody will admit it’s theirs.

The industry’s answer so far has mostly been a proxy that counts tokens. That’s useful the way a thermometer is useful in a fridge. It tells you it’s cold. It doesn’t tell you whose yoghurt it is, who’s allowed to eat it, or who had the last one.

Plenty of enterprises are adopting AI in a state of Torschlusspanik (gate-closing panic: the fear that the door is shutting and you’d better sprint through it). Sprinting is fine. Sprinting with the company card and no receipts is how you end up explaining yourself to the CFO. 2.2 is our attempt at the rest of the fridge: labels, locks, a list of who took what, and a limit on the snacks. It’s 176 commits and the longest changelog I’ve ever signed off. I read all of it. Mostly.

What’s in 2.2

Most of the governance below is in the Enterprise Edition. The unified endpoint, failover, tool filters and the new chat ship to everyone.

One front door for every model

Apps and agents now call a single OpenAI-compatible endpoint with one key and name the model as vendor/model: OpenAI, Anthropic, Bedrock, Gemini, Ollama, whatever you’ve approved. Claude Code can point at Tyk with one environment variable and run on Bedrock, with budgets, guardrails and analytics applied. Every model can carry a failover waterfall, so when a provider has a bad afternoon the request moves to the next model you chose, with the same governance at every step.

Semantic routing: the right model for the question

Not every prompt deserves your most expensive model. “What’s our holiday policy?” does not need the same brain as “refactor this billing service”, in the same way you don’t call a surgeon to put on a plaster. Semantic Routers let a client send smart/auto and have the gateway choose the model from what the prompt actually means, after authentication and inside the app’s grants.

The router works down a short ladder: explicit routes, session affinity, keyword and regex rules, embedding similarity to example utterances, then an optional LLM judge for the close calls. A route can go straight to a model or hand off to a Model Router that picks the vendor, so “complex goes to the premium model, everything else to the fast one” is a single piece of config.

If a classifier errors or times out, the request takes the default route instead of failing. Every decision lands in the response headers and the logs, the example utterances never leave the gateway, and shadow mode lets you measure a router against live traffic before it changes a single answer.

Semantic caching: the same question twice, paid for once

Your users ask the same questions all day, in slightly different words. The Advanced LLM Cache now answers a paraphrase from the cache as well as an exact repeat, using any OpenAI-compatible embedding model and an in-memory or Redis vector index. A hit never crosses apps, models, system prompts or tool sets, and anything with tool calls or images only ever gets an exact match.

A word of honesty, because embeddings are clever but not wise. They capture what a question is about, not which way round it goes, so “10 miles in km” and “10 km in miles” look like old friends (they are not). That’s why semantic matching is off by default and holds a high bar: 0.95 cosine similarity. Switch it on for FAQ-style traffic, let apps whose prompts turn on numbers or negation opt out, and check the response headers to see when a hit was semantic and how close it was.

Guardrails you don’t have to write

Guardrails are a new kind of filter: pick a detector, choose block, redact or log, and you’re done, no scripting. A built-in library catches 37 credential types, 17 kinds of personal data and the usual prompt-injection tricks. Native connectors add Lakera, Microsoft Presidio, Azure AI Content Safety, Azure AI Language and Amazon Bedrock Guardrails.

They run on prompts, on streamed responses and on tool calls in both directions, which is where the interesting attacks live now. Every hit becomes a compliance event that records what was found, never the text itself.

Governance an auditor can actually read

Fine-grained RBAC arrives with a separate publish right, so the person who writes a change isn’t the person who makes it live. An append-only audit trail records every admin action with a field-level diff. Governed metadata puts owners, risk tier and data classification on every model, tool and data source, and in enforce mode nothing saves without them. No risk tier, no save.

Team budgets pool and cap spend per team, and signed webhooks push events into whatever your workflows already use. The new Asset Catalog extends the same register to things that never touch the gateway, such as agents, prompts and skills. Each gets three named owners, a version history and an approval workflow, which answers the question every auditor opens with: who owns this agent?

MCP without the shadow MCP

The Tyk Dashboard MCP integration brings MCP servers running on Tyk Gateway into the AI Portal catalogue, next to models and tools. A developer finds an approved server, builds an app and mints their own key against policies the platform team pinned. The gateway stays the only thing on the data path. One platform, one catalogue, one key, and no MCP servers living in someone’s personal config file.

Faster, and harder to knock over

The microgateway sustains about seven times the throughput it used to: roughly 9,000 requests a second on a 4-vCPU edge in our AWS benchmark, with p50 overhead under a millisecond. Near its memory limit it now sheds load with a polite 503 instead of falling over. We also ran the biggest security and correctness sweep of the 2.x line, including a live conformance suite against every vendor we support. It found things. We fixed them. That’s rather the point of looking.

The chat got rebuilt too. Client tools pause an agent until a person approves the action, and generative UI lets a model answer with a table or a chart instead of a wall of text.

Before you upgrade

It’s a big release, so read the upgrade notes before you roll it out. Four things will catch people:

  • A budget of 0 now means zero. Existing zeros become “no limit” on first start, but scripts that send 0 to mean unlimited need to send null.
  • /metrics is closed by default. Set METRICS_AUTH_TOKEN before your Prometheus scrape starts returning 404s.
  • New tools are chat-only until you switch on REST or MCP access. Existing tools keep both.
  • Upgrade the control plane and your microgateways together. Older edges ignore the new routing, failover and budget fields.

Go and try it

2.2 is available now for Community and Enterprise, as container images, Helm charts and Linux packages. The full release notes are here: 2.2 release notes. If you’d rather see it than read about it, book a walkthrough.

AI traffic is API traffic with better PR. We’ve spent more than ten years making API traffic boring, governable and fast, and 2.2 does the same for the new stuff. Boring is the point.

Martin Buhr, Founder and CEO, Tyk

Share the Post:

Related Posts

Start for free

Get a demo

Ready to get started?

You can have your first API up and running in as little as 15 minutes. Just sign up for a Tyk Cloud account, select your free trial option and follow the guided setup.