Enterprise software, security and engineering notes

The AI Ops Layer and LLM Routers: Multi-Model, Cost and KVKK Compliance

AI and Agents ·

LLM router and AI Ops layer — multi-level highway interchange routing traffic | Aksiyon Soft

What is an LLM router, and why do companies need an AI Ops layer?

An LLM router is the middleware that inspects every AI request coming from your business applications and sends it to the language model best suited to that request. A year ago most companies started by hard-coding one provider’s API key. Today the same company wants a fast, cheap model to summarize customer emails, a stronger model for contract analysis and a model running in Türkiye for requests that contain personal data. Code that calls three models from three different places quickly becomes unmanageable.

This article covers the AI Ops layer that sits between business applications and multiple models: routing rules, fallback, personal-data masking, logging, quality evaluation and budget limits. With concrete examples, we explain how Article 9 of Türkiye’s data protection law (KVKK), amended in 2024, affects sending data to models abroad, where local options such as EVREN and Kumru plug into this architecture, and how a single OpenAI-compatible interface reduces vendor lock-in.

In short

  • An LLM router sends each request to the right model based on data sensitivity, cost, latency and quality.
  • The AI Ops layer adds fallback, PII masking, logging, evaluation and budget control around the router.
  • Since the KVKK Article 9 amendment, cross-border transfers are possible with appropriate safeguards such as standard contracts, but your rule set must reflect that.
  • EVREN’s OpenAI-compatible inference service and local models such as Kumru give requests containing personal data a strong routing option.
  • When applications talk to one OpenAI-compatible interface, switching models becomes a configuration task, not a code change.

• • •

What exactly does an AI Ops layer do?

Internative’s Türkiye software industry report for the first half of 2026 names the “AI operations layer” as its seventh trend: instead of a single model call, a middleware layer that routes each request to the right model, handles errors and provides monitoring. Its second trend notes that local language models are reaching the field and that a hybrid architecture of “local LLM + global LLM router” is spreading. That matches what we see in our own projects.

In practice, the AI Ops layer is the single endpoint your applications see. Behind it run several model providers, a rules engine, a masking service, a log store and an evaluation pipeline. The application says “classify this customer request”; the layer decides which model, in which region and on which budget, does the work.

Diagram: LLM router architecture — apps, an OpenAI-compatible AI Ops layer and on-prem, Türkiye-hosted and foreign model targets
LLM router architecture: applications only talk to the AI Ops layer, which routes each request to an on-prem, Türkiye-hosted or foreign model.

How a router differs from an API gateway

A classic API gateway handles authentication, rate limiting and routing, but it does not look inside the request. An LLM router has to understand the content: does the prompt contain a Turkish national ID number, is the task a short classification or long reasoning, which department is the user in? That is why the router is usually deployed behind the gateway as a separate, content-aware layer.

Five core jobs of the layer

Routing sends requests to models according to rules. Fallback moves a request to a backup model when a provider fails or slows down. Masking hides personal data before it reaches the model and restores it in the response. Logging and evaluation record the cost, latency and quality of every call. Budget control enforces spending limits per department, application or user.

• • •

Which criteria does an LLM router use to decide?

IBM Research compares an LLM router to an air traffic controller: it evaluates the query and sends it to the model in your library that offers the best value, choosing by price, quality, latency or any other criteria you define. In an enterprise setting two more criteria join the list: data sensitivity and where the data is processed.

CriterionWhat is measured?Example ruleWhat to watch
Data sensitivityDoes the prompt contain personal or special-category data?If a national ID, IBAN or health data is detected, route to an on-prem or Türkiye-hosted modelDetection can fail; when in doubt, pick the safe route
CostPrice per token, remaining monthly budgetSimple classification and summarization go to a small modelThe cheapest model is not always the cheapest outcome; count retries too
LatencyTime to first token and total response timeFor live chat screens, models above the target time are excludedBatch jobs tolerate latency; run them in a separate queue
QualityTask-specific evaluation scoreContract analysis only goes to models that pass the threshold on the test setMeasure with your own data, not general benchmarks
AvailabilityError rate, timeouts, quota statusA provider with repeated errors is temporarily removed and traffic moves to the backupThe backup must follow the same data rule
Data locationCountry where the model runs, contract statusThe foreign route opens only for providers with a KVKK Article 9 safeguardAsk about the provider’s sub-processors as well

Order matters. We always build the rules engine in this sequence: data sensitivity and location first, then the quality threshold, and cost and latency last. That way a cheap model can never receive a request containing personal data just because it is “faster”.

Diagram: LLM routing decision tree — personal data, masking, latency and quality rules
The decision tree: the personal-data check comes first, then whether masking is possible, and only then latency and quality rules.
This IBM Technology explainer shows with a short example how requests are routed to different models based on cost and quality.

• • •

How does KVKK Article 9 affect sending data to LLMs abroad?

Sending a prompt that contains personal data to a model running abroad counts as a cross-border transfer under KVKK. Article 9 of Law No. 6698 was amended by Law No. 7499 (Official Gazette, March 12, 2024), and the new provisions took effect on June 1, 2024. The implementing rules were set out in a regulation published in the Official Gazette on July 10, 2024.

According to the Turkish Data Protection Authority’s page on cross-border transfers, data may now be transferred if one of the processing conditions in Article 5 or 6 applies and an adequacy decision exists for the destination country. Without an adequacy decision, one of the appropriate safeguards — such as a standard contract or binding corporate rules — must be in place, provided that data subjects can exercise their rights and access effective legal remedies in the destination country. The standard contract text cannot be modified and must be notified to the Authority within five business days of signature. Limited cases for incidental, non-recurring transfers are also defined.

What this means for router design

These rules should flow directly into the LLM router’s rule set. For every model target, the configuration should record where the data is processed, which safeguard applies and whether the provider acts as a data processor; a request containing personal data must never reach a foreign model without a safeguard. The LLM provider processes data on your behalf, and we cover the contractual and audit side of that relationship in our article on supplier data breaches and KVKK. This article is not legal advice; please review your specific transfer scenarios with your legal counsel.

A word of caution about industry reports, too: the Internative report argues that for companies with KVKK constraints a local LLM is “not an option but a necessity”. Because the amended Article 9 permits transfers with appropriate safeguards, we treat the issue as a balance of risk and cost rather than an absolute ban. A local model route is the simplest way to reduce the safeguard burden, but it is not the only way.

• • •

How do local models such as EVREN and Kumru plug into the router?

The beauty of a router architecture is that adding a new model target does not affect applications. Two notable developments in Türkiye in recent weeks matter here.

EVREN: an OpenAI-compatible inference service processed in Türkiye

The large language model inference service of EVREN, a platform developed within the Presidency of Defence Industries, was announced on September 23, 2026, and the platform officially launched on September 27, 2026. As reported by Anadolu Agency and Dünya, 11 open-weight models are offered through a single OpenAI-compatible API, and prompts, responses and usage records are processed on high-performance GPU infrastructure in Türkiye. The platform uses a contribution-based credit model, access requires e-Devlet identity verification, and API calls made until November 1, 2026 are not deducted from credit balances. The service also offers its own automatic model routing.

For details, see our news posts on the EVREN LLM inference service and the launch of the national AI platform EVREN. We walk through the technical steps of connecting EVREN to existing applications in our guide on connecting EVREN API to enterprise software. From the router’s point of view, EVREN is a natural target for requests that contain personal data but only need to be processed in Türkiye.

Kumru: a Turkish model you can run in-house

Kumru, developed by VNGRS, is a 7.4-billion-parameter, decoder-only language model pre-trained from scratch for Turkish. According to VNGRS, it was trained on 500 GB of data (300 billion tokens), has an 8,192-token context window (roughly 20 A4 pages) and can run on GPUs with 16 GB of VRAM such as the RTX A4000 or 3090. The smaller Kumru-2B is published as open source on Hugging Face. In scenarios where data must never leave the company, a model like this can be the router’s “on-prem” target.

Blue circuit board with two large chips — running a local language model on in-house hardware
A single GPU with 16 GB of VRAM can be enough to run a mid-sized Turkish model such as Kumru in-house; test capacity with real traffic.

• • •

How do you design fallback, PII masking and logging?

The fallback chain

Every route needs at least one backup model, and the backup must obey the same data rule. Falling back from an on-prem model to a foreign model for a request containing personal data is not a fallback; it is a breach risk. The right chain is the on-prem model, then a model processed in Türkiye, and if neither is available, a clear “cannot be processed right now” response to the user. Health checks, timeouts and circuit-breaker thresholds should be tuned separately for each provider.

PII masking

A masking service detects fields such as national ID numbers, phone numbers, emails, IBANs and addresses in the prompt and replaces them with placeholders; when the model responds, the placeholders are filled back in with the original values. That lets tasks that do not need personal data, such as summarization, run on a strong foreign model. Detection algorithms make mistakes, though; in free text, masking alone should not count as a safeguard, and sensitive workflows should stay on the local route.

Logging and access

For every call, record the user, application, selected model, routing reason, token count, cost, latency and outcome. Prompt and response texts themselves should be stored masked and for a limited time. Who may access these logs is an authorization design question; our article on RBAC access model design is a good starting point.

• • •

How do you measure cost and quality?

Cost control is the biggest payoff of a router, but you only see it if you measure it. The core idea in the IBM Research article is simple: save the big models for high-value, complex tasks and leave easy tasks to smaller, cheaper models. To do that you need an evaluation set for each task type: real, masked sample requests and their expected outputs.

We build evaluation in three stages. First, offline testing on the evaluation set before a new model joins a route. Second, a controlled trial that sends a small percentage of live traffic to the new model. Third, continuous monitoring through user feedback and sampled human review. On the budget side, define a monthly limit per department and application, an alert as the limit approaches, and a rule that switches to a cheaper route or stops when the limit is exceeded.

These metrics belong in your existing monitoring stack, not on a separate dashboard. Seeing latency and error rates next to application metrics lets you tell quickly whether a problem sits in the model or in the integration.

Line graph on a screen — monitoring LLM router cost, latency and quality metrics
Track cost, latency and quality score per model on the same chart; low price on its own can be misleading.

• • •

How does an OpenAI-compatible interface reduce vendor lock-in?

The fourth trend in the Internative report says companies are becoming less tolerant of vendor lock-in in new investments. In AI, the most practical answer is to write applications against a single OpenAI-compatible interface rather than a specific provider’s SDK. Because many open-source servers and services such as EVREN support this interface, the AI Ops layer can offer applications the same contract, similar to /v1/chat/completions, while the model behind it changes through configuration.

Using a “task profile” instead of a model name in application code makes this even easier: the application sends a logical name such as model="summary-standard", and the router maps it to a real model according to that day’s rules. When a provider changes its prices or a new local model appears, only the rules table is updated. Compatibility is never 100%, though; features such as tool calling and structured output must be tested on every target.

The same layer also serves as the backbone for AI agents. When the agent loop calls the router instead of a provider directly, budget and data rules apply to agents as well. We cover this in our guide on taking AI agents to production, through the lens of the middleware chain.

• • •

LLM router checklist

Before your router goes live, each item below should have a written answer.

  • For every model target, are the processing country, contract status and KVKK Article 9 safeguard on record?
  • Which fields does personal-data detection cover, and is the safe route the default when in doubt?
  • Does each route’s backup model follow the same data rule, and are circuit-breaker thresholds defined?
  • Do applications write to one OpenAI-compatible interface and logical task profiles instead of provider SDKs?
  • Is there an evaluation set and quality threshold for each task type?
  • Are budget limits, alerts and overrun behavior defined per department and application?
  • Are prompt and response texts masked in logs, with retention periods and access rights set?
  • Do cost, latency and error metrics flow into your existing monitoring stack?
  • Is there a written process of offline testing and controlled trials for adding a new model to a route?

• • •

How can Aksiyon Soft help?

Through our API and integration service we design and build the AI Ops layer between your applications and language models, including routing rules, masking, fallback and logging. We connect the router to your ERP, CRM and internal portals through our enterprise software solutions service. You can review our architecture approach on the API and data integration platform page.

Our delivery model starts with discovery: which applications use models for which tasks, which data counts as personal data and which providers can be used under which safeguards are put in writing. We then build an MVP with one application and two model targets, demo the working flow at the end of every sprint, and follow go-live with a hypercare period and SLA-backed support. We are headquartered in Samsun and work remotely with teams across Türkiye, with planned on-site visits when the project needs them.

Frequently asked questions

How many models do I need to use an LLM router?

Two are enough: for example, a model processed in Türkiye for requests with personal data and a strong general model for everything else. A router’s value comes less from the number of models than from managing the rules in one place.

Does a router add latency to every request?

A rule-based router usually adds negligible overhead compared with model response time. Routers that classify content with a separate model add more latency, which is why starting with simple rules makes sense.

Can I send prompts containing personal data to a model abroad?

The amended KVKK Article 9 allows transfers when there is a processing condition plus an adequacy decision or an appropriate safeguard such as a standard contract. Review your specific scenario with legal counsel, and on the technical side, close routes without safeguards at the router level.

Can we use EVREN in enterprise applications?

Because EVREN’s inference service offers an OpenAI-compatible API, it can technically be added to the router as a target. Access requires e-Devlet identity verification; we recommend reviewing its terms of use and credit model separately for your project.

Should we buy an off-the-shelf AI gateway or build our own?

Off-the-shelf products give you a fast start, but KVKK rules, internal system integration and local model targets often require customization. During discovery we compare both options against your requirements list.

How long does it take to set up a router?

An MVP with one application and two model targets can go live in a few sprints. Masking, an evaluation pipeline and migrating many applications widen the scope; we share a written timeline after discovery.

Let's talk about your project

If you want to use several language models while balancing cost, quality and KVKK compliance, send us a short summary via the contact page. Together we will work out which applications should move behind the router, which model targets fit and what the integration scope looks like.

Sources

Subscribe to blog and news

Get an email when we publish. Unsubscribe any time.

Related posts