Enterprise software, security and engineering notes
The AI Ops Layer and LLM Routers: Multi-Model, Cost and KVKK Compliance
AI and Agents ·
Author: Mehmet DOĞAN
Editor: Mehmet DOĞAN
What is an LLM router, and why do companies need an AI Ops layer?
An LLM router is the middleware that inspects every AI request coming from your business applications and sends it to the language model best suited to that request. A year ago most companies started by hard-coding one provider’s API key. Today the same company wants a fast, cheap model to summarize customer emails, a stronger model for contract analysis and a model running in Türkiye for requests that contain personal data. Code that calls three models from three different places quickly becomes unmanageable.
This article covers the AI Ops layer that sits between business applications and multiple models: routing rules, fallback, personal-data masking, logging, quality evaluation and budget limits. With concrete examples, we explain how Article 9 of Türkiye’s data protection law (KVKK), amended in 2024, affects sending data to models abroad, where local options such as EVREN and Kumru plug into this architecture, and how a single OpenAI-compatible interface reduces vendor lock-in.
In short
- An LLM router sends each request to the right model based on data sensitivity, cost, latency and quality.
- The AI Ops layer adds fallback, PII masking, logging, evaluation and budget control around the router.
- Since the KVKK Article 9 amendment, cross-border transfers are possible with appropriate safeguards such as standard contracts, but your rule set must reflect that.
- EVREN’s OpenAI-compatible inference service and local models such as Kumru give requests containing personal data a strong routing option.
- When applications talk to one OpenAI-compatible interface, switching models becomes a configuration task, not a code change.
• • •
What exactly does an AI Ops layer do?
Internative’s Türkiye software industry report for the first half of 2026 names the “AI operations layer” as its seventh trend: instead of a single model call, a middleware layer that routes each request to the right model, handles errors and provides monitoring. Its second trend notes that local language models are reaching the field and that a hybrid architecture of “local LLM + global LLM router” is spreading. That matches what we see in our own projects.
In practice, the AI Ops layer is the single endpoint your applications see. Behind it run several model providers, a rules engine, a masking service, a log store and an evaluation pipeline. The application says “classify this customer request”; the layer decides which model, in which region and on which budget, does the work.
How a router differs from an API gateway
A classic API gateway handles authentication, rate limiting and routing, but it does not look inside the request. An LLM router has to understand the content: does the prompt contain a Turkish national ID number, is the task a short classification or long reasoning, which department is the user in? That is why the router is usually deployed behind the gateway as a separate, content-aware layer.
Five core jobs of the layer
Routing sends requests to models according to rules. Fallback moves a request to a backup model when a provider fails or slows down. Masking hides personal data before it reaches the model and restores it in the response. Logging and evaluation record the cost, latency and quality of every call. Budget control enforces spending limits per department, application or user.
• • •
Which criteria does an LLM router use to decide?
IBM Research compares an LLM router to an air traffic controller: it evaluates the query and sends it to the model in your library that offers the best value, choosing by price, quality, latency or any other criteria you define. In an enterprise setting two more criteria join the list: data sensitivity and where the data is processed.
| Criterion | What is measured? | Example rule | What to watch |
|---|---|---|---|
| Data sensitivity | Does the prompt contain personal or special-category data? | If a national ID, IBAN or health data is detected, route to an on-prem or Türkiye-hosted model | Detection can fail; when in doubt, pick the safe route |
| Cost | Price per token, remaining monthly budget | Simple classification and summarization go to a small model | The cheapest model is not always the cheapest outcome; count retries too |
| Latency | Time to first token and total response time | For live chat screens, models above the target time are excluded | Batch jobs tolerate latency; run them in a separate queue |
| Quality | Task-specific evaluation score | Contract analysis only goes to models that pass the threshold on the test set | Measure with your own data, not general benchmarks |
| Availability | Error rate, timeouts, quota status | A provider with repeated errors is temporarily removed and traffic moves to the backup | The backup must follow the same data rule |
| Data location | Country where the model runs, contract status | The foreign route opens only for providers with a KVKK Article 9 safeguard | Ask about the provider’s sub-processors as well |
Order matters. We always build the rules engine in this sequence: data sensitivity and location first, then the quality threshold, and cost and latency last. That way a cheap model can never receive a request containing personal data just because it is “faster”.
• • •
How does KVKK Article 9 affect sending data to LLMs abroad?
Sending a prompt that contains personal data to a model running abroad counts as a cross-border transfer under KVKK. Article 9 of Law No. 6698 was amended by Law No. 7499 (Official Gazette, March 12, 2024), and the new provisions took effect on June 1, 2024. The implementing rules were set out in a regulation published in the Official Gazette on July 10, 2024.
According to the Turkish Data Protection Authority’s page on cross-border transfers, data may now be transferred if one of the processing conditions in Article 5 or 6 applies and an adequacy decision exists for the destination country. Without an adequacy decision, one of the appropriate safeguards — such as a standard contract or binding corporate rules — must be in place, provided that data subjects can exercise their rights and access effective legal remedies in the destination country. The standard contract text cannot be modified and must be notified to the Authority within five business days of signature. Limited cases for incidental, non-recurring transfers are also defined.
What this means for router design
These rules should flow directly into the LLM router’s rule set. For every model target, the configuration should record where the data is processed, which safeguard applies and whether the provider acts as a data processor; a request containing personal data must never reach a foreign model without a safeguard. The LLM provider processes data on your behalf, and we cover the contractual and audit side of that relationship in our article on supplier data breaches and KVKK. This article is not legal advice; please review your specific transfer scenarios with your legal counsel.
A word of caution about industry reports, too: the Internative report argues that for companies with KVKK constraints a local LLM is “not an option but a necessity”. Because the amended Article 9 permits transfers with appropriate safeguards, we treat the issue as a balance of risk and cost rather than an absolute ban. A local model route is the simplest way to reduce the safeguard burden, but it is not the only way.
• • •
How do local models such as EVREN and Kumru plug into the router?
The beauty of a router architecture is that adding a new model target does not affect applications. Two notable developments in Türkiye in recent weeks matter here.
EVREN: an OpenAI-compatible inference service processed in Türkiye
The large language model inference service of EVREN, a platform developed within the Presidency of Defence Industries, was announced on September 23, 2026, and the platform officially launched on September 27, 2026. As reported by Anadolu Agency and Dünya, 11 open-weight models are offered through a single OpenAI-compatible API, and prompts, responses and usage records are processed on high-performance GPU infrastructure in Türkiye. The platform uses a contribution-based credit model, access requires e-Devlet identity verification, and API calls made until November 1, 2026 are not deducted from credit balances. The service also offers its own automatic model routing.
For details, see our news posts on the EVREN LLM inference service and the launch of the national AI platform EVREN. We walk through the technical steps of connecting EVREN to existing applications in our guide on connecting EVREN API to enterprise software. From the router’s point of view, EVREN is a natural target for requests that contain personal data but only need to be processed in Türkiye.
Kumru: a Turkish model you can run in-house
Kumru, developed by VNGRS, is a 7.4-billion-parameter, decoder-only language model pre-trained from scratch for Turkish. According to VNGRS, it was trained on 500 GB of data (300 billion tokens), has an 8,192-token context window (roughly 20 A4 pages) and can run on GPUs with 16 GB of VRAM such as the RTX A4000 or 3090. The smaller Kumru-2B is published as open source on Hugging Face. In scenarios where data must never leave the company, a model like this can be the router’s “on-prem” target.
• • •
How do you design fallback, PII masking and logging?
The fallback chain
Every route needs at least one backup model, and the backup must obey the same data rule. Falling back from an on-prem model to a foreign model for a request containing personal data is not a fallback; it is a breach risk. The right chain is the on-prem model, then a model processed in Türkiye, and if neither is available, a clear “cannot be processed right now” response to the user. Health checks, timeouts and circuit-breaker thresholds should be tuned separately for each provider.
PII masking
A masking service detects fields such as national ID numbers, phone numbers, emails, IBANs and addresses in the prompt and replaces them with placeholders; when the model responds, the placeholders are filled back in with the original values. That lets tasks that do not need personal data, such as summarization, run on a strong foreign model. Detection algorithms make mistakes, though; in free text, masking alone should not count as a safeguard, and sensitive workflows should stay on the local route.
Logging and access
For every call, record the user, application, selected model, routing reason, token count, cost, latency and outcome. Prompt and response texts themselves should be stored masked and for a limited time. Who may access these logs is an authorization design question; our article on RBAC access model design is a good starting point.
• • •
How do you measure cost and quality?
Cost control is the biggest payoff of a router, but you only see it if you measure it. The core idea in the IBM Research article is simple: save the big models for high-value, complex tasks and leave easy tasks to smaller, cheaper models. To do that you need an evaluation set for each task type: real, masked sample requests and their expected outputs.
We build evaluation in three stages. First, offline testing on the evaluation set before a new model joins a route. Second, a controlled trial that sends a small percentage of live traffic to the new model. Third, continuous monitoring through user feedback and sampled human review. On the budget side, define a monthly limit per department and application, an alert as the limit approaches, and a rule that switches to a cheaper route or stops when the limit is exceeded.
These metrics belong in your existing monitoring stack, not on a separate dashboard. Seeing latency and error rates next to application metrics lets you tell quickly whether a problem sits in the model or in the integration.
• • •
How does an OpenAI-compatible interface reduce vendor lock-in?
The fourth trend in the Internative report says companies are becoming less tolerant of vendor
lock-in in new investments. In AI, the most practical answer is to write applications against a
single OpenAI-compatible interface rather than a specific provider’s SDK. Because many
open-source servers and services such as EVREN support this interface, the AI Ops layer can offer
applications the same contract, similar to /v1/chat/completions, while the model
behind it changes through configuration.
Using a “task profile” instead of a model name in application code makes this even easier: the
application sends a logical name such as model="summary-standard", and the router
maps it to a real model according to that day’s rules. When a provider changes its prices or a new
local model appears, only the rules table is updated. Compatibility is never 100%, though;
features such as tool calling and structured output must be tested on every target.
The same layer also serves as the backbone for AI agents. When the agent loop calls the router instead of a provider directly, budget and data rules apply to agents as well. We cover this in our guide on taking AI agents to production, through the lens of the middleware chain.
• • •
LLM router checklist
Before your router goes live, each item below should have a written answer.
- For every model target, are the processing country, contract status and KVKK Article 9 safeguard on record?
- Which fields does personal-data detection cover, and is the safe route the default when in doubt?
- Does each route’s backup model follow the same data rule, and are circuit-breaker thresholds defined?
- Do applications write to one OpenAI-compatible interface and logical task profiles instead of provider SDKs?
- Is there an evaluation set and quality threshold for each task type?
- Are budget limits, alerts and overrun behavior defined per department and application?
- Are prompt and response texts masked in logs, with retention periods and access rights set?
- Do cost, latency and error metrics flow into your existing monitoring stack?
- Is there a written process of offline testing and controlled trials for adding a new model to a route?
• • •
How can Aksiyon Soft help?
Through our API and integration service we design and build the AI Ops layer between your applications and language models, including routing rules, masking, fallback and logging. We connect the router to your ERP, CRM and internal portals through our enterprise software solutions service. You can review our architecture approach on the API and data integration platform page.
Our delivery model starts with discovery: which applications use models for which tasks, which data counts as personal data and which providers can be used under which safeguards are put in writing. We then build an MVP with one application and two model targets, demo the working flow at the end of every sprint, and follow go-live with a hypercare period and SLA-backed support. We are headquartered in Samsun and work remotely with teams across Türkiye, with planned on-site visits when the project needs them.
Frequently asked questions
How many models do I need to use an LLM router?
Two are enough: for example, a model processed in Türkiye for requests with personal data and a strong general model for everything else. A router’s value comes less from the number of models than from managing the rules in one place.
Does a router add latency to every request?
A rule-based router usually adds negligible overhead compared with model response time. Routers that classify content with a separate model add more latency, which is why starting with simple rules makes sense.
Can I send prompts containing personal data to a model abroad?
The amended KVKK Article 9 allows transfers when there is a processing condition plus an adequacy decision or an appropriate safeguard such as a standard contract. Review your specific scenario with legal counsel, and on the technical side, close routes without safeguards at the router level.
Can we use EVREN in enterprise applications?
Because EVREN’s inference service offers an OpenAI-compatible API, it can technically be added to the router as a target. Access requires e-Devlet identity verification; we recommend reviewing its terms of use and credit model separately for your project.
Should we buy an off-the-shelf AI gateway or build our own?
Off-the-shelf products give you a fast start, but KVKK rules, internal system integration and local model targets often require customization. During discovery we compare both options against your requirements list.
How long does it take to set up a router?
An MVP with one application and two model targets can go live in a few sprints. Masking, an evaluation pipeline and migrating many applications widen the scope; we share a written timeline after discovery.
Let's talk about your project
If you want to use several language models while balancing cost, quality and KVKK compliance, send us a short summary via the contact page. Together we will work out which applications should move behind the router, which model targets fit and what the integration scope looks like.
Sources
- KVKK — Transfer of Personal Data Abroad
- KVKK — Guide on the Transfer of Personal Data Abroad
- Official Gazette — Regulation on the Procedures and Principles for Transferring Personal Data Abroad (July 10, 2024)
- Anadolu Agency — Defence industry’s national AI platform EVREN launched (September 27, 2026)
- Dünya — Local AI infrastructure grows, 11 models connected to EVREN (September 28, 2026)
- VNGRS — Kumru LLM
- Hugging Face — vngrs-ai/Kumru-2B
- Internative — Türkiye Software Industry 2026 H1 Report: 7 Key Trends
- IBM Research — An air traffic controller for LLMs (October 10, 2024)
Subscribe to blog and news
Get an email when we publish. Unsubscribe any time.
Related posts
Software Buyer Guides
Gaziantep Software Partner: Export ERP, e-Invoicing and B2B Portals
A guide to Gaziantep software needs for textile, carpet and food exporters: export ERP, e-invoice and customs integration, multi-plant production and B2B dealer portals.
Software Buyer Guides
Malatya Software Partner: Apricot Exports, Traceability and Business Continuity
How Malatya software projects can support apricot processing and exports, OIZ textiles and post-earthquake rebuilding: traceability, export documents, cloud backups and business continuity.