Running LLMs in production is not just about deploying a model. It's about managing an entire technology stack.
For DevOps and MLOps engineers, an open-source LLMOps stack architecture will feel similar because it builds on Kubernetes and other cloud-native platforms.
However, it introduces additional challenges, including hallucination control, prompt and model management, token costs, model evaluation, latency, and AI-specific observability.
So I looked at the key open-source tools that make up the LLMOps stack, what each tool does, and where it fits in the stack.
I have covered everything from high-performance inference with vLLM to Kubernetes-native serving with KServe, Model routing with LiteLLM, and observability and evaluation with Langfuse and Ragas.
Let's get started.
What is LLMOps?
Large Language Model Operations (LLMOps) is a set of practices, tools, and workflows for building, deploying, monitoring, and maintaining large language model applications in production.
The interesting part is that LLMOps is built on DevOps and MLOps practices. However, LLM applications introduce additional challenges that traditional software and machine-learning workflows may not cover.
These challenges include managing prompt versions, evaluating response quality, tracing LLM requests, protecting sensitive data, and monitoring token usage, latency, and costs.
An LLMOps stack brings together the tools needed to manage these processes throughout the application lifecycle.
The following diagram illustrates an example LLMOps stack architecture that combines LLM gateways, authentication, self-hosted LLM inference, and validation and evaluation tools.

This architecture diagram provides context on the different layers of LLMOps and how tools fit into each layer.
Best Open Source LLMOps Tools
The Open Source LLMOps stack is a set of community-driven tools that are used to deploy, manage, monitor, and optimize Large Language Models (LLMs) in production.
The following table shows the key open-source LLMOps tools.
| Tool | Use Case | When to Consider It | When Not to Use It |
|---|---|---|---|
| Guardrails AI | Validates application inputs and outputs | You need reusable content or structure checks | Existing validation meets your requirements |
| LiteLLM | Provides a common gateway to managed and self-hosted models | You need shared access, routing, or usage controls with multiple teams | A direct provider integration based on your needs |
| Gateway API Inference Extension | Routes to the suitable self-hosted models | When you need Gateway API capabilities with inference-aware routing | When you use managed LLMs or when basic routing is enough for your model |
| llm-d | For inference-aware endpoint picking and advanced distributed serving | When you need to find and route to a suitable model based on latency, KV cache, adaptor, etc | A simpler serving deployment meets your performance targets |
| KServe | Deploys and manages model servers using Kubernetes native resources | Your platform uses a consistent deployment method | You do not run model serving on Kubernetes |
| vLLM | Model server that runs a model and serves inference | You want to host a supported model yourself | You use hosted model APIs only |
| Langfuse | Records traces and manages prompts and evaluations | You need to inspect application behavior | Existing tooling already covers these needs |
| Ragas | Evaluates application quality | You need repeatable quality tests | Your existing evaluation framework is sufficient |
Now, let's have a look at them one by one.
1. Guardrails AI (Guardrails & Safety)
When LLM applications move to production, guardrails help prevent unsafe inputs, unwanted outputs, and unexpected model behavior.

Guardrails AI is a Python framework that provides reusable validators that check whether the LLM’s output meets specific requirements, such as privacy rules, safety policies, and valid formats.
Guardrails AI combines structured output and validation into a single framework. You can define a Pydantic or RAIL schema for the output and attach reusable validators to check whether the response meets your requirements.
When validation fails, Guardrails can re-ask the model, correct the output, filter out invalid fields, or raise an error, depending on how you configure it.
Here is why organizations use Guardrails AI?
Guardrails AI solves a common problem in LLM applications: models can produce harmful, unsafe outputs that may contain sensitive information, hallucinate facts, or fail required output formats.
2. LiteLLM (Model Access and Routing)
LiteLLM acts as a centralized AI gateway for LLM traffic. It provides features such as virtual API keys and model routing across different providers, making it much easier to manage, monitor, and control AI workloads across applications.

LiteLLM can be integrated in two ways. You can integrate it into application code. The other approach in a large organization is to deploy LiteLLM as a central proxy server.
One key advantage of LiteLLM is that it reduces the risk of a company becoming locked into a single technology provider.
By providing an OpenAI-compatible interface for more than 100 LLMs, engineering teams can switch between supported model providers with fewer application-level changes.
Here is why organizations use LiteLLM?
As companies start using more LLMs, managing different APIs, authentication methods, usage limits, and access rules can become complicated.
When deployed as a proxy, LiteLLM provides a central layer for managing model access, authentication, usage, and controls across teams.
This is especially useful for larger teams working with multiple models and providers.
NVIDIA uses LiteLLM to provide engineers with a consistent way to access more than 100 AI model endpoints across cloud providers and self-hosted LLMs.
3. Gateway API Inference Extension (Model-Aware Routing)
The Gateway API Inference Extension is a Kubernetes project that extends the capabilities of Gateway API to understand ML models and route traffic accordingly.
Inference Gateway uses its InferencePool resource, which selects model-serving pods by label and uses the EPP pod to point to the appropriate model server pod.
The suitable pod will be chosen based on latency, queue length, KV cache, and LoRA adapters.

The above image is an example that uses AgentGateway as the controller, and its proxy pod acts as the Inference Gateway. When the proxy receives a request, it uses the HTTPRoute to find the InferencePool and, if configured, applies a policy.
The InferencePool points to the configured EPP pod, and the proxy requests that EPP find a suitable model server.
Once the proxy gets the suitable pod, it routes traffic to it.
There are also other Gateways that support the Gateway API Inference Extension, such as Istio, NGINX Gateway Fabric, and Envoy Gateway,
Here is why organizations use Inference Gateway?
A regular Ingress or Gateway API routes traffic to healthy pods, but it doesn't know which cache is in use, the length of the processing queue, etc.
This is where this extension adds its inference-aware routing capability to select the best model-serving pod.
This helps increase the model's performance and GPU utilization.
One example is how Tesla uses a combination of Inference Gateway, llm-d, Kserve, and vLLM to implement prefix-cache-aware routing, which provides 3x more output tokens per second and a 2x reduction in time.
4. LLM-d (Distributed Inference)
llm-d is an open-source, Kubernetes-native inference framework that is designed to run LLMs efficiently at production scale.
llm-d does not replace inference engines such as vLLM. Instead, it adds distributed serving capabilities, which include LLM-aware request routing, autoscaling, KV-cache-aware scheduling, and prefill/decode disaggregation.
The image below shows how llm-d works to help you understand it better.

LLM-d can separate the prefill and decode phases of LLM inference. Prefill processes the input prompt, while decode generates the output tokens.
The prefill server processes the full prompt and creates the KV cache. Then the decode server uses that information to generate the response token by token.
By running these phases on separate model-server instances (prefill server and decode server), LLM-d can reduce interference between them and allow each phase to scale independently, improving GPU utilization and overall inference performance.
Here is why organizations use LLM-d?
Standard Kubernetes load balancing does not account for LLM-specific metrics such as prompt length, KV cache location, GPU load, or request queues.
LLM-d takes over this limitation through LLM-aware scheduling and routing.
For example, it can route a request to a model server that already contains a reusable part of the prompt in its KV cache. This avoids repeating some prompt-processing work.
This may sound like what Inference Gateway does. Because EPP is part of LLM-d and used in the Gateway API Inference Extension to find the suitable pod.
One real-world use case is that IBM Research and Red Hat demonstrated llm-d serving a 753B-parameter model on NVIDIA H100 GPUs to support thousands of coding agents.
5. KServe (Model Serving Orchestrator)
KServe is an open-source, Kubernetes-native controller that deploys, scales and manages machine learning models.
It provides model-serving runtimes like vLLM and Triton, allowing models to run without building custom infrastructure for each tool.

Here is why organizations use KServe?
Running different AI models in production can require separate deployment configurations, scaling systems, networking rules, and monitoring setups.
KServe helps overcome this complexity by providing a consistent way to deploy and manage models on Kubernetes. This reduces the need to build separate serving infrastructure for every model or framework.
One real-world use case is that Bloomberg uses KServe as its production ML Inferencing Platform.
6. vLLM (Model Server)
vLLM is an open-source inference engine that runs the model and serves its endpoint.
It improves serving performance through efficient KV-cache memory management with PagedAttention, along with techniques such as continuous batching and optimized model execution.

vLLM’s PagedAttention (memory-management technique) breaks the KV cache into small, fixed-size blocks.
Instead of requiring one large, contiguous block of GPU memory, these blocks can be placed in different available memory locations and tracked individually, similar to virtual memory in an operating system.
This reduces wasted GPU memory and lets vLLM handle more sequences at once. As a result, the GPU is used more efficiently, and overall serving speed improves.
Here is why organizations use vLLM?
Running LLMs at scale requires companies to balance throughput, latency, infrastructure costs, and hardware utilization.
vLLM helps to overcome these challenges by managing memory efficiently and batching incoming requests. This lets organizations serve more users while making better use of available hardware. As a result, several large organizations use vLLM in production.
One real-world example is, LinkedIn uses vLLM to support its wide range of generative AI use cases and serve its large, active audience.
7. Langfuse (Observability and Tracing)
Langfuse is an open-source LLM observability framework that includes prompt management, evaluations, analytics, and everything in a single self-hosted solution.

When an application is instrumented with Langfuse, it turns each LLM interaction into a structured trace so that you can inspect it step by step.
This makes it easier to find exactly where an application went wrong or became expensive.
Here is why organizations use Langfuse?
Traditional application monitoring often focuses on latency, error rates, resource usage, and uptime. However, an LLM application can return a successful 200 response while still producing an incorrect, irrelevant, or hallucinated answer.
Langfuse's current observability SDKs are built on OpenTelemetry, and Langfuse can also ingest OpenTelemetry traces directly.
Here are some companies that use Langfuse in production:
- Canva uses Langfuse to monitor and improve its AI-powered customer support system, including its multi-agent Help Assistant and OmniAgent, which handles complex, multi-step support tickets.
- Cresta uses Langfuse to trace and debug its multi-service AI agent pipelines across development and production.
8. Ragas (Evaluation and Testing)
A small prompt change or model update can silently degrade your application's quality, even when everything appears to be working.
LLM evaluation frameworks help catch these issues before they reach production, just like automated tests prevent buggy code from being deployed.
Ragas is an open-source evaluation framework built specifically for LLM applications.
It provides strong support for RAG systems, AI agents, and custom evaluation workflows, offering metrics such as faithfulness, answer relevancy, context precision, and context recall to help you measure retrieval performance and response accuracy.

The main advantage is that companies can test their applications against a dataset, compare different models, prompts, and retrieval strategies, and keep track of metrics such as faithfulness and relevance.
Here is why organizations use Ragas?
Companies use Ragas to move beyond manual testing and measure how well their LLM applications actually perform.
It helps teams evaluate retrieval and generation separately, run experiments, and use evaluation results to continuously improve their applications.
For RAG applications, teams can measure context precision, context recall, response relevance, and faithfulness. Ragas also supports AI agent evaluations, including tool-call accuracy.
These evaluations help teams compare prompts, models, retrieval configurations, and application versions before releasing changes to production.
Conclusion
LLMOps is not a replacement for MLOps and DevOps. It extends the same engineering practices to address the additional challenges of running LLM applications in production.
As we learned in this guide, each open-source tool solves a different part of the problem.
vLLM handles efficient model inference, KServe helps manage model serving on Kubernetes, LiteLLM provides a common gateway for accessing models, Langfuse gives you observability and tracing, and Ragas helps evaluate the quality of your application.
However, you don't need all of these tools to build an LLM application.
So how to choose the right stack?
Well, it depends on whether you use managed APIs or self-hosted models, whether you run on Kubernetes, your scale, and the level of control and observability you need.
The important thing is to understand what problem each layer solves.
For DevOps and MLOps engineers, many of the underlying concepts are already familiar. Kubernetes, APIs, gateways, scaling, monitoring, security, and automation.
LLMOps builds on those foundations while adding new concerns such as model quality, token usage, prompt management, evaluation, and AI-specific observability
Over to you!
Which open-source tools are you using or planning to use in your LLMOps stack?
Let me know in the comments.