Skip to content
Khalisa
All projects

Software Engineer Intern · ioNext

AI Gateway & LLM Serving Platform

Deep diving into envoy and its inference gateway to develop API Key authentication, billing, and observability for LLM serving.


Overview

The company has a platform for serving LLMs, and I worked on the inference gateway that sits in front of the models, handling request routing, authentication, and observability. The gateway is built on Envoy AI Gateway and Envoy Gateway, and the models are served by KServe and vLLM.

I was tasked with developing:

  • Observability for inference.
  • API Key authentication and billing for LLM inference requests.

The problem

Those tasks seemed simple at first, but they required deep diving into Envoy and its inference gateway, as well as understanding the model serving platform and its architecture. I had to learn about Envoy filters, how Envoy AI gateway was built, and the model serving stack, and how to integrate them with the existing platform.

Here are some of the challenges I faced in order:

  1. Understanding the company's existing LLM serving infrastructure

When I started the task, Envoy AI Gateway, Envoy Gateway, KServe, and vLLM were already in place and working, but I had to understand how they worked together and how to integrate my work with them.

It was rather confusing for the old me that the company has two Envoy named gateways, but they serve different purposes. The Envoy AI Gateway is a specialized gateway for AI inference, while the Envoy Gateway is a general-purpose API gateway.

I tried asking my colleagues, but they were busy fighting with their own task and deadlines, there was no company documentation on how the inference gateway works as it was really freshly implemented and still being tested, so I had to figure it out myself.

The documentation of Envoy AI Gateway was read by me, but I understood nothing. Knowing that doing so wouldn't get me anywhere, I decided to tinker with the gateways in the staging environment to understand how they work.

I spent a lot of time experimenting at 1AM so that I wouldn't disturb anyone working during the day, editing and reading the kubernetes manifests of all thing that smells like Envoy and AI.

Okay, I got an instinct about how it connects together, but I didn't truly know how it works internally.

  1. Learning about Envoy AI Gateway, which had limited documentation on customization.

I tried reading the documentation of Envoy AI Gateway again, and now it made much more sense.

There I found it! The observability documentation, which was the first task I was assigned to do. I read it and understood how to implement it.

Should be simple enough from now I thought...

Yes, it was simple to set it up. But my company requires a more advanced observability which was nowhere in the documentation! Asking ChatGPT and Claude was no help either.

I knew it in my heart that there will always be a way to do it, and I just need to find it.

So, I decided to read the source code of Envoy AI Gateway

  1. Deep diving into how Envoy AI Gateway was built on Envoy

With the help of Envoy's chatbot, I was able to understand its codebase, how Envoy AI Gateway was built on top of Envoy, and how it works internally. Turns out Envoy AI Gateway relies heavily on Envoy's external processing (ext_proc) filter to implement its inference processing.

ext_proc (external processing) is an Envoy filter that allows external processing of HTTP requests and responses. It can be used to implement custom logic between request and response. Envoy ext_proc flowEnvoy ext_proc service flow

Other than helping me solve my observability issue, this also enlightened me on how would I implement API Key authentication and billing, which was the second task I was assigned to do. It looks pretty straightforward, I just need to deploy and implement a new ext_proc service!

There's a reason why there are 5 challenges that I listed here...

I successfully implemented most of the observability features, however there was one thing I needed that the ext_proc didn't catch, which was telemetry during request failure before it reaches my ext_proc but after it reached Envoy AI Gateway's ext_proc.

So I needed to find a way on how to make my ext_proc service to be called before Envoy AI Gateway's ext_proc service.

However, after traversing the codebase of Envoy AI Gateway again, it explicitly states that its ext_proc service will always be called first.

I had some wild ideas, including forking Envoy AI Gateway and modifying its codebase, but before I reached to that point, I shall read more about Envoy filters first.

  1. Following through Envoy's request flow

After reading Envoy's documentation, there's another filter called ext_authz (external authorization) that is called before any ext_proc filter. Thankfully the data that I needed was a part of the request header. If it wasn't, I would have to find another way as the ext_authz filter only has access to the request header and not the body. The drawback of this approach is that there will be a slight increase in latency.

Finally, I was able to implement the observability features completely.

  1. Implementing API Key authentication and billing

After going through all that, I was able to implement API Key authentication and billing for LLM inference requests easily. I just needed to modify the external authentication service and ext_proc service to capture the AI response header and stream.

What I did

Inference Observability

  • Implemented request-level telemetry for inference traffic.
  • Integrated telemetry with the existing OpenTelemetry pipeline.
  • Captured relevant request metadata and response information.
  • Handled requests that failed before reaching the custom processing layer.

API Key Authentication

  • Implemented API-key verification within the existing authentication flow.
  • Propagated project/workspace identity through the inference pipeline.

Usage & Billing

  • Captured inference response information required for usage tracking.
  • Integrated request and response information for downstream billing.

What I learned

  1. If it isn't documented, it doesn't mean it can't be done. It just means we have to dig deeper and find a way to do it.
  2. If we don't understand something, read the source code. It is the most reliable source of truth.
  3. Writing notes is important, especially when we are working on a complex system. It helps us to remember what we did and why we did it.

Engineering Notes

Through the challenges i faced, I kept a log of my ideas, findings, problems, contemplations and solutions in a note so that i don't forget what I did. Snippet of notesSnippet of my notes, there are way more, but well NDA