Skip to content
Khalisa
All posts
1 min readDraft

Notes on observability for LLM inference

What to measure when you serve LLMs in production, and how OpenTelemetry, VictoriaMetrics, and Grafana fit together to show it.

By Khalisa Amanda Sifa GhaizaniObservabilityLLM

DRAFT: a starter outline based on my work at ioNext. Fill it in, then set draft: false. (Don't share anything confidential.)

Serving an LLM is not like serving a normal web API. Responses stream out token by token, requests vary hugely in cost, and GPUs are expensive to leave idle. That changes what you need to measure.

What to measure

The pipeline

Lessons learned