All posts
1 min readDraft
Notes on observability for LLM inference
What to measure when you serve LLMs in production, and how OpenTelemetry, VictoriaMetrics, and Grafana fit together to show it.
By Khalisa Amanda Sifa GhaizaniObservabilityLLM
DRAFT: a starter outline based on my work at ioNext. Fill it in, then set
draft: false. (Don't share anything confidential.)
Serving an LLM is not like serving a normal web API. Responses stream out token by token, requests vary hugely in cost, and GPUs are expensive to leave idle. That changes what you need to measure.