Co-author · BINUS University research
Utility-Aware Inference Scheduler
Researching how to decide which LLM request gets served first, by weighing each request's priority, value, expected output length, and uncertainty. In trace-driven simulations it delivered 47.2% more total utility than first-come, first-served.
Overview
When an LLM server gets more requests than it can handle at once, something has to decide who goes first. That decision sounds small, but it changes a lot: a user waiting for a bug fix and a user saying "hi" to a chatbot end up in the same queue, fighting for the same GPU.
Together with my co-authors at BINUS University, I researched a utility-aware scheduler: instead of serving requests in arrival order, it estimates how valuable and how expensive each request is, and serves the ones that give the most value for the compute they use.
The problem
The common scheduling policies each optimise for only one thing:
- First-come, first-served (FCFS) is fair, but it treats an urgent debugging request the same as casual chat.
- Shortest-cost first keeps the queue moving, but long and valuable tasks (debugging, multi-step reasoning) keep getting pushed to the back.
So just sort by value and cost, right? Not quite. Here's what made it tricky:
-
We don't know the cost of a request until it's done
The cost of an LLM request is mostly how many tokens it generates, and that's only known after the model has finished answering. A scheduler has to decide before that. So we had to predict it from the prompt alone.
-
We don't know how valuable a request is either
Nobody labels their prompt with "this one is important". The value had to be inferred from the prompt itself: is it coding or debugging (high value), or casual conversation (lower)?
-
Predictions can be wrong, and a wrong guess is expensive
If the scheduler thinks a request will be short and it turns out to be very long, it blocks everyone behind it (head-of-line blocking). So we didn't only want a prediction, we wanted to know how confident that prediction was.
-
Testing it fairly
A scheduler can look great on made-up traffic. We needed real prompts and realistic traffic patterns, including the sudden bursts that real services get.
What we did
Predicting value
- Classified each prompt into one of six task types (from coding/debugging to casual chat) with a zero-shot classifier (
facebook/bart-large-mnli), and mapped each type to a value.
Predicting cost, and how unsure we are about it
- Trained a Random Forest on TF-IDF features of the prompt to predict the output-token length.
- Used how much the forest's trees disagree with each other as an uncertainty score, and turned it into a risk penalty.
Scheduling
- Combined priority, predicted value, predicted cost, and uncertainty into one risk-adjusted utility-to-cost score, and served the highest-scoring request first.
Evaluating
- Ran trace-driven simulations on 400 real conversations from the LMSYS-Chat-1M dataset.
- Compared against FCFS and shortest-cost scheduling under steady, bursty, and concurrent (Zipfian) workloads.
Results
- 47.2% more total utility than FCFS (49,295 vs. 33,495), and well ahead of shortest-cost scheduling (39,325).
- Most of that gain came from simply accounting for*request value.
And the uncertainty part? Honestly, under normal load it didn't make a statistically significant difference (p = 0.72). But under sudden traffic bursts, it gave a small, stabilising edge (48,290 vs. 47,310) by helping avoid head-of-line blocking. So it's not useless, it just earns its keep when things get chaotic.
Read the paper
Here's the full paper
Utility-Aware LLM Scheduling: Maximizing Compute Efficiency Under Priority Workload Bursts