DARS

EMNLP 2026 · Main Conference

Distribution-Aware Routing Supervision

From Sampled Outcomes
to Capability Distributions

Rethinking Supervision for LLM Routing

Guannan LaiHaoran HuLong ChenZhenguo LiHan-Jia Ye

Nanjing University · HKUST · SinapisAI · Frontier Robotics

One sampled answer can change which model looks best.
DARS builds routing supervision from repeated observations of model capability.

40/42

utility comparisons improve

7

router architectures evaluated

3 × 6

reasoning tasks × candidate LLMs

131,328

released scored generations

01 / The supervision question

How much can one
answer tell us?

LLM responses vary with query wording and stochastic decoding. A single observation can change a model's score, the preferred routing label, and the policy learned from that label.

DARS combines semantics-preserving query rewrites with repeated decoding to estimate expected quality, expected cost, and performance variability for each query–model pair.

Those estimates provide risk-aware targets for existing routers. Sampling is part of offline supervision construction; the trained router selects a model from the original query at inference.

See the label-instability analysis ↗

02 / Method

Observe. Estimate. Route.

A common supervision source
across different router families.

Figure 3. Single-shot supervision records one response per model. DARS uses query rewrites and repeated decoding to estimate mean quality, mean cost, and quality variability, then constructs risk-aware routing targets.
Figure 3 from the paper. Repeated observations reveal a distribution of capability.
1

Collect observations

Five query rewrites × five decodes per rewrite, for each candidate model in the full training protocol.

2

Estimate capability

Summarize mean quality μq, mean cost μc, and quality standard deviation σq.

3

Build supervision

Use capability estimates as regression targets or convert risk-aware utilities into routing preferences.

Risk-aware utility

U(x, m) = μq(x, m) − λ μc(x, m) − β σq(x, m)

Balance expected quality, cost, and performance variability.
Paper defaults: λ = 0.05 and β = 0.2.

03 / Results

Better supervision.
Across router families.

Paper Table 1 ↗

DARS improves utility in 40 of 42 router–task–evaluation settings. Explore every comparison below. Values are reported in the paper, with single-shot supervision averaged over 100 independently sampled training sets.

GPQA · Table 1 reported routing utility (higher is better)
RouterRewrite: single-shot → DARSΔ utilityDecoding: single-shot → DARSΔ utility
MLP0.382 → 0.399+0.0170.400 → 0.413+0.013
kNN0.441 → 0.461+0.0200.469 → 0.489+0.020
EmbedLLM0.469 → 0.473+0.0040.496 → 0.480-0.016
AvengersPro0.464 → 0.480+0.0160.490 → 0.496+0.006
GraphRouter0.462 → 0.496+0.0340.487 → 0.516+0.029
MIRT0.409 → 0.444+0.0350.433 → 0.468+0.035
RM-Softmax0.453 → 0.477+0.0240.475 → 0.512+0.037

λ = 0.05. Rewrite and decoding utility evaluate two separate test views. Two settings decrease: EmbedLLM / GPQA decoding and MLP / MATH-500 decoding. Download all 42 comparisons as CSV.

Sampling can be moderate

The 3 × 3 protocol performs close to the full 5 × 5 observation grid in the paper's sample-efficiency analysis.

Figure 4 ↗

Single-shot labels can flip

GPQA has a 97.0% winner flip rate in the diagnostic analysis, illustrating how sampled outcomes alter model preferences.

Section 3.3 ↗

Quality and cost matter together

DARS generally improves cost–quality trade-offs across the representative routers studied on GPQA.

Figure 5 ↗

04 / Open benchmark

Start with the observations.

Hugging Face dataset ↗

600 training queries and 1,148 held-out test queries, evaluated with a heterogeneous pool of six LLMs. Download the precomputed scores to study routing without paid generation calls.

Released benchmark · one row is one scored generation
DatasetTaskTrain queries / rowsTest queries / rows
GPQAGraduate-level science QA200 / 30,000248 / 8,928
MATH-500Mathematical reasoning200 / 30,000300 / 10,800
DROP-800Discrete reasoning over text200 / 30,000600 / 21,600

Candidate models Gemma-3-12B-IT · Mistral-Small-3.2-24B-Instruct · Qwen3-32B · Llama-3.3-70B-Instruct · Gemini-2.5-Flash-Lite · DeepSeek-Chat-V3.1

Explore in Python

from datasets import load_dataset

data = load_dataset("AIGNLAI/DARS", "gpqa")
print(data["train"][0]["score"])

Install datasets first. Also available: math-500 and drop-800.

Reproduce routing experiments

Run the repository's download and router commands with the released JSONL files. Seven router implementations share the same supervision comparison protocol.

Follow the quick start ↗

Exact paper comparisons require matching the encoder, hyperparameters, seeds, and observation settings. TF-IDF is provided for an easy pipeline check.

05 / Citation

Build on DARS.

@inproceedings{lai2026dars,
  title     = {From Sampled Outcomes to Capability Distributions:
               Rethinking Supervision for {LLM} Routing},
  author    = {Lai, Guannan and Hu, Haoran and Chen, Long
               and Li, Zhenguo and Ye, Han-Jia},
  booktitle = {Proceedings of the 2026 Conference on Empirical
               Methods in Natural Language Processing},
  year      = {2026},
  url       = {https://arxiv.org/abs/2606.06924}
}

Accepted to EMNLP 2026 Main. The citation URL points to the arXiv version. Machine-readable citation ↗

Research scope

This study covers single-step routing with three English reasoning tasks and a fixed six-model pool. Repeated observations add offline collection cost. See the paper and dataset card for limitations, scoring conventions, and source-data terms.