Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation
Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.