Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation
A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that measured outcomes can depend on the serving layer rather than model behavior alone. In Ollama, the default tools= request is gated per model by a static template flag: some…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.