Skip to main navigation Skip to search Skip to main content

Rerouting LLM Routers

Research output: Contribution to conferencePaperpeer-review

Abstract

LLM routers balance quality and cost of responding to queries by routing them to a cheaper or more expensive LLM depending on the query's estimated complexity. Routers are a type of what we call ``LLM control planes,'' i.e., systems that orchestrate multiple LLMs.

In this paper, we investigate adversarial robustness of LLM control planes using routers as a concrete example. We formulate LLM control-plane integrity as a distinct problem in AI safety, where the adversary's goal is to control the order or selection of LLMs employed to process users' queries. We then demonstrate that it is possible to generate query-independent ``gadget'' strings that, when added to any query, cause routers to send this query to a strong LLM. In contrast to conventional adversarial inputs, gadgets change the control flow but preserve or even improve the quality of outputs generated in response to adversarially modified queries.

We show that this attack is successful both in white-box and black-box settings against several open-source and commercial routers. We also show that perplexity-based defenses fail, and investigate alternatives.
Original languageEnglish
Number of pages33
StatePublished - 2025
EventSecond Conference on Language Modeling, COLM 2025 - Montreal, Canada
Duration: 7 Oct 2025 → …
Conference number: 2

Conference

ConferenceSecond Conference on Language Modeling, COLM 2025
Abbreviated titleCOLM 2025
Country/TerritoryCanada
CityMontreal
Period7/10/25 → …

Keywords

  • LLMs
  • Routers
  • Machine Learning

Fingerprint

Dive into the research topics of 'Rerouting LLM Routers'. Together they form a unique fingerprint.

Cite this