Not a stuffing problem.
A routing problem. RECAST landed on arXiv October 7, 2026, filed under cs.AI as arXiv:2610.10507v1, formally titled “Learning to Compute the Right Context through Adaptive Evidence Routing,” and the whole argument compresses to one line: a small model should decide what evidence gets fetched, computed, or discarded before your expensive model ever reads the prompt. Adaptive evidence routing is the name for that decision layer. The transferable part is the split of duties. A lightweight RouterLM assembles evidence over multiple rounds. And a frozen AnswerLM generates exactly once, after the evidence passes inspection.
Steal the shape even if you never train a router of your own.
How Adaptive Evidence Routing Works Inside RECAST
The architecture reads like a build log.
No manifesto anywhere in it.
RouterLM runs the loop: it formulates evidence operations, either invoking a parameterized lexical, semantic, or relational primitive or requesting a custom operation that a frozen CompilerLM translates into executable code.
Then it looks at what came back.
The resulting evidence and execution outcome inform the next routing decision, which means the router refines its constructed evidence over multiple rounds instead of firing one retrieval call and hoping.
The stopping rule matters more. RouterLM assesses the outcomes of its operations and determines when the constructed evidence is sufficient. A condition you can check, not a feeling. Only after acceptance does the context pass to the frozen AnswerLM for final solution generation.
Training follows the current recipe. Supervised fine-tuning first. Group Relative Policy Optimization after that.
One line in the paper deserves more attention than it will get: RECAST explicitly treats retrieval and computation as complementary means of obtaining task-relevant evidence from heterogeneous sources. Evidence is not only things you find. It is also things you compute. Freight depots behave the same way. Nobody loads the entire warehouse onto a truck when the manifest says three pallets go to Denver.
Model Routing Grew Up First
Routing was already respectable before this paper, just at a different layer. GraphRouter builds a heterogeneous graph comprising task, query. And LLM nodes, with the interactions represented as edges. And its edge-prediction mechanism predicts the effect and cost of potential LLM responses. Which is why it adapts to newly introduced LLMs without retraining.
The code is on GitHub. BEST-Route pushes the same instinct toward thrift with a cost-efficient multi-headed router that dynamically assesses query difficulty to select the appropriate model and allocate computational resources.
Both of those decide where a request goes. RECAST decides what the request contains. The progression is easy to miss in the noise: route the model, then route the compute budget, then route the context itself. When all three layers route, the prompt your frontier model sees stops being a landfill and starts being an argument. Every component in this space now carries the -LM suffix, which is starting to read like a Pokémon evolution chart.
What Adaptive Evidence Routing Means for Small Builders
Three reads from the operator side, none of which require a research budget.
Cheap parts doing expensive judgment.
RouterLM is lightweight by design, and both CompilerLM and AnswerLM stay frozen. The pattern: spend small-model tokens on deciding, big-model tokens on answering.
Anyone billing by the token should recognize that division of labor as the entire pitch.
Computation as evidence kills a whole class of hallucination. A pipeline that only retrieves ends up answering arithmetic and join questions by paraphrasing whatever prose resembles the answer.
Let the router request code that computes the evidence and the answer model quotes a result instead of imitating a vibe.
The stopping rule is the governance. “Evidence is sufficient” is a checkable condition, which turns context selection into a decision process you can debug instead of a stuffing heuristic you can only feel.
Honest caveat, attached: this is an arXiv paper, not a product.
The mechanics are published and the packaging is not, so copy the shape and skip the waiting.
Metrics Before You Copy the Pattern
If you rebuild any of this, measure the routing itself, not just the final answers.
Evidence relevance. Of everything the router accepted, count how much the final answer actually used. Omission risk is the quiet one: how often “sufficient” fired while the answer still needed something the router never fetched. That failure mode wrecks reliability without announcing itself. Rounds to sufficiency matter as well.
Because every refinement round buys precision with seconds, so set a ceiling before your users set it for you.
Watch cost per answered question with the router included, since a lightweight router stops being lightweight once every answer needs a long negotiation.
Grade answer reliability on a fixed sample, never on the router’s confidence in itself.
That list separates actually adopting adaptive evidence routing from renaming your existing retrieval loop.
Numbers first. Aesthetics second.
Adaptive Evidence Routing FAQ
What is adaptive evidence routing?
A small model placed in charge of deciding what evidence to fetch, compute, or discard before a larger model sees the prompt. In RECAST that role belongs to RouterLM, working in rounds against a stopping condition, with a frozen AnswerLM waiting downstream.
How does RECAST differ from RAG?
RAG retrieves. RECAST retrieves and computes, treating the two as complementary means of obtaining task-relevant evidence from heterogeneous sources, and it loops.
The router inspects execution outcomes and refines its constructed evidence before anything reaches the answer model.
Do these routers retrain when new models appear?
GraphRouter does not. Its edge-prediction mechanism predicts the effect and cost of potential LLM responses, which lets it adapt to newly introduced LLMs without retraining.
Which metric fails first?
Omission risk. “Sufficient” firing while the answer still needs unfetched evidence quietly wrecks reliability. So instrument that before anything else.
Stop treating the context window as a budget you fill and treat it as a decision you make. Read the RECAST paper, poke the GraphRouter repo, then audit one real query in your own stack end to end this week. If nothing in your pipeline can say “stop, this context is enough,” you have found your next build. Ship it.
Sources
– RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing. ArXiv:2610.10507v1
– GraphRouter — arXiv:2410.03834
– GraphRouter code — GitHub
– BEST-Route — arXiv:2506.22716v1
