01 / Give the router a narrow job
A useful first experiment is deciding which part of an AI guide a question belongs to: models, agents, RAG, token costs or interview practice. Keep an other category for requests outside that list. Classifying a request should select a candidate workflow; it should not authorize every action that workflow can perform.
02 / Describe the boundaries
This is an illustrative task definition, not a live response or an executable provider request. Each label has a meaning. An overlapping request such as 'reduce the cost of my RAG agent' should be represented in the evaluation set instead of being silently forced into a convenient bucket.
{
"state": "How can I compare models for coding?",
"task": "Choose the main destination for this question",
"labels": {
"models": "Model capabilities, pricing and selection",
"agents": "Agent setup and tool workflows",
"rag": "Retrieval, documents and answer evidence",
"tokens": "Token budgets, caching and spending",
"interview": "AI and FDE interview preparation",
"other": "No clear match or more information needed"
}
}03 / Keep code in charge of the next step
Validate the response shape and selected label. Match only to a server-owned set of destinations or tools. Apply an uncertainty rule measured on labelled examples; if it fails, show ordinary search results or ask for clarification. Treat timeout, unavailable service and invalid output as separate fallback reasons.
Tool selection is only a recommendation. Before executing a tool, validate its arguments and permissions and keep the application's existing approval rules. A high probability is not an authorization token.
Request → Decision model → Validate label and uncertainty rule
Accepted → Allowed destination → Validate any tool action
Unknown / uncertain / failed → Search or clarification04 / Test ambiguous and out-of-scope requests
Label examples with one clear intent, multiple intents, missing context and no supported destination. Include typos, short messages and the languages your users actually write. Record route accuracy and fallback frequency together. This site currently provides this learning example; it does not call a decision model to route your searches.