Query Rewriting for Production RAG in October 2026: Multi-Query, HyDE, and Conversational Condensing Without Drifting From the User's Intent

Most RAG failures blamed on "bad retrieval" start one step earlier: the query itself. Users type "what about the EU one?", "pricing??", or a three-paragraph rant with the actual question buried in the middle. Your embedding model dutifully turns that text into a vector, the vector store returns the nearest chunks, and the answer is confidently wrong because the search never had a real question to work with.

Magnifying glass on a yellow background, a search concept photo
Image: Jernej Furman from Slovenia via Wikimedia Commons (CC BY 2.0)

Query rewriting is the layer that fixes this. Before you hit the index, a small model (or a few deterministic rules) turns the raw user message into one or more search queries that actually describe what the user wants. Done well, it lifts recall on vague and conversational questions. Done carelessly, it adds latency, burns tokens, and quietly changes the question so the system answers something the user never asked.

This guide walks through the main rewriting techniques, when each one earns its cost, and how to ship them in a production pipeline without drifting from the user's intent.

Where query rewriting sits in the pipeline

A typical production RAG request looks like this:

  1. Receive the user message plus conversation history.
  2. Rewrite: produce one or more standalone search queries, and optionally filters.
  3. Retrieve: run hybrid search (keyword plus vector) for each query.
  4. Fuse and rerank the combined candidates.
  5. Generate the answer from the top chunks, using the original user message as the question.

That last point matters. The rewritten query is a search tool, not a replacement for what the user said. Keep the original text for the generation step so the final answer responds to the real question, tone and all.

Technique 1: Conversational condensing

This is the single most valuable rewrite for chat-style products. Follow-up messages lean on context: "Does it support SSO?" only makes sense if you know "it" is the product discussed two turns ago. Embedding "Does it support SSO?" alone retrieves generic SSO content from every product you document.

Condensing asks a small model to turn the latest message plus recent history into one standalone question:

System: Rewrite the user's latest message as a single standalone
search query. Resolve pronouns and references using the conversation.
Do not answer the question. Do not add facts that are not in the
conversation. If the message is already standalone, return it unchanged.

History:
User: How does the Team plan handle audit logs?
Assistant: The Team plan keeps audit logs for 90 days...
User: Does it support SSO?

Output: Does the Team plan support SSO?

Practical rules that keep condensing safe:

  • Pass only the last few turns. Long histories make the model more likely to drag in stale topics. Three to five turns is usually enough.
  • Skip it on the first turn. If there is no history, there is nothing to resolve, so save the call.
  • Return unchanged when possible. Tell the model explicitly that "no change" is a valid output, and check for it in logs.
  • Never let it answer. A rewriter that starts writing answers will leak its guesses into retrieval.
Card catalog drawers at the Indiana State Library
Image: TBurmeister (WMF) via Wikimedia Commons (CC BY-SA 4.0)

Technique 2: Multi-query expansion

Some questions can be phrased many ways, and your documents may use different vocabulary than your users. A user asks about "getting a refund"; the policy page says "chargebacks and credit reversals." Multi-query expansion generates a handful of alternative phrasings, runs retrieval for each, and merges the results.

A reasonable starting configuration:

  • Generate three to four variants, not ten. More variants mean more retrieval calls and more near-duplicate chunks, with shrinking returns.
  • Ask for variants that differ in vocabulary, not in meaning: synonyms, the formal term, the product-specific term.
  • Always include the original (or condensed) query as one of the searches, so expansion can only add candidates, never remove the obvious ones.

Fusing results with reciprocal rank fusion

Each query returns its own ranked list. The simplest robust way to merge them is reciprocal rank fusion (RRF): every chunk gets a score of 1 / (k + rank) from each list it appears in, and the scores are summed. A constant around 60 is the common default for k. RRF only uses ranks, so you do not have to normalize similarity scores across queries.

def rrf(result_lists, k=60):
    scores = {}
    for results in result_lists:
        for rank, chunk_id in enumerate(results, start=1):
            scores[chunk_id] = scores.get(chunk_id, 0) + 1 / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)

After fusion, send the top candidates to your cross-encoder reranker using the condensed query. The reranker gets the final say on relevance; fusion just makes sure good candidates reach it.

Technique 3: HyDE (hypothetical document embeddings)

HyDE flips the problem around. Instead of embedding the question, you ask a model to write a short hypothetical passage that would answer it, then embed that passage and search with it. The intuition is that an answer-shaped paragraph sits closer in embedding space to real answer paragraphs than a terse question does.

HyDE can help when:

  • Questions are short and documents are long and formal, such as research papers, legal text, or technical specs.
  • Your embedding model was not trained with separate query and document modes.

It is risky when:

  • The model guesses wrong in a confident way. A hypothetical passage that invents a wrong product name or version number pulls retrieval toward the wrong documents.
  • Your corpus is private or niche. The model has no idea what your internal docs say, so its "hypothetical answer" reflects public knowledge instead.
  • Latency is tight. HyDE requires a full generation before retrieval can even start.

A sensible pattern is to run HyDE as one extra retrieval path alongside the plain query and fuse the results with RRF, instead of replacing the plain query entirely. Never show the hypothetical passage to users or pass it into the final prompt as if it were evidence.

Technique 4: Decomposition for multi-part questions

"Compare the retention limits and export formats for the Team and Enterprise plans" is really four lookups. A single embedding averages them into a blurry vector that matches none of them well. Decomposition splits the question into sub-queries, retrieves for each, and hands the generator a combined context.

Keep decomposition bounded. Cap the number of sub-queries (four to six is plenty for most products), and fall back to a single query if the model returns something that does not parse. Without a cap, a chatty rewriter can turn one question into a dozen retrieval calls.

Technique 5: Extracting structured filters

Some of the most useful rewriting is not about text at all. Questions often contain constraints that belong in metadata filters: "release notes from last month," "the Python SDK docs," "policies for Germany." Extracting those into structured fields lets you filter before similarity search instead of hoping the embedding picks them up.

{
  "query": "rate limit changes",
  "filters": {
    "doc_type": "release_notes",
    "published_after": "2026-09-01"
  }
}

Use structured outputs with a strict JSON schema so the filter object always parses, and validate every field against an allowlist of known values. A filter that does not match any real metadata value should be dropped, not passed through, or you will return zero results for a question you could have answered.

Keeping latency and cost under control

Every rewrite is an extra model call in the hot path, before retrieval starts. A few habits keep that cost honest:

  • Use a small, fast model. Rewriting is a narrow task; it rarely needs your flagship model.
  • Route by need. Skip condensing on first turns, skip expansion for queries that already contain exact identifiers like error codes or SKUs, and reserve HyDE and decomposition for question types where your evals show they help.
  • Combine steps. One call can return the condensed query, two variants, and filters in a single JSON object instead of three sequential calls.
  • Cache rewrites. Identical first-turn queries ("reset password") produce identical rewrites. An exact-match cache keyed on the normalized message avoids repeat calls.
  • Set a timeout with a fallback. If the rewriter is slow or errors, search with the original message. A degraded answer beats a hung request.

Guarding against intent drift

The quiet failure of query rewriting is drift: the rewritten query is fluent and reasonable but asks something slightly different. "Can I cancel without a fee?" becomes "How do I cancel my subscription?", and the retrieved chunks explain the cancel button while ignoring fees.

Ways to catch and limit drift:

  • Preserve key terms. Extract entities, numbers, and quoted phrases from the original message and check that they survive in at least one rewritten query. If they do not, add the original query back as a search.
  • Log original and rewritten side by side. Reviewing a sample of pairs each week surfaces drift patterns faster than any metric.
  • Answer from the original question. As noted above, the generator should see what the user typed, so even imperfect retrieval gets judged against the real question.
  • Prefer adding over replacing. Multi-query plus RRF with the original included is far more forgiving than swapping the query out.

Measuring whether rewriting helps

Do not ship rewriting because it sounds smart; ship it because your evaluation set says it helps. A practical process:

  1. Build a labeled set of real user queries, including follow-ups with their conversation history, each mapped to the chunks that should be retrieved.
  2. Measure retrieval recall at your context size (for example, recall@10) and mean reciprocal rank with no rewriting as the baseline.
  3. Turn on one technique at a time and rerun. Slice results by query type: first-turn, follow-up, multi-part, identifier lookups.
  4. Track the cost side too: added p50 and p95 latency and extra tokens per request.
  5. Keep a technique only for the slices where it clearly improves recall without hurting others, and route accordingly.

It is common to find that condensing helps follow-ups a lot, expansion helps vocabulary-mismatch queries, and HyDE helps on one corpus while hurting another. That is fine. The point of measuring is to turn each technique on exactly where it pays off.

A sensible default stack

If you are adding query rewriting to an existing RAG system this month, start here:

  1. Conversational condensing on every follow-up turn, with a small model and a strict "do not answer" prompt.
  2. Filter extraction with a JSON schema and allowlisted values.
  3. Two or three multi-query variants plus the original, fused with RRF, then reranked.
  4. A timeout that falls back to the raw message.
  5. HyDE and decomposition only after your eval set shows a specific slice that needs them.

Query rewriting is cheap to prototype and easy to overdo. Treat the rewritten query as a search aid, keep the user's words at the center of the answer, and let your evaluation numbers decide how much rewriting you actually need.

Comments

Popular posts from this blog

Structured Outputs and Constrained Decoding for Reliable LLM APIs in Production (September 2026)

Grok Bot - a step closer to AGI

GPT-Live-1 for Developers (September 2026): A Practical Guide to OpenAI’s Full-Duplex Voice API