B2B questions often require several facts. Adding every available document can also introduce obsolete specifications, adjacent models and irrelevant background. This guide concerns a controlled, self-built retrieval workflow, not undisclosed settings inside commercial AI search products.
Retrieval and generation have different budgets
Retrieving twenty candidates does not require sending twenty full documents to a generator. Filter candidates by entity, version, scope and provenance before allocating context. Instructions, the question, conversation history, tool wrappers and reserved output all consume capacity.
usable_evidence_budget = context_limit
- instructions - question - history - reserved_output
Measure every term with the relevant model tokenizer. Characters are not tokens. Source IDs and location records also cost space, but removing them can make an otherwise correct answer impossible to verify. When essential conditions do not fit, narrow the answer rather than silently dropping restrictions.
Inspect a reproducible truncation exercise
The context section of the execution record places one synthetic AX-220 fact first, in the middle or last among BX-110 distractors. Two additional variants represent faithful and misleading compression. The program keeps the first 160 characters and checks whether the model identifier and complete prohibition survive together.
| Variant | Changed variable | Observable check | Invalid inference |
|---|---|---|---|
| first | Evidence at the beginning | Restriction survives truncation | All models trust the beginning |
| middle | Four distractors precede evidence | Budget may run out first | Every model has identical position bias |
| last | Evidence follows all distractors | Tail evidence may disappear | Long-context products are useless |
| compressed_bad | Scope becomes everywhere | Meaning is broadened | Shorter is always better |
| compressed_good | Model, voltage and prohibition remain | Complete fact fits the window | Automated summarization is validated |
This is executed literal-retention logic, not LLM inference. The 160-character limit is a teaching setting, not a production recommendation. Change it in controls.py to observe boundary effects. The output is not a ChatGPT benchmark or a customer performance report.
Protect meaning during compression
“AX-220 uses 220 V indoors only. Do not use outdoors.” contains an entity, a supply value and operating restrictions. “The device supports 220 V” loses the entity. “AX-220 works indoors and outdoors” contradicts the source. Both may look plausible in a short summary.
Maintain a protected-field checklist: subject, attribute, value and unit, permitted conditions, exclusions and source version. Compare these fields after compression. Similarity alone is an inadequate acceptance rule because two statements can be semantically close while disagreeing about a safety-relevant restriction.
Resolve conflicts before combining sources
A newer document is not necessarily authoritative for the selected configuration. Establish the model, region, configuration and document purpose before interpreting dates. A recent general marketing page should not silently replace an older model-specific manual.
| Conflict | Defensible action | Shortcut to reject |
|---|---|---|
| Different configurations | Separate values and sources | Pick the more attractive value |
| Explicit replacement version | Use current version; preserve history | Delete the provenance chain |
| Unresolved disagreement | Show uncertainty and escalate | Use retrieval score as truth |
| Withdrawn evidence | Remove eligibility and review dependants | Continue publishing cached claims |
Design a generation experiment separately
Hold the question, model snapshot, instructions, sampling settings and answer key constant. Vary evidence count, order, repetition and compression independently. Retain complete inputs and outputs, citations, refusals and errors. Score supported claims and omitted restrictions; aggregate repeated answers at the question level.
Lost in the Middle found position sensitivity in its evaluated tasks and models. That motivates controlling position; it does not establish a universal arrangement for current products. Our downloadable exercise does not reproduce the paper or run a generation model. Keeping that distinction makes a subsequent experiment interpretable.
What this means for Zhihe Growth clients
Zhihe Growth organizes Chinese exporters' product facts so identifiers and conditions remain intact across languages and evidence fragments. A useful supplier review asks for discarded evidence, compression differences and conflict decisions, not only a fluent final answer. This improves the inspectability of source material without promising control over an external platform's context window.
Continue with document chunking, fact provenance and citation support. The technical FAQ provides short answers; the advanced lab package contains the executable exercise.