Agents · Retrieval
- passages
- 92 this site’s glossary and legend
- embedder
- all-MiniLM-L6-v2 384 dims
- retrieve
- top 3 by cosine
- model
- Qwen3-1.7B
All steps
The knowledge, in passages
The knowledge to draw on is this site’s own text: 85 glossary definitions and 7 lines of the legend, 92 passages. A model trained elsewhere has never read them.
92 passages · 85 glossary + 7 legendText becomes a vector
An embedding model reads each passage once and outputs one vector: all-MiniLM-L6-v2, a 6-layer BERT, averages its last layer over the tokens and scales the result to length 1. The question is embedded the same way.
text → 384 numbers, length 1Close in meaning, close in space
Passages about related things get similar vectors. Squeezed onto two dimensions (their top two principal components), related terms sit together and the question lands near the legend line that answers it. Two dimensions keep only part of the picture, so other close passages can look far.
2 of 384 dimensions, for the eye onlyRank by cosine similarity
Retrieval is a dot product: the question’s vector against every passage’s, highest first. The top 3 are kept. For the second question the answer is only at rank 2; a word overlap put “Few-shot prompting” first.
score = q · d (both length 1) · keep the top 3Passages into the prompt
The program pastes the top passages into the prompt above the question, with an instruction to answer from them only. The model sees retrieved text exactly as it sees anything else in its context.
211 tokens: instruction + 3 passages + questionWith and without retrieval
Without the passages, Qwen3-1.7B guesses what a hatched cell might mean anywhere. With them, it answers from this site’s legend. Retrieval did not change the model; it changed what was in front of it.
same model · with and without the passagesWhen retrieval misses
Retrieval fails quietly. “Which word came first” shares no words with the passages about position, and none comes back. A question can also find the right passage when that passage lacks the answer. Told to use only the context, the model says so, where on its own it made a name up.
wrong passages in, weak answer out
Code
embedder = SentenceTransformer("all-MiniLM-L6-v2")
P = embedder.encode(passages, normalize_embeddings=True) # once, ahead of time: (84, 384)
q = embedder.encode(question, normalize_embeddings=True) # (384,)
scores = P @ q # cosine similarity, since both have length 1
top = scores.argsort()[::-1][:3]
context = '\n'.join(f'[{i + 1}] {passages[j]}' for i, j in enumerate(top))
answer = llm(f'Answer using only the context below.\n\nContext:\n{context}\n\nQuestion: {question}')