Agents · Loops
- model
- Qwen3-1.7B thinking off
- tools
- lookup, calculate, finish
- stop string
- “\nObservation”
- this run
- 4 calls 658 tokens
All steps
The prompt sets up the loop
The system prompt teaches a format: a Thought, then one Action from a list of tools written in plain text, then an Observation. A worked example shows it once. That is 362 tokens before the question is even read.
362 prompt tokens · 3 actions · 1 worked exampleThink, then act
Qwen3-1.7B writes a Thought, then an Action. Naming the action is an ordinary next-token choice: after “Action 1:” it puts 100% on “lookup”.
Thought k: … → Action k: tool[argument]Stop, run the tool, append
The model never runs anything. The program around it watches the text, stops it where “Observation” would start, reads the action with a regular expression, runs the tool and appends the real result. Then it hands the context back.
stop → parse → run → append → continueRound and round
The real run: two lookups, one calculation, then finish[20]. Each observation came from the program, not the model. The model only decided what to look up and what to do with it.
4 model calls · 3 tool results · answer 20The context grows
Every step adds the model’s text and the tool’s result to the context, which is read again at the next call. Kept in a KV cache between calls, each token is processed once (658). Sent afresh each time, the calls process 2200.
658 tokens with a cache kept · 2200 re-sentWithout the stop
Without the stop, the model goes on and writes an Observation itself, a guess in the right format. That is why the loop must cut the text at “Observation” and insert the real result.
the stop string keeps results real
Code
while True:
out = llm.generate(context, stop=['\nObservation']) # a Thought and an Action, then stop
context += out
name, arg = re.search(r'Action \d+: (\w+)\[(.*)\]', out).groups()
if name == 'finish': return arg
result = TOOLS[name](arg) # the program runs the tool, not the model
context += f'\nObservation {k}: {result}\nThought {k + 1}:'