Agents · Tools
- model
- Qwen3-1.7B thinking off
- tools
- count_letter, tokenize
- call format
- <tool_call> JSON </tool_call>
- result
- 3
All steps
Why call a tool
Without tools, Qwen3-1.7B spells “strawberry” out and still answers 2, after 117 tokens; there are 3. It never sees letters: “ strawberry” is a single token in its vocabulary. Counting is a job for a few lines of code.
“ strawberry” = strawberry · answer 2, truth 3Tools as JSON schemas
The chat template writes each tool into the system prompt as a JSON schema: a name, a description and typed parameters, followed by the format to call them in. Two tools cost 241 tokens on every request.
2 tools · 241 tokens of system promptThe model writes a call
The model answers with a call instead of text: JSON between <tool_call> tags. Its first token is <tool_call> at 100%; at the name it puts 100% on “count”. The arguments are copied from the question.
count_letter(word="strawberry", letter="r")Check it, run it, return it
The program parses the JSON, checks it against the schema, runs the function and writes the result back as a <tool_response> in a new turn. The model does not run anything; it only reads what comes back.
checks pass · result 3The answer, and when not to call
With the result in its context, the model answers: “The letter "r" appears 3 times in "strawberry".”. The same tools do not force a call: asked “What is the capital of France?”, it puts <0.1% on <tool_call> and answers directly.
call when the tools help, answer when they do notConstrained decoding
A server can also force valid calls. At the name, only tokens that begin a declared tool name are allowed: 8 of 151,669. The rest get probability zero, as in a causal mask, and the allowed ones are renormalised.
8 of 151,669 tokens allowed at the name
Code
tools = [{'type': 'function', 'function': {'name': 'count_letter', 'parameters': {...}}}]
text = tokenizer.apply_chat_template(messages, tools=tools, add_generation_prompt=True, enable_thinking=False, tokenize=False)
out = generate(text) # '<tool_call>\n{"name": "count_letter", "arguments": {...}}\n</tool_call>'
call = json.loads(re.search(r'<tool_call>(.*?)</tool_call>', out, re.S).group(1))
jsonschema.validate(call['arguments'], schema) # a malformed call never reaches the function
result = FUNCTIONS[call['name']](**call['arguments'])
messages += [{'role': 'assistant', 'tool_calls': [{'function': call}]}, {'role': 'tool', 'content': str(result)}]