Agents · Tools

Tool Calling

schema → call → result

model
Qwen3-1.7B thinking off
tools
count_letter, tokenize
call format
<tool_call> JSON </tool_call>
result
3

All steps

  1. Why call a tool

    Without tools, Qwen3-1.7B spells “strawberry” out and still answers 2, after 117 tokens; there are 3. It never sees letters: “ strawberry” is a single token in its vocabulary. Counting is a job for a few lines of code.

    “ strawberry” = strawberry · answer 2, truth 3

  2. Tools as JSON schemas

    The chat template writes each tool into the system prompt as a JSON schema: a name, a description and typed parameters, followed by the format to call them in. Two tools cost 241 tokens on every request.

    2 tools · 241 tokens of system prompt

  3. The model writes a call

    The model answers with a call instead of text: JSON between <tool_call> tags. Its first token is <tool_call> at 100%; at the name it puts 100% on “count”. The arguments are copied from the question.

    count_letter(word="strawberry", letter="r")

  4. Check it, run it, return it

    The program parses the JSON, checks it against the schema, runs the function and writes the result back as a <tool_response> in a new turn. The model does not run anything; it only reads what comes back.

    checks pass · result 3

  5. The answer, and when not to call

    With the result in its context, the model answers: “The letter "r" appears 3 times in "strawberry".”. The same tools do not force a call: asked “What is the capital of France?”, it puts <0.1% on <tool_call> and answers directly.

    call when the tools help, answer when they do not

  6. Constrained decoding

    A server can also force valid calls. At the name, only tokens that begin a declared tool name are allowed: 8 of 151,669. The rest get probability zero, as in a causal mask, and the allowed ones are renormalised.

    8 of 151,669 tokens allowed at the name

Code

tools = [{'type': 'function', 'function': {'name': 'count_letter', 'parameters': {...}}}]
text = tokenizer.apply_chat_template(messages, tools=tools, add_generation_prompt=True, enable_thinking=False, tokenize=False)
out = generate(text)                    # '<tool_call>\n{"name": "count_letter", "arguments": {...}}\n</tool_call>'
call = json.loads(re.search(r'<tool_call>(.*?)</tool_call>', out, re.S).group(1))
jsonschema.validate(call['arguments'], schema)   # a malformed call never reaches the function
result = FUNCTIONS[call['name']](**call['arguments'])
messages += [{'role': 'assistant', 'tool_calls': [{'function': call}]}, {'role': 'tool', 'content': str(result)}]

Go deeper