All concepts
Agent Tool Calling
The model reasons, calls a tool, observes the result, and repeats until it can answer.
Agentic AI · Intermediate · ~8 min
In plain English
Give the model a menu of functions with clear names and descriptions. It picks one and fills in the arguments; your code runs it and hands back the result.
Why it's worth your time
It's the mechanism that turns a text predictor into something that can actually do things — and the tool description is the prompt.
If you remember three things
- The model chooses and fills arguments; it never executes anything
- Tool descriptions are prompt text and deserve prompt-level care
- Validate every argument before executing — always
Overview
Tool-calling turns an LLM into an agent. The model chooses a function and arguments; your code runs it and returns the result; the model incorporates the observation and decides the next action. The ReAct loop — reason, act, observe — repeats until the task is done.
How it works
- Task arrives The user's task goes to the LLM, which drives the whole loop.
- Reason (Thought) The model reasons about what to do next — the 'Thought' in the reason–act–observe loop.
- Choose a tool (Act) If it needs outside info, it emits a structured tool call with typed arguments — the 'Act'.
- Execute & observe Your runtime validates and runs the tool, then returns the result — the 'Observation'.
- Loop back to reason The observation feeds back to the model, which reasons again. The loop is bounded by a step limit to control cost.
- Final answer When no tool is needed, the model produces the final answer.
In an interview
Tool calling lets an LLM invoke functions: it emits a structured call with arguments, your runtime executes it, and the result is fed back so the model can reason and act again. This reason–act–observe (ReAct) loop lets agents use search, databases, and APIs instead of relying only on parametric memory.
Production defaults
- Count
- 5–8 tools. Past that, selection accuracy drops — split into sub-agents
- Naming
- verbs describing intent: search_orders, not GET /v2/orders
- Schema
- few required fields, enums for categoricals, descriptions on every parameter
- Errors
- return actionable text the model can use to retry differently
What breaks
- Picks the wrong tool on real traffic — Overlapping descriptions. Make each tool's 'when to use this' explicitly exclusive of the others.
- Passes invalid arguments — Validate and return a specific error naming the bad field. Never execute unvalidated arguments.