The tool surface IS the agent's API: precise descriptions, strict schemas, errors that teach, and fewer tools than you think.
Writing the menu the agent orders from. The names, the descriptions and the error messages are all prompt — they decide whether it picks right.
Most 'the model is bad at tool use' complaints are actually badly designed tools, and this is the fastest fix available.
The model never sees your implementation. Name, description, input schema, return shape and error strings are the entire interface it reasons over, which makes tool design a prompt-engineering problem wearing an API-design costume. Four things decide whether an agent picks the right tool and calls it correctly: descriptions that state the one job plus the trigger and the anti-trigger; schemas strict enough that a hallucinated argument is rejected before anything side-effecting runs; error messages written for a model that will read them and retry; and restraint, because every additional overlapping tool measurably lowers selection accuracy and eats window space.
The model can't read my code, so the tool surface is the whole API: name, description, schema, returns and errors. I write each description as one job plus a trigger and an anti-trigger, because if two tools could plausibly answer the same sentence the model will guess. I make schemas strict — enums, regex patterns, required fields, ranges — so a hallucinated argument is rejected at the boundary before anything side-effecting runs. Errors say what was wrong, what's expected and what was received, which lets the agent self-correct on the next lap. And I keep the belt small: overlapping tools measurably lower selection accuracy, so I curate and treat tool-selection accuracy as a tracked metric.
What is MCP? Integrate AI Agents with Databases & APIs — IBM Technology, 3:46