All concepts

Tool Design

The tool surface IS the agent's API: precise descriptions, strict schemas, errors that teach, and fewer tools than you think.

Agentic Engineering · Intermediate · ~6 min

In plain English

Writing the menu the agent orders from. The names, the descriptions and the error messages are all prompt — they decide whether it picks right.

Why it's worth your time

Most 'the model is bad at tool use' complaints are actually badly designed tools, and this is the fastest fix available.

If you remember three things

  • Name tools by intent, not by endpoint
  • The description IS the prompt
  • Errors should teach the next attempt

Overview

The model never sees your implementation. Name, description, input schema, return shape and error strings are the entire interface it reasons over, which makes tool design a prompt-engineering problem wearing an API-design costume. Four things decide whether an agent picks the right tool and calls it correctly: descriptions that state the one job plus the trigger and the anti-trigger; schemas strict enough that a hallucinated argument is rejected before anything side-effecting runs; error messages written for a model that will read them and retry; and restraint, because every additional overlapping tool measurably lowers selection accuracy and eats window space.

How it works

  1. A tool is a contract Name, description, schema, returns and errors are the whole interface. Write them for a new teammate.
  2. Naming and descriptions State the one job, the trigger, and the anti-trigger. Overlapping descriptions make the model coin-flip.
  3. Strict input schemas Enums, patterns, required fields, ranges — reject bad arguments before anything side-effecting runs.
  4. Errors that teach Say what was wrong, what's expected, and what was received, so the next lap self-corrects.
  5. Fewer, sharper tools Disjoint beats exhaustive. Measure selection accuracy; every extra tool competes for attention.

In an interview

The model can't read my code, so the tool surface is the whole API: name, description, schema, returns and errors. I write each description as one job plus a trigger and an anti-trigger, because if two tools could plausibly answer the same sentence the model will guess. I make schemas strict — enums, regex patterns, required fields, ranges — so a hallucinated argument is rejected at the boundary before anything side-effecting runs. Errors say what was wrong, what's expected and what was received, which lets the agent self-correct on the next lap. And I keep the belt small: overlapping tools measurably lower selection accuracy, so I curate and treat tool-selection accuracy as a tracked metric.

Production defaults

Count
5–8 per agent. Selection accuracy falls off past that
Names
verbs: search_orders, refund_order. Not GET /v2/orders
Descriptions
state when to use it AND when not to. Exclusivity is what prevents mis-selection
Schema
few required fields, enums for categoricals, examples in the description
Errors
'no order matched; try a wider date range' — actionable, not a stack trace

What breaks

  • Two tools get confused with each other — Overlapping descriptions. Make each say explicitly what the other is for.
  • The agent retries identically after an error — The error gave it nothing to change. Rewrite it as an instruction.

Watch it explained

What is MCP? Integrate AI Agents with Databases & APIs — IBM Technology, 3:46

Related