Trace, evaluate, and monitor LLM apps: nested spans, datasets, prod feedback
A place to see what your LLM app actually did: the prompts, the chains, the tool calls, the costs, and how a change scored on your test set.
Tracing and evals are the two things that turn prompt work from guessing into engineering.
LangSmith is the observability and evaluation platform for LLM apps, from the LangChain team but usable with or without LangChain. It captures every chain or agent run as a trace of nested spans with latency, token, and cost data, so you can see exactly what happened inside a request. It also provides datasets and evaluators for offline testing, a prompt hub for versioning prompts, and production monitoring with feedback capture.
LangSmith is observability and evals for LLM apps. Every request becomes a trace of nested spans showing latency, tokens, and cost, so you can debug what the model and tools actually did. Offline, you build datasets and run evaluators — LLM-as-judge, heuristics, or human — to score changes before shipping; online, you monitor production and collect feedback that loops back into new test cases. It works with any stack, not just LangChain.
What Is LangSmith? Explained in 5 Minutes — LangChain, 5:24