Product Updates
[ What's launched at LaunchDarkly ]
May 19, 2026
AI Engineering
A new Playground for offline evaluations
Compare prompt and model variations side by side, score them against a dataset of real inputs, and promote the winner.
Building a great agent is a cycle, not a one-time event. The new Playground and Offline Evaluations experience in AgentControl gives teams a dedicated offline environment for that cycle, where you can experiment freely, compare variations side by side, validate against real inputs, and promote a winning configuration when you are ready.
You start in the Playground with a single prompt and model. When you want to compare approaches, you click Compare to add columns and run multiple variations in parallel. When you are ready to validate at scale, you attach a dataset of representative inputs and run a bulk evaluation across every row. Dataset columns map to template variables in your prompt, so each row flows through as a real, contextual request.
Acceptance criteria powered by built-in LLM evaluation metrics score each row on dimensions like answer relevancy, faithfulness, and hallucination, with pass and fail results, scores, and per-row reasoning in the results table. Token cost is surfaced alongside quality so you can make informed tradeoffs.
When a configuration earns it, save it as a variation and bring it into your product. When you want to improve a live agent, click Open in Playground from any configuration to pull it back into the same environment and start the loop again.
Read the Playground and Offline Evaluations documentation to get started.
