[Virtual event] Automate Your Software Factory - Nov 5 — Save my seat

Skip to:
BlogRight arrowAI Agents
Right arrow3 reasons teams can’t trust their AI agents with more

Oct 1, 2026

3 reasons teams can’t trust their AI agents with more

Here’s why the teams that are trying to build capable, trustworthy agents without adjusting their development process are struggling.

Kelvin Yap
Senior Product Marketing Manager

There are many limitations to how AI agents operate today. Maybe an agent is good at its job—say, drafting refunds—but even after a year of doing this, a human still has to review and approve the refund draft. This means that the reviewer has to spend every morning reading outputs, and sometimes they can’t get to everything. Or maybe an internal agent has been nearly ready for customers for a long time, but the team can’t quite get it into production.

Some of these guardrails exist on purpose. For instance, a human should make the final call whenever a mistake can’t be walked back. But a lot of this friction is unintentional. Teams want their agents to reach production and do as much there as safely possible, but every step in that direction requires sign-off—typically from someone on an entirely different team. And any change to an agent’s model, prompt, tools, and context changes its behavior, often in unpredictable ways. 

The teams that are trying to build capable, trustworthy agents without adjusting their development process are struggling because AI agents are fundamentally different from deterministic software. Here are three reasons the traditional approach to software development and delivery is holding AI agents back—and what teams can do about it.

Reason 1: Dashboards don’t measure output quality

When building AI agents, teams typically create dashboards that measure cost, latency, token counts, and failed tool calls. These metrics are easy to define and alert on, but none of them say whether the agent actually got it right.

The traditional software development process is built around tests that either pass or fail. This approach holds up when the output is code, fits a schema, or produces a number that can be checked downstream. But most agent output isn't like that. It's right or wrong by degrees, in ways that are specific to the agent's work and domain. An agent-written refund draft, for example, can comply with policy and quote the account history correctly but still be for the wrong amount.

This is due, in part, to the way agents are built and operate. A third-party model provider might ship a new version, or a tool might start returning slightly different data, and both of these shifts can cause the agent's output to get worse without anyone on the team changing anything.

When there’s no numerical way to measure quality, a human must review the output—and this approach doesn’t hold up at scale. Some failures also don’t surface at review time. For instance, the refund drafted for the wrong amount turns up when the chargeback lands a month later.

Teams that get past this do so by writing down what “good” means for their agent in terms a model can apply. They then use an LLM-as-a-judge to score output on live traffic against examples from their own production environment. This assigns a number to quality, which is something the person in charge of sign-off can easily look at. The human reviewer only has to read whatever the judge flagged, as well as a sample of what it passed. 

Reason 2: When quality drops, the fix must wait for sign-off

The next roadblock shows up after a decline in quality is detected—whether by an LLM-as-a-judge or a customer.

Before a person starts investigating, the only safety net is whatever the team set up in advance. This is usually a kill switch, a validator that catches obvious failure modes, and a queue for someone to pick up. None of these mechanisms can detect that quality has dropped, which means the path to an approved and deployed fix is slow. The signal arrives in real time, the change arrives whenever the process gets to it, and whatever happens in between happens in front of users.

This is why many teams only allow agents to perform actions that are cheap to get wrong. The agent that could theoretically issue a refund is only responsible for drafting it because a bad draft only costs a review, while a bad refund costs a customer. This is a fair rule for a production system, but it sets a ceiling on what the agent is worth to the business.

Some teams have addressed this challenge by deciding in advance what happens when quality dips below a certain threshold. For instance, they might revert to the last-good config, fall back to a simpler path, or hand off to a human. This happens automatically, within seconds. Whoever is on call then gets paged, but the fix doesn’t block on review and approval. Many of these teams also have policies in place that govern who is allowed to change an agent's configuration in production, and where that change gets written down.

Reason 3: Improving the agent costs more than leaving it alone

Most teams have every intention to improve their agent after launch. They might mean to write a tighter prompt, remove a step from the workflow, or use a newer model or a tool that returns cleaner data. But these plans often stall out because it’s very difficult to prove that any change hasn’t quietly broken something.

The traditional software development process puts every change through the entire cycle—from offline tests and review to a staged ship—which means something that took an afternoon to write can take weeks to trust. And this approach only confirms that the change addressed yesterday’s problem, since offline tests were built for an agent that isn’t running anymore.

These limitations remove the incentive for improving the agent, so as long as it’s good enough, it stays that way. Everyone knows it could probably be better, but proving it costs more than the outcome is likely to be worth.

The teams that are able to continuously improve their agents are the ones that have made the loop shorter than the agent’s own rate of change. Production traffic becomes the evaluation set, and is kept wherever the team keeps its tests. A new variation gets scored against it, then runs beside the current version on a slice of traffic with the same judge watching. That slice grows so long as the score holds. If the score drops, the variation is automatically pulled back out. Each change is small and reversible, which is what makes the next one easy to approve.

Evolving the process for more capable agents 

At many organizations today, engineers build agents while someone else decides how much autonomy they should be given. And that “someone else” is making decisions about a system that nobody can measure, correct while it’s running, or improve without a significant amount of work. The result is less trustworthy and less capable agents. 

Every team that wants to solve this problem must ask themselves the following questions:

  1. Have we written down what “good” means for this agent, and how do we know the output meets this standard?
  2. When the output stops being good, what happens before a human catches it?
  3. How quickly can we make the agent better, and how quickly can we reverse a change if it turns out to be bad?

Scoring every output against what “good” means, instantly moving to a fallback when a score drops, and keeping a record of every change are capabilities a team either has or doesn’t have. Most teams assembled their agents (and their process) with whatever tools and capabilities were close at hand at the time. The limitations of their agents show up where those capabilities run out. 

Progress starts when a team can show how the last change affected agent quality—and take it back quickly when the answer is bad.

See how LaunchDarkly AgentControl can help.

Like what you read?
Flaming pointer icon
See what's right for you.

Get started in minutes and scale seamlessly as your project grows.

Letter in envelope icon
Sign up for our newsletter

Get all the content, tips, and news you can use.

By supplying my contact information, I authorize LaunchDarkly to contact me with personalized marketing communications about our products and services. See our Privacy Policy for more details, or Opt-Out at any time.