[Virtual event] Automate Your Software Factory - Nov 5 — Save my seat

Skip to:
BlogRight arrowRuntime Control
Right arrowAI Model Management in Production: A Practical Workflow

Oct 3, 2026

AI Model Management in Production: A Practical Workflow

Manage AI models, prompts, evaluations, rollouts, and production performance with a practical AI model management workflow.

Scarlett Attensil
Scarlett Attensil
LaunchDarkly
AI model management in production: managing prompts, models, evaluations, targeted releases, guarded rollouts, and monitoring as one closed loop

Deploying an AI feature is a different challenge from deploying application code. When your application calls a third-party hosted model like GPT-6 Astra or Claude Fable 5.1 through a provider API, you don't own or operate the model itself. However, you do own every decision about which model gets called, with what prompt, for which user, and what happens when the output is wrong. Those decisions need to change frequently as models improve, prompts get refined, and edge cases surface in production.

Most software deployment tooling wasn't built for this. Shipping a new prompt or swapping a model shouldn't require a code deployment, and rolling back a misbehaving AI feature shouldn't take longer than the incident itself. Effective AI model management follows a consistent pattern: externalize model decisions from code, test variations before they ship, release them safely with guardrails, evaluate them continuously in production, and feed what you learn back into the next iteration.

This article walks through that pattern capability by capability, then maps it to how LaunchDarkly's AgentControl implements each step.

Summary of key AI model management capabilities

Platform capability

What it covers

Runtime model and prompt management

Externalizing model selection, prompts, and parameters from code so teams can update them at runtime without redeploying

Pre-production evaluation

Testing prompt and model variations against sample datasets with automated scoring before any production traffic is affected

Targeted release and segmentation

Serving specific model variations to specific user segments, environments, or contexts

Guarded rollout with automatic rollback

Progressively shifting traffic to new variations while monitoring metrics, with optional automatic rollback when monitored metrics regress

Production evaluation, monitoring, and Adaptive Triggers

Measuring live model output against quality criteria and operational metrics such as cost, latency, and token usage, with automated responses to configured metric thresholds

Experimentation

Measuring the impact of model variations on end-user behavior and business outcomes, alongside output quality scores

Continuous iteration

Feeding production findings back into pre-production evaluation, so real failure cases become the test set for the next round of variations

Runtime model and prompt management

The ability to update prompts, swap models, or adjust inference parameters without shipping new code is the foundation on which everything else depends. If changing a prompt requires a deployment, then testing variations, gradual rollout, and rollback become slow and risky by default.

The right abstraction is a single resource that bundles the model choice, prompt content, and parameters into a versioned configuration. Applications retrieve and apply that configuration at runtime via an SDK. Changes take effect immediately and can be reverted just as quickly.

This also changes who can make changes. Product managers and ML engineers shouldn't need to file a ticket or wait for a release cycle to iterate on a prompt. When model decisions live in configuration rather than code, anyone with the right access can update them without touching the deployment pipeline.

The versioning piece matters as much as the speed. When a prompt change causes a quality regression, the fix needs to be a configuration rollback, not a code revert that drags unrelated changes along with it. That requires model decisions to be versioned independently from application code, with a clear history of what changed, when, and in which environment.

Model registries and prompt stores solve adjacent pieces of this lifecycle. A model registry can track deployable model versions, while a prompt store can version prompt templates. AI model management also needs to connect those configuration changes to runtime targeting, evaluation, guarded rollout, and rollback for applications consuming hosted-model APIs.

AgentControl in LaunchDarkly externalizes model selection, prompts, and parameters into a managed resource, the AgentControl config, that applications retrieve through LaunchDarkly AI SDKs. Non-engineers can iterate on prompts and model settings without deploying, and changes take effect immediately across whichever environments and segments the team targets.

On the application side, the current Python AI SDK can evaluate an AgentControl config, select the configured provider, model, and prompt, invoke the model, and record metrics through a provider handler. The openai_messages convenience function wraps that flow in a single call.

The current Python AI SDK is in open beta; LaunchDarkly does not recommend beta SDKs for production use.

The LaunchDarkly Python AI SDK evaluates the summarization-v2 AgentControl config using the supplied user context. The openai_messages handler applies the model and messages selected by LaunchDarkly, invokes OpenAI, and records metrics as part of the call. The application supplies the input and context; LaunchDarkly controls the runtime config, and the provider handler owns the model invocation.

Pre-production evaluation

Before a new prompt or model variation reaches a real user, teams need a way to run it against representative inputs and measure whether the output meets quality standards. Strong pre-production evaluation catches many regressions early, when the cost of a fix is lowest. It catches them only to the extent that the dataset and scoring method cover the failure modes that matter, so some issues still surface only in production.

A pre-production evaluation environment needs three things:

  • A sandbox for experimenting with prompts, models, and parameters
  • Sample datasets that represent realistic production inputs
  • Automated scoring that uses a separate model as a judge to assess output quality against criteria like accuracy, relevance, or groundedness

The LLM-as-judge approach fills a gap that unit tests can't. Unit tests are good at deterministic checks like exact outputs, schemas, and golden-output comparisons, but they don't judge subjective quality well, like whether a response is accurate, relevant, or grounded. For language model output, quality exists on a spectrum that a traditional assertion can't capture. An evaluator model can score completions at scale against custom or built-in criteria, giving teams a way to compare variations before any production traffic is affected.

The quality of pre-production evaluation depends mostly on two design decisions: what goes into the dataset and how the judge is instructed.

A dataset that only contains typical, well-formed requests will pass almost any variation. Useful test sets mix three kinds of inputs: representative requests sampled from real traffic, edge cases that have caused failures before, and adversarial inputs such as prompt injection attempts, off-topic requests, and inputs in unexpected languages or formats. In practice, a test set is a plain file, one case per row, in the range of hundreds to a few thousand rows. Each row pairs an input with the expected output and any supporting context:

The variables field holds values that get substituted into the prompt template at evaluation time, so one test case can exercise different tones, personas, or locales without duplicating the row.

The test set should also grow over time. Each production failure captured as a test case closes a gap the original dataset missed.

The judge needs the same scrutiny. An evaluator model scores what the rubric tells it to score, so vague criteria like “the response should be helpful” produce noisy results. Effective rubrics define each criterion concretely, with a scoring scale and examples of passing and failing outputs. Judges also carry known biases: they tend to favor longer responses and outputs that resemble their own style. Calibrating judge scores against 50 to 100 hand-labeled responses, checking that judge scores track the human labels, then re-checking that calibration whenever the judge model changes, keeps automated scores meaningful.

The practical difference with an integrated platform is that evaluation criteria and datasets stay connected to the deployment configuration rather than living in a separate tool you have to sync manually.

The AgentControl Playground provides a sandbox for testing prompts and models with attached evaluation criteria. Datasets let teams build reusable test sets from real production inputs, uploaded as CSV or JSONL files smaller than 32 MB and with fewer than 10,000 rows. An evaluator LLM scores each completion against custom or built-in criteria so teams can compare variations before promoting them to production.

Evaluation results in the AgentControl Playground, with judge scores, token usage, and latency for each run (source: LaunchDarkly docs)

Targeted release and segmentation

A new model or prompt variation is rarely ready for every user at once. Take the ticket summarization feature from the earlier code snippet: an updated model might be right for internal support agents or paying customers before it’s ready for a free tier. A higher-cost model might be justified for enterprise accounts but not for trial users. Different regions may have latency or compliance requirements that make model choice non-uniform.

Targeting controls sit between the AgentControl config and the application, deciding which variation each request receives based on user attributes, account tier, region, environment, or any custom context the application provides. Expanding access from internal users to a beta cohort to general availability happens through targeting rule updates, not code changes.

This is also what makes targeting AI features different from feature flags on application code. With AI features, the variation being controlled is a bundle containing the model, prompt, skills to attach and parameters. The targeting layer controls that full bundle for each user, which is a different scope than gating a code path on or off.

When model configuration, audience targeting, and inference traffic are managed in separate systems, teams must keep their rules synchronized. Connecting targeting directly to the AgentControl config reduces the risk that the application retrieves a variation different from the one intended for that user or context.

AgentControl config targeting uses the same engine as the rest of the LaunchDarkly platform. Teams can target by individual context, segment, or custom attribute, and rules evaluate in the same order as feature flags: individual targets first, then custom rules, then the default rule, per environment. Targeting rules remain specific to each config and environment, while segments can be reused across configs and flags.

Guarded rollout with automatic rollback

A guarded rollout routes a small percentage of traffic to the new model variation while the rest stays on the current one. The platform monitors metrics like error rate, latency, and quality scores during the rollout. If monitored metrics remain healthy, traffic shifts incrementally. If LaunchDarkly detects a regression, the rollout can pause, and automatic rollback can restore the previous variation when that option is enabled.

This matters more for AI than for traditional features because model output quality can degrade in ways that error rates don't capture. Responses might be slower, more expensive, or lower quality without throwing a single exception. Guardrails need to watch AI-specific signals alongside operational ones.

Speed of rollback matters too. Redeploying during an incident is too slow. When rollback is a configuration change, recovery happens quickly without touching deployment pipelines or waiting for a CI run.

The threshold configuration is also worth thinking through carefully. A model variation that improves average quality scores might still degrade performance for a specific user segment or context type. Good guarded rollout implementations let teams attach multiple metrics with independent thresholds, so a regression in cost or latency can trigger a halt even when quality scores look fine. In practice, the thresholds read like an incident policy: halt the rollout if p95 latency exceeds 800 ms, cost per request rises more than 20%, or the quality score drops below 0.7. The numbers come from a baseline: measure current p95 latency and cost per request before the rollout, then set thresholds as tolerances on what you measured.

Inference-layer traffic splitting and application-layer release control solve different problems. Splitting requests between model endpoints does not by itself connect user targeting, AI quality signals, and rollback decisions. An AI model management workflow needs those signals associated with the configuration variation being served.

Guarded rollouts in LaunchDarkly progressively shift traffic to a new AgentControl config variation while monitoring attached metrics. Regression detection uses sequential testing: a regression is flagged when the confidence interval falls entirely on the worst side of a metric’s success criteria, and a rollout that does not serve the new variation to enough contexts by its end is rolled back because LaunchDarkly cannot reliably detect a regression. Sequential testing here means the check runs continuously as evaluations accumulate, rather than once at a fixed sample size. If a regression is detected, the platform can automatically roll back the change. Because rollback changes the served variation rather than redeploying application code, teams can restore the previous configuration without waiting for a new CI/CD deployment.

Production evaluation, monitoring, and Adaptive Triggers

Pre-production testing confirms that a variation works on known inputs. Production evaluation tests whether it works on the inputs your users actually send, which are messier and broader than any test set you can build in advance.

Production monitoring covers two layers. The first is quality scoring, e.g., an LLM as a judge running on a sample of live traffic, measuring output against criteria like accuracy, relevance, and toxicity. The second is operational metrics: cost, latency, token usage, and error rates, broken down by variation so teams can see how each model or prompt is performing relative to the others.

Both layers matter. Quality scoring catches regressions in output that operational metrics miss. Operational metrics catch issues like cost spikes or latency increases that quality scoring misses. Together, quality and operational metrics give a fuller picture of how a variation behaves in production. They are not the whole picture, since signals like user feedback and data drift matter too.

Running an LLM judge on production traffic adds another model request for each evaluated response, so sampling strategy matters. Operational metrics such as latency and token usage can be collected broadly, while judge-based quality evaluation can use a configurable sample. Teams can start with lower sampling for stable configurations and increase coverage for new or higher-risk variations, balancing evaluation cost against the amount of quality signal they need.

Attribution matters as much as collection. A dashboard that averages metrics across all variations will hide a regression that only affects one of them, so every metric needs a breakdown by variation and version. Monitoring that lives in a separate system also shows you what happened, but the remediation step still requires going elsewhere to update the model configuration or trigger a rollback. Keeping observation and action in the same place shortens that path.

Online evaluations in LaunchDarkly run an LLM as a judge on a configurable percentage of production traffic. Three built-in judges cover accuracy, relevance, and toxicity, and teams can define custom judges for domain-specific criteria.

The Monitoring tab breaks down cost, latency, token usage, and satisfaction metrics by variation (source: LaunchDarkly docs)

LaunchDarkly Adaptive Triggers can automatically switch an AgentControl config to a different variation when a monitored production metric crosses a configured threshold. For example, an Adaptive Trigger can fall back to another model or provider after an error spike.

AgentControl adaptive trigger templates.

Experimentation

Offline evaluation, online evaluation, and experimentation are related but distinct activities:

  • Offline evaluation scores variations against a fixed dataset before release.
  • Online evaluation scores live production responses with an LLM judge.
  • Experimentation goes one step further, measuring whether the model choice actually moves the metric the feature was built to improve, such as conversion rate, session depth, or task completion.

A model can produce higher-quality output by every evaluation criterion and still fail to improve user outcomes. Experimentation is what closes that gap, telling you whether a quality improvement in the model translates to a product improvement for users. An experimentation framework needs a way to assign users to variation groups, instrumentation for the business metrics being measured, statistical significance tracking, and tight integration with the runtime management layer so variation assignment and metric collection stay in sync.

The primary metric should be the business outcome the feature exists to improve, with quality scores and operational metrics tracked as guardrails rather than as the goal. Sample size math still applies: a lower-traffic AI feature may need a substantially longer experiment to detect a modest change in conversion, which is longer than most prompt iteration cycles. Teams handle that tension by reserving experiments for changes large enough to move business metrics, such as a model swap, and relying on evaluation scores for smaller prompt refinements. When an experiment has to run on thin traffic, variance reduction helps: adjusting results with pre-experiment user data, or measuring a more sensitive proxy metric like task completion instead of a rare outcome like conversion, cuts the runtime needed to detect a real lift.

The mechanics also punish loose integration. If the system assigning users to variations is separate from the AgentControl config layer, the same user can receive one variation while their metrics get attributed to another, which quietly corrupts results. Variation assignment and metric attribution need to share the same context identity.

LaunchDarkly Experimentation works directly on AgentControl configs, using the same variation assignment as targeted releases and feeding into statistical analysis with a choice of frequentist or Bayesian methods. Teams can measure model and prompt impact on end-user behavior alongside output quality, with no separate integration to maintain. Additionally, once a winner is determined, the same flag that allowed you to run the experiment gives you a one click path to rolling it out.

Continuous iteration

The previous six capabilities compound when findings flow between them. Production failures should become pre-production test cases. Successful experiments should inform the next round of variations.

Here’s the iteration loop in practice:

A closed-loop AI model management workflow connects production, capture, evaluation, and release.

What makes this harder than it sounds is that the data structures, judgment criteria, and metrics need to be shared across the lifecycle. A team that logs production responses in one schema and stores offline test cases in another ends up writing conversion scripts just to reuse a failure, and the loop quietly stops turning. The prevention is unglamorous: agree on one schema for inputs, outputs, and metadata at project start, and use it for both production logging and test case storage. If the metrics that trigger a rollback in production aren't the same ones being tracked in the evaluation sandbox, teams end up optimizing against different targets before and after release.

The value of a closed loop compounds over time. Each production failure that gets captured as a test case makes the pre-production dataset more representative of real user inputs. After several cycles, the evaluation environment reflects the actual distribution of what users send, which is what you want in order to catch regressions before they ship. A team using separate tools for evaluation and production monitoring has to move data manually to get there. An integrated platform can reduce the manual work required to connect production findings with subsequent evaluation.

In LaunchDarkly, the same judges can support both pre-production and online evaluation, while Datasets provide reusable test cases for offline evaluation. Findings from the Monitoring tab can inform new cases added to a Dataset and the next round of variations tested in the Playground. This connects what teams learn in production with what they test next.

Building a complete AI model management workflow

The argument of this article can be compressed into one loop. Externalizing model decisions from code enables fast iteration. Pre-production evaluation and targeted, guarded releases make new variations safe to ship. Production evaluation and experimentation tell you whether a change helped, and continuous iteration feeds what production taught you back into the next test set. Together, these seven capabilities form a complete AI model management lifecycle. The integration points that matter most in practice are:

  • Whether configuration changes take effect at runtime or require a redeploy, and how the platform propagates them, for example, through streaming versus polling
  • Whether datasets, judges, rubrics, versions, and scores are traceable across pre-production and production
  • Whether experimentation and targeting share the same variation assignment logic or require manual synchronization

The tighter those integrations, the less time teams spend reconciling data between tools, and the faster the feedback loop between production findings and the next iteration. When assessing an implementation, verify these integrations on a live feature rather than relying only on an architecture diagram.

For teams building on hosted models, AgentControl manages model and prompt configuration, offline and online evaluation, and AI performance data, while LaunchDarkly's targeting, release, and Experimentation capabilities control how config variations reach users and measure their impact.

Like what you read?
Flaming pointer icon
See what's right for you.

Get started in minutes and scale seamlessly as your project grows.

Letter in envelope icon
Sign up for our newsletter

Get all the content, tips, and news you can use.

By supplying my contact information, I authorize LaunchDarkly to contact me with personalized marketing communications about our products and services. See our Privacy Policy for more details, or Opt-Out at any time.