Set up LaunchDarkly observability plugin
Before you can manually record LLM spans, you must initialize your LaunchDarkly SDK with the observability plugin in your application. This enables your application to emit spans that LaunchDarkly can ingest and display in monitoring and trends views. All examples in the following sections assume this setup is already complete. For setup instructions for other SDKs and environments, read the Observability SDKs documentation. Complete the following steps to install and initialize LaunchDarkly observability for a Python application.-
Install the required LaunchDarkly SDK and observability package.
-
Initialize the LaunchDarkly SDK with the observability plugin enabled. This configuration is typically done once at application startup and enables span ingestion for your service.
Manually instrument LLM spans
Manual instrumentation involves creating a span around an LLM invocation and attaching structured attributes that describe the request and response. SDKs send these spans to LaunchDarkly observability and associate them with AgentControl monitoring and trends views. The examples below use Python. The attribute conventions shown are language agnostic, but the APIs and helpers used in these examples are specific to the Python SDK.Record an LLM span
Use this pattern to record a single LLM request as a span with basic model metadata, token usage, and one prompt and completion. Some attributes use thegen_ai namespace, while others use the llm namespace, as supported by LaunchDarkly observability. Token usage attributes include both the gen_ai and llm namespaces.
Token usage values should be read from the LLM provider response. Because token counts depend on provider-specific tokenization and system-injected content, they are not known ahead of time and should not be calculated by the application. Instead, record the values returned by the provider SDK, such as response.usage.prompt_tokens and response.usage.completion_tokens.
Wrap each model call in a span created with observe.start_span. The span lifecycle should align closely with the actual model invocation.
The following example records a single LLM request using this minimal attribute set. In practice, you typically start the span before invoking the model and populate attributes after the response is received so that span duration reflects model latency. Token usage values usually come from the model provider response.