An agent failed, so the team enables full prompt and tool-response logging. The next investigation becomes easier, but the logs now contain customer documents, access tokens, private instructions, and data the team did not need to retain.

Observability is the ability to explain system behavior from evidence. It does not require preserving every byte the system handled. The trace should answer a defined operating question under a deliberate data policy.

Start with outcome and event identity

Give the request a run identifier and connect model calls, tool calls, validations, approvals, and external effects. Record timestamps, component versions, operation names, status, and cost metadata where available. A trace without a final outcome may explain activity without explaining whether the user succeeded.

For a synthetic briefing agent, the event record can show that three sources were retrieved, one was excluded as expired, a summary was drafted, and publication remained pending approval. That is more useful than a complete transcript with no structured state.

Use conventions, but pin their version

The OpenTelemetry GenAI semantic-conventions project defines structures for generative calls, tools, agents, and related telemetry. Some retrieved conventions are explicitly in Development. Attribute names and recording recommendations can change.

Pin the instrumentation and schema version, and preserve that version with exported traces. Do not describe an evolving convention as a permanent interoperability guarantee. Test whether your backend can read old and new events during an upgrade.

Separate metadata from content

Tool name, duration, status, model identifier, and token usage may be sufficient for some performance investigations. Diagnosing a source-support failure may require a controlled sample of passages or outputs. Those are different recording needs.

Use content capture selectively with access rules, redaction, retention, and consent appropriate to the product. Avoid storing raw credentials or secrets. Redaction can itself miss fields, so test it on the actual payload shapes rather than assuming a generic filter is complete.

Measure retries as part of one request

If one user request produces five model calls, the request-level trace should retain the total cost and final status. Reporting only successful individual calls hides failed attempts. Cancellation, timeouts, and uncertain external effects need explicit events.

Distinguish attempted writes from completed writes and status reconciliation. A tool returning an error after accepting an operation should not be logged as proof that nothing happened. Keep the external operation identifier where it can be recorded safely.

Sampling changes what an investigation can conclude

Sampling reduces cost and exposure but can omit rare consequential failures. Decide which events are always retained, which content is sampled, and which critical outcomes trigger a restricted diagnostic record. State the sample policy when interpreting rates.

The NIST Generative AI Profile provides a risk-management framework rather than a single logging recipe. Use it to organize responsibilities and evidence, while designing the actual telemetry around the system's task and data.

Exercise the investigation path

Inject a missing source, a duplicate tool request, a revoked permission, and a delayed outcome in a safe environment. Ask an operator to find the cause using the proposed telemetry. Check whether they can distinguish source failure, model misuse, and integration failure without unnecessary access to personal content.

A trace is successful when it enables a proportionate response. Logging everything can create a new risk surface; recording the right evidence makes both the product and the investigation more accountable.

Sources and further reading