Evaluation and observability
OpenTelemetry
Open standard and ecosystem for traces, metrics, and logs increasingly used in LLM app observability.
Advanced · Language-specific SDK packages (opentelemetry-sdk, @opentelemetry/sdk-node) with an OpenTelemetry Collector deployed as a sidecar or centralized gateway service
Editorial review
Tool categories, pricing, source status, deployment options, and product claims can change quickly. Verify the official source before production or commercial use.
About OpenTelemetry
Open standard and ecosystem for traces, metrics, and logs increasingly used in LLM app observability.
Best for: Engineering teams that want LLM application traces, metrics, and logs to flow into existing observability infrastructure — Datadog, Grafana, Jaeger, Honeycomb, or any OTLP-compatible backend — using the same instrumentation standard as the rest of the service stack.
Deployment: Language-specific SDK packages (opentelemetry-sdk, @opentelemetry/sdk-node) with an OpenTelemetry Collector deployed as a sidecar or centralized gateway service
Skill level: Advanced
Tradeoffs
OpenTelemetry is an instrumentation and export standard, not an LLM-specific evaluation or debugging product — a backend (Datadog, Honeycomb, Grafana, etc.) is required to visualize and query the exported data. LLM-specific features like prompt management, RAG evaluation, and dataset comparison require additional tooling like Phoenix or Langfuse layered on top.
Related guides and resources
Explore step-by-step setup guides, comparisons, and stack recipes for this tool category.
Best for
Engineering teams that want LLM application traces, metrics, and logs to flow into existing observability infrastructure — Datadog, Grafana, Jaeger, Honeycomb, or any OTLP-compatible backend — using the same instrumentation standard as the rest of the service stack.
Why use it
OpenTelemetry is the emerging standard for AI application observability through the GenAI and OpenLLMetry semantic conventions, which define how to represent LLM spans (model, prompt, completion, token counts, latency) in vendor-neutral OTLP format. Teams already operating OTel infrastructure can add LLM traces to existing dashboards and alerts without adopting a separate AI-specific platform.
Key features
- Vendor-neutral OTLP trace, metric, and log export to any OpenTelemetry-compatible backend via the OpenTelemetry Collector pipeline
- GenAI and OpenLLMetry semantic conventions for standardized LLM span attributes: model, prompt tokens, completion tokens, and latency
- Collector pipeline for trace sampling, batching, attribute transformation, and fan-out export to multiple observability backends simultaneously
- Language SDKs for Python, JavaScript/TypeScript, Go, Java, .NET, and Rust with automatic instrumentation for common web frameworks
Tradeoffs
OpenTelemetry is an instrumentation and export standard, not an LLM-specific evaluation or debugging product — a backend (Datadog, Honeycomb, Grafana, etc.) is required to visualize and query the exported data. LLM-specific features like prompt management, RAG evaluation, and dataset comparison require additional tooling like Phoenix or Langfuse layered on top.
Alternatives
- Langfuse
- Phoenix
- Ragas
Official verification sources
Direct official links used to verify pricing, features, security claims, and product packaging.
OpenSourcesAI ecosystem connections
Use these next-step links to move from this profile into related tools, comparisons, guides, stacks, and curated shortlists.
Alternative solutions
Guides, comparisons, and resources
Directory paths