Opik
by Comet ML
Traces LLM and agent runs, scores them with LLM-as-a-judge and heuristic metrics, and monitors production quality via Python and TypeScript SDKs and 60+ integrations.
Skills
Agent Trace Logging
Captures trace trees of LLM calls, tool executions, and retrieval steps via a track decorator or framework integrations.
LLM As Judge Metrics
Scores outputs for hallucination, moderation, answer relevance, and context precision with built-in or custom judge metrics.
Datasets And Experiments
Evaluates applications against versioned datasets as experiments and compares the results in the Opik dashboard.
Production Monitoring Rules
Applies online evaluation rules to live traces and charts feedback scores, trace counts, and token usage over time.
PyTest CI Evaluation
Tests LLM pipelines on every commit through the PyTest integration so evaluation results are recorded per CI run.
Related Agents
OpenSandbox
Runs AI-agent workloads in isolated Docker or Kubernetes sandboxes, exposing sandbox lifecycle, command, filesystem, an…
Braintrust
Evaluates and observes LLM applications end to end: capture traces, score them with evals, iterate in a prompt playgrou…
AgentMail
Email inbox API built for AI agents. Create, send, receive, search, and manage email programmatically with SDKs for Pyt…
Claude MCP
Anthropic's Model Context Protocol — open standard for connecting AI models to tools, data sources, and services with u…