DeepEval
by Confident AI
Unit-tests LLM and agent outputs in a pytest-style workflow, scoring answer relevancy, faithfulness, hallucination and tool correctness in CI.
Skills
Evaluation Metric Suite
Scores outputs on faithfulness, relevancy, hallucination and bias using research-backed metric implementations.
Pytest-Style Assertions
Wraps evaluations as familiar unit tests so quality checks run in existing CI pipelines and fail loudly.
Agent Trace Evaluation
Evaluates multi-step agent traces including tool selection and task completion, not just the final answer.
Related Agents
Crawl4AI
Crawls and scrapes web pages into clean Markdown for RAG pipelines, with CSS, XPath, and LLM-driven structured extracti…
OpenSandbox
Runs AI-agent workloads in isolated Docker or Kubernetes sandboxes, exposing sandbox lifecycle, command, filesystem, an…
promptfoo
Tests and red-teams LLM apps, agents and RAG pipelines from declarative config, scanning for prompt injection, jailbrea…
Opik
Traces LLM and agent runs, scores them with LLM-as-a-judge and heuristic metrics, and monitors production quality via P…