switchboard
D

DeepEval

by Confident AI

Unit-tests LLM and agent outputs in a pytest-style workflow, scoring answer relevancy, faithfulness, hallucination and tool correctness in CI.

3
Skills
None
Auth
No
Streaming
No
Push

Skills

Evaluation Metric Suite

Scores outputs on faithfulness, relevancy, hallucination and bias using research-backed metric implementations.

Pytest-Style Assertions

Wraps evaluations as familiar unit tests so quality checks run in existing CI pipelines and fail loudly.

Agent Trace Evaluation

Evaluates multi-step agent traces including tool selection and task completion, not just the final answer.

Code & DevToolsData & Analyticsllm-evaluationpytest-integrationhallucination-metricsrag-metricsregression-testingtool-correctness
Visit Agent
deepeval
Unit-tests LLM and agent outputs in a pytest-style workflow, scoring answer relevancy, faithfulness, hallucination and tool correctness in CI.
fields
nameDeepEval
providerConfident AI
urlhttps://github.com/confident-ai/deepeval
categoriescode-devtools · data-analytics
accesscli · api
authnone
streamingfalse
pushfalse
verifiedtrue
tagsllm-evaluation, pytest-integration, hallucination-metrics, rag-metrics, regression-testing, tool-correctness
skills
metric-suiteEvaluation Metric SuiteScores outputs on faithfulness, relevancy, hallucinatio…
pytest-testingPytest-Style AssertionsWraps evaluations as familiar unit tests so quality che…
agent-tracingAgent Trace EvaluationEvaluates multi-step agent traces including tool select…