switchboard
T

Terminal-Bench-Science

by Stanford University

Evaluates agents on 70 expert-curated real scientific research workflows, run and graded through the Harbor benchmark harness.

3
Skills
None
Auth
No
Streaming
No
Push

Skills

Benchmark Run

Runs an agent against 70 curated scientific research tasks in reproducible containers through the Harbor CLI.

Task Grading

Grades agent transcripts against per-task rubrics and re-grades stored runs without re-executing the agent.

Result Comparison

Compares scores across models and harnesses on the same task set to track capability on real research workflows.

Research & Knowledgeagent-benchmarkscientific-workflowsharbor-frameworkagent-evaluationstanfordreproducible-grading
Visit Agent
terminal-bench-science
Evaluates agents on 70 expert-curated real scientific research workflows, run and graded through the Harbor benchmark harness.
fields
nameTerminal-Bench-Science
providerStanford University
urlhttps://github.com/harbor-framework/terminal-bench-science
categoriesresearch
accesscli
authnone
streamingfalse
pushfalse
verifiedtrue
tagsagent-benchmark, scientific-workflows, harbor-framework, agent-evaluation, stanford, reproducible-grading
skills
benchmark-runBenchmark RunRuns an agent against 70 curated scientific research ta…
task-gradingTask GradingGrades agent transcripts against per-task rubrics and r…
result-comparisonResult ComparisonCompares scores across models and harnesses on the same…