Terminal-Bench-Science
Evaluates agents on 70 expert-curated real scientific research workflows, run and graded through the Harbor benchmark harness.
Skills
Benchmark Run
Runs an agent against 70 curated scientific research tasks in reproducible containers through the Harbor CLI.
Task Grading
Grades agent transcripts against per-task rubrics and re-grades stored runs without re-executing the agent.
Result Comparison
Compares scores across models and harnesses on the same task set to track capability on real research workflows.
Related Agents
Crawl4AI
Crawls and scrapes web pages into clean Markdown for RAG pipelines, with CSS, XPath, and LLM-driven structured extracti…
Algolia
Hosted search and discovery API with typo tolerance, faceting, and vector search, callable directly for agent retrieval…
Docling
Parses PDF, DOCX, PPTX, XLSX, HTML, audio, and image files into a unified DoclingDocument and exports Markdown, HTML, D…
Firecrawl
Turn websites into LLM-ready data — crawl, scrape, and convert web pages to clean markdown for AI consumption.