Catalog

Explore agents, frameworks & tools

Filter by category, growth, license, capabilities — sort by what matters today. Indexed from GitHub and Hugging Face, refreshed every 30 minutes.

Quick filter:
Activebenchmark
31 results
AgentBench
StalePython
61

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

benchmarkOpen Source
3.7k
#1
WebArena
StalePython
61

Self-hosted realistic web environment for evaluating autonomous agents.

benchmarkOpen Source
1.6k
#2
ALFWorld
StalePython
55

Text and interactive environment for embodied/household agent tasks.

benchmarkOpen Source
847
#3
WebShop
StalePython
55

E-commerce environment with 1.18M products for grounded web interaction.

benchmarkOpen Source
587
#4
VisualWebArena
StalePython
55

Benchmark for multimodal web agents requiring visual understanding.

benchmarkOpen Source
485
#5
Deep-Research-Survey
Stale
50

A Systematic Survey of Deep Research

benchmark
323
#6
VisualAgentBench
StalePython
55

Towards Large Multimodal Models as Visual Foundation Agents

benchmarkOpen Source
275
#7
mcp-agent-trajectory-benchmark
Stale
55

A benchmark dataset for evaluating agent trajectory generation using the MCP (Multi-Context Planning) protocol, hosted on Hugging Face.

benchmarkHF
0
#8
RVMS-Bench
Stale
55

RVMS-Bench is a benchmark suite hosted on Hugging Face for evaluating the performance of multimodal models on a variety of vision‑language tasks.

benchmarkHF
0
#9
AgentWorldBench
Stale
55

AgentWorldBench is a benchmark for evaluating AI agents in a simulated world environment.

benchmarkHF
0
#10
ATBench
Stale
55

ATBench is a benchmark suite hosted on Hugging Face for evaluating AI agents across a range of tasks, providing standardized tasks and metrics for comparison.

benchmarkHF
0
#11
CityCube-Bench
Stale
55

CityCube-Bench is a benchmark dataset for evaluating AI models on tasks related to city cube data, intended for researchers and developers working on urban data analysis.

benchmarkHF
0
#12
Mobile-RobustBench
Stale
55

Mobile‑RobustBench is a benchmark suite for assessing the robustness of AI models on mobile devices, providing standardized tests and metrics for mobile‑friendly deployments.

benchmarkHF
0
#13
ZClawBench
Stale
55

ZClawBench is a benchmark dataset hosted on Hugging Face for evaluating AI agents’ performance on a variety of tasks.

benchmarkHF
0
#14
RVMSBench
Stale
55

RVMSBench is a benchmark dataset hosted on Hugging Face for evaluating the performance of vision-language models on a range of multimodal tasks.

benchmarkHF
0
#15
WorldCoder-Bench
Stale
55

WorldCoder-Bench is a benchmark for evaluating code generation models across multiple programming languages.

benchmarkHF
0
#16
EvoAgentBench
Stale
55

EvoAgentBench is a benchmark dataset hosted on Hugging Face for evaluating the performance of evolutionary AI agents. It provides standardized tasks and metrics to compare different agent designs.

benchmarkHF
0
#17
AgencyBench
Stale
55

AgencyBench is a benchmark dataset for evaluating the performance of AI agents on a variety of tasks, providing standardized tasks and metrics for comparison.

benchmarkHF
0
#18
RVMS-Bench
Stale
55

RVMS-Bench is a benchmark dataset hosted on Hugging Face for evaluating the performance of vision‑language models on a variety of multimodal tasks.

benchmarkHF
0
#19
VitaBench
Stale
55

VitaBench is a benchmark dataset hosted on Hugging Face for evaluating vision-language models, providing standardized tasks and metrics for performance comparison.

benchmarkHF
0
#20
PersonalizedDeepResearchBench
Stale
55

PersonalizedDeepResearchBench is a benchmark dataset hosted on Hugging Face for evaluating models that perform personalized deep research tasks, such as tailored literature search and recommendation.

benchmarkHF
0
#21
Deep-Research-Benchmarks
Stale
55

Deep-Research-Benchmarks is a collection of benchmark datasets for evaluating AI models on research-oriented tasks, hosted on Hugging Face.

benchmarkHF
0
#22
PR-Review-Bench
Stale
55

PR-Review-Bench is a benchmark dataset for evaluating AI agents on pull request review tasks, providing examples of code changes and review feedback.

benchmarkHF
0
#23
claude-agent-skills-benchmark
Stale
55

A benchmark dataset for evaluating the skill set of Claude agents, hosted on Hugging Face Hub.

benchmarkHF
0
#24