Filter by category, growth, license, capabilities — sort by what matters today. Indexed from GitHub and Hugging Face, refreshed every 30 minutes.
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Self-hosted realistic web environment for evaluating autonomous agents.
Text and interactive environment for embodied/household agent tasks.
E-commerce environment with 1.18M products for grounded web interaction.
Benchmark for multimodal web agents requiring visual understanding.
A Systematic Survey of Deep Research
Towards Large Multimodal Models as Visual Foundation Agents
A benchmark dataset for evaluating agent trajectory generation using the MCP (Multi-Context Planning) protocol, hosted on Hugging Face.
RVMS-Bench is a benchmark suite hosted on Hugging Face for evaluating the performance of multimodal models on a variety of vision‑language tasks.
AgentWorldBench is a benchmark for evaluating AI agents in a simulated world environment.
ATBench is a benchmark suite hosted on Hugging Face for evaluating AI agents across a range of tasks, providing standardized tasks and metrics for comparison.
CityCube-Bench is a benchmark dataset for evaluating AI models on tasks related to city cube data, intended for researchers and developers working on urban data analysis.
Mobile‑RobustBench is a benchmark suite for assessing the robustness of AI models on mobile devices, providing standardized tests and metrics for mobile‑friendly deployments.
ZClawBench is a benchmark dataset hosted on Hugging Face for evaluating AI agents’ performance on a variety of tasks.
RVMSBench is a benchmark dataset hosted on Hugging Face for evaluating the performance of vision-language models on a range of multimodal tasks.
WorldCoder-Bench is a benchmark for evaluating code generation models across multiple programming languages.
EvoAgentBench is a benchmark dataset hosted on Hugging Face for evaluating the performance of evolutionary AI agents. It provides standardized tasks and metrics to compare different agent designs.
AgencyBench is a benchmark dataset for evaluating the performance of AI agents on a variety of tasks, providing standardized tasks and metrics for comparison.
RVMS-Bench is a benchmark dataset hosted on Hugging Face for evaluating the performance of vision‑language models on a variety of multimodal tasks.
VitaBench is a benchmark dataset hosted on Hugging Face for evaluating vision-language models, providing standardized tasks and metrics for performance comparison.
PersonalizedDeepResearchBench is a benchmark dataset hosted on Hugging Face for evaluating models that perform personalized deep research tasks, such as tailored literature search and recommendation.
Deep-Research-Benchmarks is a collection of benchmark datasets for evaluating AI models on research-oriented tasks, hosted on Hugging Face.
PR-Review-Bench is a benchmark dataset for evaluating AI agents on pull request review tasks, providing examples of code changes and review feedback.
A benchmark dataset for evaluating the skill set of Claude agents, hosted on Hugging Face Hub.