Filter by category, growth, license, capabilities — sort by what matters today. Indexed from GitHub and Hugging Face, refreshed every 30 minutes.
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Self-hosted realistic web environment for evaluating autonomous agents.