Explore/benchmark/CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
C

Chenmu Zhang, Levi Felix, Jun-Jie Zhang, Xingfu Li, Xuelian Jiang, Tao Jiang, Subhendu Mishra, Xixi Qin, Boris Yakobson/CompMat-Bench: Benchmarking AI Agents for Computational Materials ScienceUnknown

Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational materials research, repeating the same expensive simulations across agents and trials can make evaluation impractical. We introduce CompMat-Bench, a benchmark of 94 tasks derived from recently published computational materials studies, each asking agents to complete a step toward achieving the study's scientific goal. We reproduce the research steps in advance and assess agents on preparing inputs and analyzing outputs for expensive simulations, so expensive simulations can be avoided during evaluation. The reproduced inputs and results serve as ground truth for grading agents with fixed rules, without an LLM judge. The benchmark supports four evaluation conditions: single tasks and workflows composed of related tasks, each with full or reduced methodological guidance. With full guidance on single tasks, agents based on three LLMs demonstrate the ability to complete individual materials research steps, with pass rates of 66.0-90.4% across 94 tasks. Both longer workflows and reduced guidance can limit agent performance, but in different ways for different agents: they lower the pass rates of the weaker agents, whereas the strongest agent falls only when a long workflow is combined with reduced guidance. Failure analysis attributes most failures to scientific errors rather than to errors in software usage. CompMat-Bench provides a basis for comparing agents on the steps of real materials research and for analyzing agent failure modes.

benchmark
GitHubCompare
Refreshed 1h ago
OverviewActivity52wAlternativesDocs
Stars0
Forks0
HF Downloads—30d
Last commit—
Refreshed1h ago
Project healthUnknownNo activity data.
Production readinessResearch / EarlyBest for exploration and prototyping.
Risk notesUnknown licenseVerify license before production use.
AgentHub Score
55 / 100
Composite score from 6 signals. How we score →
Active project
55Score
Growth
40C
Activity
30C
Documentation
70C+
Maturity
45C
Community
42C
Production
58C
GitHub stars · 0 days observed0 not enough history
snapshots
not enough history
Repository activity · 0 days observednot enough history from pushed_at
inactivepushed
not enough history
not enough history
Practical assessment
Should you use it?

✓ Best for

  • Research and experimentation
  • Prototype development
  • Learning agentic patterns

◎ Strengths

  • Active community
  • Open source
  • Well-documented API

✕ Not ideal for

  • Untested at scale without validation
  • Teams without AI/ML expertise

⚠ Watch-outs

  • Review changelog before updating
  • Verify license for commercial use
Technical details
What's inside
Language—
License—
Sourcearxiv
Open source✗ No
Commercial use—
Docs—
Demo—

AgentHub Score

55
Score 55/100
Below average

Alternatives

R
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
0 · benchmark
55
C
Commitment To Cooperation With Self-Negotiated Contracts
0 · benchmark
55
E
Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?
0 · benchmark
55
C
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
0 · benchmark
55
Compare all →

Recent activity

Latest commit ——
Indexed by AgentHub crawler1h ago
Monitor for new releasesongoing