Explore/benchmark/Reconstructing Implicit Scientific Knowledge: Evaluating LLM Agents through End-to-End Reproduction of Astronomy
R

Yuehui Wang, Xinyu Qi, Guirong Xue, Cheng Wang, Yangbin Xie, Xiaoyu Tang, Cong Sun/Reconstructing Implicit Scientific Knowledge: Evaluating LLM Agents through End-to-End Reproduction of AstronomyUnknown

The integration of large language models (LLMs) into scientific workflows is accelerating, yet their ability to reconstruct the reasoning underlying published research remains unexplored. Papers specify explicit procedures while leaving many methodological dependencies-data selection, calibration corrections, priors, and domain assumptions-implicit. This ambiguity complicates the evaluation of LLM-based agents, since a failure to reproduce a result may reflect either limitations of the agent or underspecification in the source. We present a framework that evaluates agents through end-to-end reproduction, separating execution from verification and computational failure from methodological ambiguity. We apply it to fourteen astronomy studies: a case study from The Astrophysical Journal and thirteen papers published in Nature. Eleven of the thirteen contained an ambiguity preventing a uniquely specified reproduction path. In a controlled case study, twelve predefined paths, a 3x2x2 sensitivity analysis over sample definition, sky masking, and parallax zero-point treatment-gave estimates from 2.16 to 3.53 kpc for the same quantity, with only one recovering the published value (about 2.70 kpc). The published value was never used as an optimization target, selection criterion, or stopping condition; the matching path was found only after all twelve had run. Crucially, the decisive information (a +0.02 mas parallax zero-point correction) was already in the paper, but the agents did not recognize its causal relevance until the analysis made the effect visible. Matching a published outcome therefore does not validate reconstruction of the underlying reasoning, and the bottleneck is as often a failure to connect relevant information as to retrieve it. End-to-end reproduction thus serves both as a test of reproducibility and as a framework for evaluating implicit scientific knowledge in AI systems.

benchmark
GitHubCompare
Refreshed 23h ago
OverviewActivity52wAlternativesDocs
Stars0
Forks0
HF Downloads—30d
Last commit—
Refreshed23h ago
Project healthUnknownNo activity data.
Production readinessResearch / EarlyBest for exploration and prototyping.
Risk notesUnknown licenseVerify license before production use.
AgentHub Score
55 / 100
Composite score from 6 signals. How we score →
Active project
55Score
Growth
40C
Activity
30C
Documentation
70C+
Maturity
45C
Community
42C
Production
58C
GitHub stars · 0 days observed0 not enough history
snapshots
not enough history
Repository activity · 0 days observednot enough history from pushed_at
inactivepushed
not enough history
not enough history
Practical assessment
Should you use it?

✓ Best for

  • Research and experimentation
  • Prototype development
  • Learning agentic patterns

◎ Strengths

  • Active community
  • Open source
  • Well-documented API

✕ Not ideal for

  • Untested at scale without validation
  • Teams without AI/ML expertise

⚠ Watch-outs

  • Review changelog before updating
  • Verify license for commercial use
Technical details
What's inside
Language—
License—
Sourcearxiv
Open source✗ No
Commercial use—
Docs—
Demo—

AgentHub Score

55
Score 55/100
Below average

Alternatives

R
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
0 · benchmark
55
C
Commitment To Cooperation With Self-Negotiated Contracts
0 · benchmark
55
E
Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?
0 · benchmark
55
C
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
0 · benchmark
55
Compare all →

Recent activity

Latest commit ——
Indexed by AgentHub crawler23h ago
Monitor for new releasesongoing