Explore/agent app/Herschel: Continuous Optimization of Production LLM Inference through On-Demand Profiling
H

Luping Wang, Weigao Chen, Yifei Wu, Yonghe Zhang, Rui Zhang, Wenchao Wu, Jiyu Luo, Haoran Geng, Xin Yang, Chen Cao, Yuemin Wu, Cheng Huang, Guodong Yang, Liping Zhang/Herschel: Continuous Optimization of Production LLM Inference through On-Demand ProfilingUnknown

Model-as-a-service platforms call for continuous optimization as complex serving conditions expose inefficiencies missed before deployment. Detailed always-on profiling can incur substantial overhead, while lightweight collection omits information needed for diagnosis. We present Herschel, a continuous optimization system for production large language model (LLM) inference. Our key insight is that adaptive, on-demand profiling can provide rich full-stack evidence without continuous collection. Herschel safely attaches to and detaches from selected running processes without engine changes or restarts, and adapts coverage as investigations reveal missing evidence. Herschel reconstructs operator executions and cross-process dependencies to identify inefficiency mechanisms and suggest solutions using applicable reference fixes. AI agents implement and test engine and kernel changes under controlled conditions that preserve the triggering workload and dependencies, with expert review before deployment. Controlled tests show active-collection overhead below 0.5% for time to first token and 7% for time per output token. Bounded windows, typically 30 s, avoid the continuous cost of always-on tracing. Over six months, Herschel collected approximately 17,000 traces across over 120 model variants and more than 10 accelerator types, identifying inefficiency patterns in 23% of the traces. Representative findings guide widely deployed optimizations, including restructured synchronization, removal of unused computation, and improved operator implementations.

agent app
GitHubCompare
Refreshed 1h ago
OverviewActivity52wAlternativesDocs
Stars0
Forks0
HF Downloads—30d
Last commit—
Refreshed1h ago
Project healthUnknownNo activity data.
Production readinessResearch / EarlyBest for exploration and prototyping.
Risk notesUnknown licenseVerify license before production use.
AgentHub Score
55 / 100
Composite score from 6 signals. How we score →
Active project
55Score
Growth
40C
Activity
30C
Documentation
70C+
Maturity
45C
Community
42C
Production
58C
GitHub stars · 0 days observed0 not enough history
snapshots
not enough history
Repository activity · 0 days observednot enough history from pushed_at
inactivepushed
not enough history
not enough history
Practical assessment
Should you use it?

✓ Best for

  • Research and experimentation
  • Prototype development
  • Learning agentic patterns

◎ Strengths

  • Active community
  • Open source
  • Well-documented API

✕ Not ideal for

  • Untested at scale without validation
  • Teams without AI/ML expertise

⚠ Watch-outs

  • Review changelog before updating
  • Verify license for commercial use
Technical details
What's inside
Language—
License—
Sourcearxiv
Open source✗ No
Commercial use—
Docs—
Demo—

AgentHub Score

55
Score 55/100
Below average

Alternatives

C
ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information
0 · agent app
55
A
Agent2UCB: Agentic System for Generative Engine Optimization
0 · agent app
55
A
Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments
0 · agent app
55
I
IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations
0 · agent app
55
Compare all →

Recent activity

Latest commit ——
Indexed by AgentHub crawler1h ago
Monitor for new releasesongoing