Explore/framework/JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation
J

Mufeng Yang, Junwei Yu, Yepeng Ding/JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent EvaluationUnknown

Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet a single judge is unreliable and even a panel of judges leaves a hard residue: when judges disagree, majority voting discards the conflict instead of resolving it. We present JuryFlow, a disagreement-guided, human-in-the-loop multi-agent evaluation framework that treats inter-judge disagreement not as noise to be averaged away, but as a precise, claim-level signal indicating where an evaluation is uncertain. JuryFlow decomposes each candidate response into atomic claims, has a panel of heterogeneous judges assign per-claim verdicts, and builds a disagreement graph whose nodes are scored by verdict entropy and whose edges encode structural similarity between claims. A human acts as a structural guide, selecting which disagreement to resolve through a single, minimal intervention rather than re-labeling the response, after which the focal claim is re-evaluated, the correction propagates along graph edges and to historically similar cases, and is crystallized into reusable rubric entries that all judges inherit, making the evaluator progressively self-refining. To enable large-scale, reproducible benchmarking without human studies, we evaluate JuryFlow in an automatic configuration in which focal selection is made by entropy ranking. On MT-Bench and LLMBar, JuryFlow improves agreement with gold labels over single-judge and majority-vote panel baselines, and ablations isolate the contributions of disagreement-targeted re-evaluation, propagation, and rubric induction. We contribute (1) a human-in-the-loop paradigm that recasts the human from labeler to structural guide, (2) the JuryFlow framework operationalizing it through a disagreement graph, focal re-evaluation, and closed-loop rubric induction, and (3) an evaluation protocol with ablations that isolate where the gains originate.

framework
GitHubCompare
Refreshed 1h ago
OverviewActivity52wAlternativesDocs
Stars0
Forks0
HF Downloads—30d
Last commit—
Refreshed1h ago
Project healthUnknownNo activity data.
Production readinessResearch / EarlyBest for exploration and prototyping.
Risk notesUnknown licenseVerify license before production use.
AgentHub Score
55 / 100
Composite score from 6 signals. How we score →
Active project
55Score
Growth
40C
Activity
30C
Documentation
70C+
Maturity
45C
Community
42C
Production
58C
GitHub stars · 0 days observed0 not enough history
snapshots
not enough history
Repository activity · 0 days observednot enough history from pushed_at
inactivepushed
not enough history
not enough history
Practical assessment
Should you use it?

✓ Best for

  • Multi-agent orchestration
  • Production agentic workflows
  • Stateful long-running tasks

◎ Strengths

  • Stable API
  • Active release cadence
  • Strong GitHub community

✕ Not ideal for

  • Simple single-step automation
  • Teams without Python/ML expertise

⚠ Watch-outs

  • Breaking changes between minor versions
  • Ecosystem lock-in if tightly coupled
Technical details
What's inside
Language—
License—
Sourcearxiv
Open source✗ No
Commercial use—
Docs—
Demo—

AgentHub Score

55
Score 55/100
Below average

Alternatives

O
OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Interactive Buildings
0 · framework
55
C
CHAL: Council of Hierarchical Agentic Language
0 · framework
55
T
Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents
0 · framework
55
S
Structuring agentic AI for HPC code modernization
0 · framework
55
Compare all →

Recent activity

Latest commit ——
Indexed by AgentHub crawler1h ago
Monitor for new releasesongoing