Leaderboard
Composite score of each system on the 88 test papers, all 30 fields weighted equally.
| Rank | Model | Design | Core | RAI | Composite | Paper | Cost / paper |
|---|---|---|---|---|---|---|---|
| 1 | Single-pass | 0.752 | 0.661 | 0.692 | 0.709 | $0.13 | |
| 2 | Single-pass | 0.676 | 0.682 | 0.680 | 0.699 | $0.27 | |
| 3 | Single-pass | 0.653 | 0.653 | 0.653 | 0.665 | $0.09 | |
| 4 | ReAct | 0.734 | 0.592 | 0.639 | 0.652 | $0.26 | |
| 5 | Parallel Specialists | 0.699 | 0.605 | 0.636 | 0.647 | $0.73 | |
| 6 | CroissantMiner team | Single-pass | 0.597 | 0.634 | 0.622 | $0.03 | |
| 7 | open | Single-pass | 0.698 | 0.584 | 0.622 | 0.634 | self-hosted |
| 8 | ReAct | 0.688 | 0.582 | 0.617 | 0.631 | $0.54 | |
| 9 | CroissantMiner team | Single-pass | 0.612 | 0.616 | 0.614 | $0.03 | |
| 10 | Triage + Critique | 0.675 | 0.578 | 0.610 | 0.624 | $0.21 | |
| 11 | open | Single-pass | 0.675 | 0.573 | 0.607 | 0.625 | $0.06 |
| 12 | Single-pass | 0.615 | 0.585 | 0.595 | 0.616 | $0.01 | |
| 13 | Single-pass | 0.561 | 0.589 | 0.580 | 0.596 | $0.03 | |
| 14 | Single-pass | 0.577 | 0.574 | 0.575 | 0.590 | $0.07 | |
| 15 | Parallel Specialists | 0.627 | 0.548 | 0.574 | 0.589 | $0.53 | |
| 16 | CroissantMiner team | Single-pass | 0.465 | 0.628 | 0.574 | $0.07 | |
| 17 | ReAct | 0.723 | 0.498 | 0.573 | 0.582 | $0.30 | |
| 18 | open | Single-pass | 0.617 | 0.525 | 0.555 | 0.575 | $0.01 |
| 19 | Locator-Extractor | 0.643 | 0.498 | 0.546 | 0.566 | $0.20 | |
| 20 | Triage + Critique | 0.592 | 0.522 | 0.545 | 0.557 | $0.16 | |
| 21 | Parallel Specialists | 0.607 | 0.477 | 0.520 | 0.539 | $0.57 | |
| 22 | open | Single-pass | 0.616 | 0.460 | 0.512 | 0.527 | self-hosted |
| 23 | Locator-Extractor | 0.523 | 0.455 | 0.478 | 0.502 | $0.11 | |
| 24 | Locator-Extractor | 0.560 | 0.413 | 0.462 | 0.478 | $0.08 | |
| 25 | Locator-Extractor | 0.500 | 0.413 | 0.442 | 0.448 | $0.09 | |
| 26 | Triage + Critique | 0.513 | 0.370 | 0.418 | 0.436 | $0.14 | |
| 27 | open | Single-pass | 0.506 | 0.310 | 0.375 | 0.391 | self-hosted |
Click a row to see the system's score on each of the 30 fields; the page address then links straight to it. Systems added after the paper show who ran them under their name.
Composite weights all 30 fields equally; point at it for its 95% interval. Core covers the 10 core fields, scored by rules; RAI the 20 Responsible AI fields, scored by GLM-5 at Z.AI (OpenRouter), which judged every row in Oct 2026. Paper is the composite in the paper, from the judge run of May 2026 (empty for systems added later). The panel of each system also shows its interval and when its outputs were made. Cost per paper at list prices of April and May 2026 for the paper's systems, and of the day the outputs were made for later ones.
* Built on a Claude model, which may have an advantage because Claude Sonnet 4.5 drafted the gold answers. Claude Sonnet 4.5 itself is not listed: it is scored against its own drafts, so its score is not comparable.