CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

Berke Arda1*, Ahmetcan Yavuz1*, Paul Gerry1,3, Sebastian Lobentanzer4, Nobin Sarwar5, Joan Giner-Miguelez6, Kongtao Chen7, Luyao Zhang8, Mrinmaya Sachan1,2, Mubashara Akhtar1,2

NeurIPS 2026 · Evaluations and Datasets Track
Affiliations
1ETH Zurich 2ETH AI Center 3CSAIL, MIT 4Helmholtz Zentrum München 5University of Maryland, Baltimore County 6Barcelona Supercomputing Center 7Google 8Duke Kunshan University * Lead authors
102
dataset papers
30
metadata fields
3,060
gold answers
22
annotators

Leaderboard

Composite score of each system on the 88 test papers, all 30 fields weighted equally.

CSV: scores · field scores
RankModelDesignCoreRAICompositePaperCost / paper
1Single-pass0.7520.6610.6920.709$0.13
2Single-pass0.6760.6820.6800.699$0.27
3Single-pass0.6530.6530.6530.665$0.09
4ReAct0.7340.5920.6390.652$0.26
5Parallel Specialists0.6990.6050.6360.647$0.73
6CroissantMiner teamSingle-pass0.5970.6340.622$0.03
7openSingle-pass0.6980.5840.6220.634self-hosted
8ReAct0.6880.5820.6170.631$0.54
9CroissantMiner teamSingle-pass0.6120.6160.614$0.03
10Triage + Critique0.6750.5780.6100.624$0.21
11openSingle-pass0.6750.5730.6070.625$0.06
12Single-pass0.6150.5850.5950.616$0.01
13Single-pass0.5610.5890.5800.596$0.03
14Single-pass0.5770.5740.5750.590$0.07
15Parallel Specialists0.6270.5480.5740.589$0.53
16CroissantMiner teamSingle-pass0.4650.6280.574$0.07
17ReAct0.7230.4980.5730.582$0.30
18openSingle-pass0.6170.5250.5550.575$0.01
19Locator-Extractor0.6430.4980.5460.566$0.20
20Triage + Critique0.5920.5220.5450.557$0.16
21Parallel Specialists0.6070.4770.5200.539$0.57
22openSingle-pass0.6160.4600.5120.527self-hosted
23Locator-Extractor0.5230.4550.4780.502$0.11
24Locator-Extractor0.5600.4130.4620.478$0.08
25Locator-Extractor0.5000.4130.4420.448$0.09
26Triage + Critique0.5130.3700.4180.436$0.14
27openSingle-pass0.5060.3100.3750.391self-hosted

Click a row to see the system's score on each of the 30 fields; the page address then links straight to it. Systems added after the paper show who ran them under their name.

Composite weights all 30 fields equally; point at it for its 95% interval. Core covers the 10 core fields, scored by rules; RAI the 20 Responsible AI fields, scored by GLM-5 at Z.AI (OpenRouter), which judged every row in Oct 2026. Paper is the composite in the paper, from the judge run of May 2026 (empty for systems added later). The panel of each system also shows its interval and when its outputs were made. Cost per paper at list prices of April and May 2026 for the paper's systems, and of the day the outputs were made for later ones.

* Built on a Claude model, which may have an advantage because Claude Sonnet 4.5 drafted the gold answers. Claude Sonnet 4.5 itself is not listed: it is scored against its own drafts, so its score is not comparable.

Which fields are hard?

Average score of the 27 ranked systems on each field, hardest first. Point at a field to see what it asks for and the best system on it.

Core field (scored by rules)Responsible AI field (scored by the judge)
Imputation0.14
Missing data0.29
Manipulation0.32
License0.33
Personal and sensitive information0.35
Annotations per item0.40
Publisher0.42
Live dataset0.44
Citation0.45
Maintenance plan0.47
Biases0.48
Annotation analysis0.48
Annotator demographics0.54
Machine annotation tools0.56
Preprocessing0.58
Description0.60
Collection type0.60
Annotation platform0.62
Collection timeframe0.62
Social impact0.62
Language0.64
Limitations0.64
Annotation protocol0.66
Creator0.77
Raw data0.77
URL0.81
Date published0.83
Use cases0.85
Data collection0.86
Name0.90

Citation

@inproceedings{arda2026croissantminer,
  title     = {CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets},
  author    = {Arda, Berke and Yavuz, Ahmetcan and Gerry, Paul and Lobentanzer, Sebastian and
               Sarwar, Nobin and Giner-Miguelez, Joan and Chen, Kongtao and Zhang, Luyao and
               Sachan, Mrinmaya and Akhtar, Mubashara},
  booktitle = {Advances in Neural Information Processing Systems (Evaluations and Datasets Track)},
  year      = {2026}
}