davanstrien/bpl-ocr-bench-results
Rankings are computed using Bradley-Terry MLE from pairwise comparisons judged by a vision-language model. The judge sees the original document image alongside two anonymised OCR outputs and picks the more faithful transcription. Browse the comparisons to see the evidence — and vote yourself to build a Human ELO column. Human votes are stored locally for this session only and will reset when the server restarts.
| # | Model | Params | Judge ELO | 95% CI | Wins | Losses | Ties | Win% |
|---|---|---|---|---|---|---|---|---|
| 1 | LightOnOCR-2-1B | 1B | 1559 | 1497–1630 | 39 | 25 | 0 | 61% |
| 2 | GLM-OCR | 0.9B | 1535 | 1471–1591 | 48 | 35 | 1 | 57% |
| 3 | dots.ocr | 1.7B | 1453 | 1385–1515 | 26 | 37 | 0 | 41% |
| 4 | DeepSeek-OCR | 4B | 1452 | 1388–1514 | 33 | 49 | 1 | 40% |
Smaller models can win on the right documents. Error bars show 95% confidence intervals.