Leaderboard

davanstrien/bpl-ocr-bench-results

Rankings are computed using Bradley-Terry MLE from pairwise comparisons judged by a vision-language model. The judge sees the original document image alongside two anonymised OCR outputs and picks the more faithful transcription. Browse the comparisons to see the evidence — and vote yourself to build a Human ELO column. Human votes are stored locally for this session only and will reset when the server restarts.

# Model Params Judge ELO 95% CI Wins Losses Ties Win%
1 LightOnOCR-2-1B 1B 1559 1497–1630 39 25 0 61%
2 GLM-OCR 0.9B 1535 1471–1591 48 35 1 57%
3 dots.ocr 1.7B 1453 1385–1515 26 37 0 41%
4 DeepSeek-OCR 4B 1452 1388–1514 33 49 1 40%

ELO vs Parameter Count

Smaller models can win on the right documents. Error bars show 95% confidence intervals.