General-purpose OCR for images and scanned PDFs, on your own hardware. We're making it especially good at Arabic, dots and diacritics included.
- Input
- PNG, JPEG and scanned PDF
- Output
- UTF-8 text, plus JSON with regions and reading order
- Runs on
- Your own CPU, with no network
- Stage
- Pretrained models; no fine-tuning yet
See the measurements
Quarterly report 2026
Order no. ABC-123
{ "text": "Quarterly report 2026", "reading_order": 0 }Where pretrained models stand.
| Configuration | Strict CER | Median s/image |
|---|
| Tesseract fast, ara+eng | 52.80% | 0.199 |
|---|
| Tesseract best, ara+eng | 42.03% | 0.288 |
|---|
| PaddleOCR PP-OCRv5 + Arabic recognizer | 15.23% | 1.888 |
|---|
Arabic, 14 synthetic pages, CPU, 6 September 2026. Diagnostics, not accuracy on your documents.