PaddleOCR is an OCR toolkit for text and document-structure recognition across PDFs, images, multilingual OCR, and LLM-ready parsing.
Open Source: #ocr
Catalog projects marked with #ocr. Tags work as dedicated landing pages, so related tools are easier to find and connect.
This collection holds 6 projects with a combined 350,725 GitHub stars. Main languages: Python, C++, JavaScript.
Repositories
Tesseract OCR is an engine for recognizing text in images and scans, with support for many languages and OCR scenarios.
MinerU is a document-analysis tool that turns PDFs and office files into Markdown/JSON for search, RAG, and agent workflows.
Umi-OCR is a free offline OCR application for recognizing text in images and PDFs.
paperless-ngx is a document management system for scanning, OCR, indexing, and archiving files.
Tesseract.js is a JavaScript OCR library for recognizing text in images in the browser and Node.js.