PaperPress
by Venom X Technology

OCR for Scanned PDFs & Photos — English, தமிழ் & සිංහල

To extract text from a scanned PDF or photo for free: drop the file below, choose the language (English, Tamil, Sinhala, or mixed), and click Read text. Recognition runs locally with Tesseract — one of the few online OCR tools that handles Tamil and Sinhala without uploading your document anywhere.

Drop a scanned PDF or photo here

or browse files — PDF, JPG, or PNG

First run downloads the language data (a few MB, then cached). Clear, high-contrast scans give the best results. OCR runs locally — your document still never leaves the device.

How to use this tool

  1. Drop a scanned PDF, JPG, or PNG into the tray.
  2. Select the document language — including Tamil, Sinhala, and mixed modes.
  3. Click Read text, then review, copy, or export the result as .txt or .docx.

OCR for Sri Lankan documents

Most free OCR sites support only Latin scripts, and the ones that handle Tamil or Sinhala require uploading your document to a server. PaperPress runs Tesseract — the open-source engine originally developed at HP and Google — directly in your browser via WebAssembly, with trained models for both languages. That makes it suitable for digitising old letters, government forms, certificates, and textbook pages without privacy concerns.

Frequently asked questions

Which languages are supported?

English, Tamil (தமிழ்), Sinhala (සිංහල), plus English+Tamil and English+Sinhala mixed modes.

Why does the first run take longer?

The language model (a few MB) downloads on first use and is then cached by your browser. The recognition itself always runs locally.

How do I get the best accuracy?

Use a clear, well-lit, high-contrast scan, keep the page straight, and pick the correct language mode.