A scanned or photographed document may look perfectly readable, but inside the file it is just an image: you can't search it or copy its words. OCR (optical character recognition) means letting the computer read the letters in the image and turn them into text.
Can OCR read Thai?
Yes. Tesseract, an open-source OCR engine, includes Thai (code "tha") in its language data. Our tool uses Tesseract.js, which runs in the browser, with the tessdata_best language data released under the Apache-2.0 licence.
Runs on your device, nothing uploaded
Thai OCR runs the recogniser as WebAssembly in your browser:
- The first time, your browser loads the engine and language data (about 0.9 MB for Thai and 3 MB for English) from our site.
- Your file is not sent anywhere, and it is cleared after you download.
- It suits private documents such as contracts, receipts or official letters.
What you get
- A searchable PDF: looks exactly like the original, with an invisible text layer on top, so you can search and select text.
- A .txt file: paste straight into Word, Google Docs or a chat.
- Or both, and you can copy the text from the page as soon as it's done.
Tips for better accuracy
- Sharp: good light, no blur, no hand shadows.
- Straight: if the photo is skewed, straighten the page first with Scan to PDF, then run OCR.
- Right language: Thai with some English → "Thai + English"; English only → "English".
- Not too small: we render pages at about 300 dpi before reading, but very small print or low-resolution originals can still be misread.
Limits
- Handwriting, decorative fonts and text on patterned backgrounds read poorly.
- Tables may come out with text out of column order.
- Always proofread, especially numbers and names.
If the result is too large, shrink it with Compress PDF.