SciExpert service

Document Extraction & OCR

Nougat, Tesseract and Camelot in the browser

Turn scanned pages and PDFs into structured Markdown, LaTeX and clean tables without uploading anything to a third-party service.

  • Layout-aware PDF to Markdown and LaTeX conversion
  • Tesseract 5 OCR across twelve languages
  • Stream-mode table detection with CSV export
  • Math and heading recovery for scanned articles
Open Document Extraction & OCR in workspace