Only RK

Price: 50000

Number of applications: 8

Decision acceptance deadline

29.09.26 (inclusive)

Form of award

Payment

Product status

Finished product

Task type

ICT tasks

Сфера применения

Robotics

Область задачи

Information processing and transformation

Type of product

Software/ IS

Problem description

What is it about In our service (verification of packages of documents for the state The text is extracted from the downloaded PDF for further analysis. Most files have a text layer and are parsed locally. But some of them come in scans without a text layer — they are now recognized by the cloud-based Google Gemini 2.5 Flash. We want to get away from the cloud and do it locally, on our hosting.

Expected effect

Increasing the independence of the product from third-party services

Full name of responsible person

Zhantas Asylbek

Purpose and description of task (project)

A task for the community: local OCR of PDF scans (RU/KK) instead of Gemini What is it about In our service (verification of packages of documents for the state text is extracted from the downloaded PDF files for further analysis. Most files have a text layer and are parsed locally. But some of them come in scans without a text layer — they are now recognized by the cloud-based Google Gemini 2.5 Flash. We want to get away from the cloud and do it locally, on our hosting. What needs to be done A module that accepts a PDF scan and returns the extracted text (plain text, UTF-8), verbatim and in the correct reading order. It should work offline, without external APIs. Important caveat on the type of documents This is not the recognition of passports, identification cards, complex forms or manuscripts. Scans are ordinary printed text documents: charters, certificates, contracts, business plans, letters. That is: mostly solid printed text, 1 column; machine font (not a manuscript); there may be simple tables, a stamp/signature at the bottom — but this is not the main content. This greatly simplifies the task: you don't need the "intelligence" of the multimodal LLM level for complex layout — you need reliable OCR of printed RU/KK text. Specificity Languages: Russian + Kazakh (ә ғ қ ң ө))))), sometimes in the same document. Quality: scans of different quality (distortion, noise, DPI 150-300). Critical: numbers, amounts, BIN/IIN (12 digits), dates — an error in a digit is unacceptable. Input: PDF up to 15 MB, 1-30 pages. Output: UTF-8 text, top—to-bottom / left-to-right order; tables - line by line. Integration: A PHP module or a local CLI/HTTP service called from PHP (CodeIgniter 4 / PHP 8.2 stack). Allowed stack Everything that is hosted: Tesseract 5 (rus + kaz LSTM), PaddleOCR / EasyOCR, docTR; rasterization of PDF→image (pdftoppm / Ghostscript); preprocessing (OpenCV / ImageMagick: deskew, denoise, binarize). Combinations are welcome. Acceptance criteria Benchmark: we will provide a set of ~150-200 real impersonal PDFs (RU/KK) with a reference text. Metrics: CER / WER vs. benchmark; Field-accuracy for critical fields (BIN, amounts, dates). The goal: accuracy is no worse than Gemini 2.5 Flash on the same set (we will publish its results as a baseline). Performance: less than N seconds per document (to specify the hardware), completely local. The format of the submission Repository with code + README (installation, dependencies, launch). A benchmark run script that outputs CER/WER and field-accuracy. A brief description of the pipeline. Deadline: 01.10.26

Note