Products · 04 · Document Intelligence
← All products

Classifix

Turn unstructured documents into categorised, machine-readable data — without sending them anywhere.

Upload documents, read them with on-premise OCR in 129 languages, then classify and extract fields using language models that can run entirely on your own hardware. People verify the results in a keyboard-driven review station with the recognised words drawn straight onto the page.

129Bundled OCR languages
3OCR output formats
17Tracked job types
0Bytes that must leave
The problem

The documents you most need to process are the ones you cannot send away.

Cloud document-understanding services are excellent, but for many buyers they are simply not an option: the documents are patient records, case files, tender responses or personnel data, and the moment they cross a border or a tenancy boundary the project stops. So the work stays manual — sorting, keying, filing — at a cost nobody measures because it has always been that way.

01
Data cannot leave
Regulation, contract or policy rules out any service that processes documents off-site.
02
Mixed inbound batches
A scan run containing five document types, and someone paid to split it by hand.
03
Keying from paper
Fields retyped into a system that could have read them, with the errors that implies.
04
No feedback loop
Corrections vanish into a person's head instead of improving the next batch.
The solution

The entire pipeline runs on your hardware — including the model.

01
Local recognition
OCR runs in-process with 129 language packs already included. No download step, no per-page service, no outbound call.
02
Local intelligence
Nine model providers behind one configuration, five of which run on your own machines — classification and extraction with zero data egress is a setting, not a project.
03
People in the loop
A review station built for speed: the document on one side, the decision on the other, with the whole queue navigable from the keyboard.

Capabilities

01Optical character recognition
  • 129 bundled language packs, including vertical Chinese, Japanese and Korean, Fraktur and script detection.
  • Three outputs kept per document: plain text, word-positioned markup and an archival XML layer.
  • PDF pages rasterised at 300 dots per inch; images accepted directly.
  • Configurable resolution, language, page-segmentation mode and output format.
  • Thread pool sized to the machine, with per-page progress and cached results so re-running is a choice, not a cost.
  • Encrypted and zero-page PDFs are detected and rejected rather than silently producing nothing.
02Classification & extraction
  • Define your own classifiers and categories, with bulk category creation from a pasted list.
  • Classification by language model returns the category, a confidence score and its reasoning.
  • Per-page classification with automatic splitting: a mixed inbound batch is separated into one output file per category.
  • Category keywords — required, optional and prohibited terms — score a document and can re-file or flag it automatically.
  • Extractors defined per document type with four prompt slots — system, extraction, schema and refinement — plus an optional refinement pass.
  • Output as JSON or XML, stored with field count, confidence, method, model and processing time.
  • Single, sequential batch and parallel batch extraction endpoints.
03Models, local or hosted
  • Nine model providers — five that run on your own hardware and four hosted.
  • Named model profiles per organisation with full generation-parameter control and a connectivity test.
  • Reusable, versioned prompts with templated variables for both classification and extraction.
  • A model pass corrects recognition output before anything downstream reads it.
  • A regex path handles simple, fixed patterns where a model is unnecessary.
04Validation & evaluation
  • A three-panel review screen — queue, document, decision — with recognised word, line and block boxes overlaid on the page image.
  • Keyboard-driven navigation: next, previous, accept and page turning without touching the mouse.
  • Accept a classification or re-file it against the category list, one document at a time.
  • Queue statistics so a supervisor can see the backlog rather than guess at it.
  • A model bake-off runs named models over the same corpus and reports accuracy, average confidence and processing time.
  • Per-category precision, recall and F-measure, with test runs persisted as history for comparison.
05Operations & integration
  • Seventeen tracked job types covering recognition, training, classification, extraction, reclassification, restore and housekeeping.
  • Every job reports progress, items processed, estimated finish and a detailed log; jobs can be cancelled, re-run or bulk-cleared.
  • 236 REST endpoints across twenty-three controllers, described by an OpenAPI document.
  • Scoped API keys with a per-key permission map, expiry, revocation and per-day usage accounting.
  • Export or stream a whole classifier — files and data — to another Classifix instance, or copy it within one.
  • Guided first-run setup, including generation of a service unit.
■Single self-contained Java 21 / Spring Boot JAR with OCR engine and language packs included; PostgreSQL database.
■Server-rendered front end, no framework or build step; in-browser PDF viewer with word-position overlay.
■236 REST endpoints across 23 controllers, documented via OpenAPI.
■Nine model providers, five self-hosted; storage on local disk or a connected cloud folder.
■Roughly 139,000 lines of code across application and tests, 24 entities, 168 test classes.

What it does

  • 01Everything runs on your own hardware — OCR, classification and the language model — with zero required data egress.
  • 02129 bundled OCR languages and nine model providers, five of which run entirely on-premise.
  • 03A keyboard-driven review station overlays recognised words directly on the page, so reviewers verify rather than guess.

At a glance

For whomAccounts Payable · Legal · Public Sector · Document Bureaux
LicenceLicensed per installation, not per page. Classifix is bought once and run by you, priced as an annual fee on your own turnover (from €900 minimum, with yearly increases capped at 25%). Every tranche includes unlimited users, projects and organisations in one tenant, plus self-hosting and support. Hosting, hardware and AI model costs are your own; consultancy and training are quoted separately.

Classifix in your organisation?