Classifix
Turn unstructured documents into categorised, machine-readable data — without sending them anywhere.
Upload documents, read them with on-premise OCR in 129 languages, then classify and extract fields using language models that can run entirely on your own hardware. People verify the results in a keyboard-driven review station with the recognised words drawn straight onto the page.
The documents you most need to process are the ones you cannot send away.
Cloud document-understanding services are excellent, but for many buyers they are simply not an option: the documents are patient records, case files, tender responses or personnel data, and the moment they cross a border or a tenancy boundary the project stops. So the work stays manual — sorting, keying, filing — at a cost nobody measures because it has always been that way.
The entire pipeline runs on your hardware — including the model.
Capabilities
- 129 bundled language packs, including vertical Chinese, Japanese and Korean, Fraktur and script detection.
- Three outputs kept per document: plain text, word-positioned markup and an archival XML layer.
- PDF pages rasterised at 300 dots per inch; images accepted directly.
- Configurable resolution, language, page-segmentation mode and output format.
- Thread pool sized to the machine, with per-page progress and cached results so re-running is a choice, not a cost.
- Encrypted and zero-page PDFs are detected and rejected rather than silently producing nothing.
- Define your own classifiers and categories, with bulk category creation from a pasted list.
- Classification by language model returns the category, a confidence score and its reasoning.
- Per-page classification with automatic splitting: a mixed inbound batch is separated into one output file per category.
- Category keywords — required, optional and prohibited terms — score a document and can re-file or flag it automatically.
- Extractors defined per document type with four prompt slots — system, extraction, schema and refinement — plus an optional refinement pass.
- Output as JSON or XML, stored with field count, confidence, method, model and processing time.
- Single, sequential batch and parallel batch extraction endpoints.
- Nine model providers — five that run on your own hardware and four hosted.
- Named model profiles per organisation with full generation-parameter control and a connectivity test.
- Reusable, versioned prompts with templated variables for both classification and extraction.
- A model pass corrects recognition output before anything downstream reads it.
- A regex path handles simple, fixed patterns where a model is unnecessary.
- A three-panel review screen — queue, document, decision — with recognised word, line and block boxes overlaid on the page image.
- Keyboard-driven navigation: next, previous, accept and page turning without touching the mouse.
- Accept a classification or re-file it against the category list, one document at a time.
- Queue statistics so a supervisor can see the backlog rather than guess at it.
- A model bake-off runs named models over the same corpus and reports accuracy, average confidence and processing time.
- Per-category precision, recall and F-measure, with test runs persisted as history for comparison.
- Seventeen tracked job types covering recognition, training, classification, extraction, reclassification, restore and housekeeping.
- Every job reports progress, items processed, estimated finish and a detailed log; jobs can be cancelled, re-run or bulk-cleared.
- 236 REST endpoints across twenty-three controllers, described by an OpenAPI document.
- Scoped API keys with a per-key permission map, expiry, revocation and per-day usage accounting.
- Export or stream a whole classifier — files and data — to another Classifix instance, or copy it within one.
- Guided first-run setup, including generation of a service unit.
What it does
- 01Everything runs on your own hardware — OCR, classification and the language model — with zero required data egress.
- 02129 bundled OCR languages and nine model providers, five of which run entirely on-premise.
- 03A keyboard-driven review station overlays recognised words directly on the page, so reviewers verify rather than guess.