01 / Document AI

Insureflow

I built fully local insurance coverage extraction with PDF review. Rewriting the core with Rust and Python reduced RAM usage by roughly 70%.

Deployment constraint100% local
STACKRust / Python / XGBoost / TATR / CamemBERT
HITL inspection: extracted data and highlighted PDF cells

01

The problem

Health insurance coverage tables do not follow a single format. Providers change layouts, terminology and how options are presented. Text extraction alone is not enough: the system must recover structure, understand domain terms and let an operator verify every result.

02

My role

I designed the vision / NLP pipeline and built the Human-in-the-Loop (HITL) review interface. I then rewrote the Python MVP’s application core in Rust, retaining Python workers for the ML models.

03

Constraints

  • Execution on the client’s computer, without cloud services.
  • Highly variable document and table structures.
  • Results traceable to cells in the original PDF.
  • A memory budget compatible with users’ machines.

04

Architecture & pipeline

Conceptual diagram
  1. XGBoost selects relevant pages to avoid extraction processing on pages without coverage data.
  2. DETR / TATR detect tables and recover their rows, columns and cells.
  3. CamemBERT classifies subheadings and maps domain terms to the expected coverage schema.
  4. Bounding boxes are reconstructed on the original PDF so operators can locate each value’s source cell, review it and make corrections before export.

05

Decisions & iterations

Fully local execution requires controlling RAM usage, model loading and orchestration on the client’s computer, without remote services.

I separated the Rust application core from the Python ML workers to reduce memory usage while retaining existing models and libraries.

I linked extracted values to their source PDF cells so value and alignment errors can be checked in the HITL interface.

Limited real documents led to fine-tuning TATR on synthetic tables. I tracked rows, columns and merged cells separately to identify regressions.

06

Results & limitations

Rewriting the application core reduced RAM usage by roughly 70%: from 3–10 GB to 1–3 GB, depending on documents and models used.

The application extracts coverage data locally and lets operators verify each value against its source PDF cell and correct it before export.

Fine-tuning on synthetic data improves rows and columns but degrades merged-cell handling. Structural gains do not extend to every cell type.

Unusual tables and ambiguous domain terms remain sources of error requiring human review.

07

Next iterations

Improve synthetic table generation, particularly merged cells, then reevaluate each structural element before increasing dataset size. Next, compare CamemBERT with a small CPU-based LLM, measuring mapping quality, latency and RAM usage.

08

In pictures

From document processing to Human-in-the-Loop review and export.

Other projects

Available for freelance projects

Need to improve an ML prototype?

Describe the errors you observe, the data available and your deployment constraints. We can define the next steps and the results to measure.

Contact meSaint-Étienne · Lyon · Remote · Available for freelance work · quotes on request