01
The problem
Health insurance coverage tables do not follow a single format. Providers change layouts, terminology and how options are presented. Text extraction alone is not enough: the system must recover structure, understand domain terms and let an operator verify every result.
02
My role
I designed the vision / NLP pipeline and built the Human-in-the-Loop (HITL) review interface. I then rewrote the Python MVP’s application core in Rust, retaining Python workers for the ML models.
03
Constraints
- Execution on the client’s computer, without cloud services.
- Highly variable document and table structures.
- Results traceable to cells in the original PDF.
- A memory budget compatible with users’ machines.
04
Architecture & pipeline
- XGBoost selects relevant pages to avoid extraction processing on pages without coverage data.
- DETR / TATR detect tables and recover their rows, columns and cells.
- CamemBERT classifies subheadings and maps domain terms to the expected coverage schema.
- Bounding boxes are reconstructed on the original PDF so operators can locate each value’s source cell, review it and make corrections before export.
05
Decisions & iterations
Fully local execution requires controlling RAM usage, model loading and orchestration on the client’s computer, without remote services.
I separated the Rust application core from the Python ML workers to reduce memory usage while retaining existing models and libraries.
I linked extracted values to their source PDF cells so value and alignment errors can be checked in the HITL interface.
Limited real documents led to fine-tuning TATR on synthetic tables. I tracked rows, columns and merged cells separately to identify regressions.
06
Results & limitations
Rewriting the application core reduced RAM usage by roughly 70%: from 3–10 GB to 1–3 GB, depending on documents and models used.
The application extracts coverage data locally and lets operators verify each value against its source PDF cell and correct it before export.
Fine-tuning on synthetic data improves rows and columns but degrades merged-cell handling. Structural gains do not extend to every cell type.
Unusual tables and ambiguous domain terms remain sources of error requiring human review.
07
Next iterations
Improve synthetic table generation, particularly merged cells, then reevaluate each structural element before increasing dataset size. Next, compare CamemBERT with a small CPU-based LLM, measuring mapping quality, latency and RAM usage.
08


