Skip to content

Work / DocExtract

DocExtract: document extraction with an eval gate in CI

Turning PDFs and scans into validated, searchable records, with an extraction eval that runs as a check on every relevant pull request.

Problem

Document extraction looks easy in a demo and drifts in production. A prompt tweak or a model swap can quietly drop field accuracy on invoices, receipts and statements, and nobody notices until a downstream report is wrong. DocExtract was built under a paid client engagement that needed structured fields it could trust, plus evals that run as checks rather than as a one-off demo.

Approach

Architecture

DocExtract request path from upload through FastAPI, Redis and ARQ, classification, two-pass extraction, validation, embedding and pgvector storage to agentic RAG, plus a CI-only offline replay checked against an accuracy floor.
Simplified from the mermaid diagram in the README.

Results

ResultValueSource
Weighted field-level score (critical fields 2×)95.5% on a 28-fixture offline replayevidence note
CI floor0.85; the check also fails on a drop of more than 0.03 from baselineeval-gate.yml
Regression caughtA demo branch that corrupted eight fixtures failed the checkeval-gate proof
Eval authoring corpus200 cases (150 golden, 50 adversarial)evals/
Dot chart of weighted field-level accuracy per document type from the 28-fixture offline replay, with the 0.85 CI floor marked.
Replay score by document type, generated from the replay output by scripts/render_eval_chart.py.

What I'd do next

Methodology & limits

What the numbers cover
  • The score is a deterministic replay of committed prediction fixtures, not a live-model grade and not F1. It gives partial credit: strings by normalized Levenshtein similarity, numbers within ±1%, and half credit for a value where null was expected. The 200-case authoring corpus is a separate population from the 28 scored fixtures.
  • The replayed outputs are frozen, so a prompt or model change does not move the replay score. The paid live eval job measures those changes and runs only when an API key is configured in CI.
  • Retrieval quality (recall, support, abstention) is not measured yet.
  • Branch protection on main requires the CI test check (since 2026-10-03). The weighted replay runs on eval-relevant pull requests and is not a universal required merge check.
  • Cost and latency figures in the repo are modeled until a metered run is committed.
  • DocExtract was built under a paid client engagement; an authorized walkthrough is available on request.
  • Full detail: the Methodology & limits section of the README.

Contact

Open to full-time AI Engineer roles

US remote. Applied AI, AI backend and LLM platform teams.