Skip to content
VDAI with VD

Project

Document Extraction Pipeline

Turn Documents into Structured Data with AI

Self-Built ProjectDemonstrated

A production-ready FastAPI service that extracts structured data from PDFs and images using OCR and LLM technology. Features async processing, JWT authentication, and support for multiple document types including invoices, legal documents, and ESG reports.

  • FastAPI
  • Celery
  • PostgreSQL
  • Redis
  • MinIO
  • OpenAI
  • Ollama
  • Docker
  • MinerU

Key Features

Multi-Format Support

Upload PDF or image files (PNG, JPG, TIFF) up to 10MB. Automatic format detection and preprocessing.

Async Processing

Non-blocking API calls with Celery workers. Submit documents and poll for results without waiting.

OCR + LLM Extraction

MinerU (vlm-auto-engine) converts documents into structured Markdown preserving tables and layouts, then OpenAI GPT-4o-mini or local Ollama models extract structured fields.

Structured Output

Pre-built schemas for Invoice, Legal, and ESG documents with customizable extraction templates.

Architecture

  • FastAPI for high-performance async API endpoints
  • Celery workers for background document processing
  • PostgreSQL with JSONB for flexible result storage
  • Redis as Celery broker and caching layer
  • MinIO for S3-compatible object storage
  • JWT authentication with per-user rate limiting

Diagrams

Document Extraction Pipeline architecture

FastAPI + JWT accepts uploads, persists to MinIO, enqueues Celery jobs in Redis. Workers run MinerU OCR, hand off to an LLM for schema-driven extraction, and store results as JSONB in Postgres.

Document Extraction request flow

Async request life-cycle: upload returns a job id immediately; a worker fetches the file, runs OCR + LLM extraction, and persists the structured result. Clients poll the result endpoint until status = completed.

API Usage

Upload & Extract
# Upload a document
curl -X POST http://localhost:8000/documents/upload \
  -H "Authorization: Bearer <token>" \
  -F "file=@invoice.pdf" \
  -F "schema=invoice"

# Response: {"job_id": "uuid", "status": "queued"}

# Poll for results
curl http://localhost:8000/documents/result/{job_id} \
  -H "Authorization: Bearer <token>"

# Result: Structured JSON with extracted fields

Impact

Demonstrates async document processing, schema-validated JSON output, and local or managed LLM backends. No comparative speed or extraction-accuracy benchmark is claimed.

Async

Queued Processing

Typed

JSON Outputs

3

Supported Schemas

Want something like this built for you?

Tell me about it. I reply within one working day with a first take and no sales pitch.