Project
Document Extraction Pipeline
Turn Documents into Structured Data with AI
A production-ready FastAPI service that extracts structured data from PDFs and images using OCR and LLM technology. Features async processing, JWT authentication, and support for multiple document types including invoices, legal documents, and ESG reports.
- FastAPI
- Celery
- PostgreSQL
- Redis
- MinIO
- OpenAI
- Ollama
- Docker
- MinerU
Key Features
Multi-Format Support
Upload PDF or image files (PNG, JPG, TIFF) up to 10MB. Automatic format detection and preprocessing.
Async Processing
Non-blocking API calls with Celery workers. Submit documents and poll for results without waiting.
OCR + LLM Extraction
MinerU (vlm-auto-engine) converts documents into structured Markdown preserving tables and layouts, then OpenAI GPT-4o-mini or local Ollama models extract structured fields.
Structured Output
Pre-built schemas for Invoice, Legal, and ESG documents with customizable extraction templates.
Architecture
- FastAPI for high-performance async API endpoints
- Celery workers for background document processing
- PostgreSQL with JSONB for flexible result storage
- Redis as Celery broker and caching layer
- MinIO for S3-compatible object storage
- JWT authentication with per-user rate limiting
Diagrams
FastAPI + JWT accepts uploads, persists to MinIO, enqueues Celery jobs in Redis. Workers run MinerU OCR, hand off to an LLM for schema-driven extraction, and store results as JSONB in Postgres.
Async request life-cycle: upload returns a job id immediately; a worker fetches the file, runs OCR + LLM extraction, and persists the structured result. Clients poll the result endpoint until status = completed.
API Usage
# Upload a document
curl -X POST http://localhost:8000/documents/upload \
-H "Authorization: Bearer <token>" \
-F "file=@invoice.pdf" \
-F "schema=invoice"
# Response: {"job_id": "uuid", "status": "queued"}
# Poll for results
curl http://localhost:8000/documents/result/{job_id} \
-H "Authorization: Bearer <token>"
# Result: Structured JSON with extracted fieldsImpact
Demonstrates async document processing, schema-validated JSON output, and local or managed LLM backends. No comparative speed or extraction-accuracy benchmark is claimed.
Async
Queued Processing
Typed
JSON Outputs
3
Supported Schemas
Want something like this built for you?
Tell me about it. I reply within one working day with a first take and no sales pitch.