Skip to content
#

document-data-extraction

Here are 4 public repositories matching this topic...

Language: All
Filter by language

Part of the KDAN ecosystem, DocSlight offers document parsing, OCR, and data extraction that turn PDFs, scans, images, and Office files into structured outputs for RAG pipelines, AI agents, and enterprise document automation.

  • Updated Jul 20, 2026
  • Vue

Benchmarked invoice and document extraction pipeline: PII redaction before any model call, layout-aware extraction into a strict schema, field-level confidence scoring, validation rules, a human review queue, and corrections fed back as exemplars. Runs offline, no API key.

  • Updated Jul 26, 2026
  • HTML

Improve this page

Add a description, image, and links to the document-data-extraction topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the document-data-extraction topic, visit your repo's landing page and select "manage topics."

Learn more