Document AI Engineer
We are looking for a Document AI & RPA Engineer to own the document side of our automation practice - one of the core service categories on the REWORK platform. Client businesses process invoices, intake forms, contracts, and compliance paperwork by hand; you will replace that with engineered pipelines: multi-channel ingestion (mail parsing, SFTP drops, uploads), document classification, OCR/layout extraction, LLM-based structured extraction validated against JSON schemas, confidence-thresholded human-in-the-loop review, and idempotent, audit-logged writes into accounting, CRM, and ERP systems - including browser-level RPA where no API exists. This is a deeply technical role: you should care about extraction accuracy on golden document sets, dedup on re-ingest, and pipelines that fail loudly and recover cleanly.
Core Responsibilities
Ingestion & Extraction
- Build multi-channel document intake: parsed mailboxes, SFTP/watched folders, webhook and portal uploads
- Classify document types and route them through OCR/layout models plus LLM extraction constrained by JSON schemas
- Score extraction confidence per field and design the thresholds that decide auto-post vs human review
RPA & System Writes
- Automate legacy and no-API systems with UiPath, Power Automate, or Playwright using resilient selectors and recovery paths
- Design idempotent, audit-logged writes into accounting/CRM/ERP targets - no double-posted invoices, ever
Reliability, Review & Evaluation
- Build human-in-the-loop review queues with clear SOPs for the exceptions your thresholds surface
- Engineer retries, dead-letter handling, and alerting; measure extraction accuracy and drift against golden document sets
Required Qualifications
Required Experience
- Shipped document-processing or RPA systems that ran in production for real businesses
- 1+ years hands-on with at least one Document AI/OCR stack AND one RPA toolchain, plus strong Python
System Design & Problem-Solving
- Schema-first extraction design; sound judgment on accuracy vs cost tradeoffs across OCR, layout models, and LLM passes
- Rigorous about idempotency, dedup on re-ingest, PII-safe handling, and retention - able to debug a mis-parsed field back to its pixel
Documentation & Nice-to-Haves
- Documents extraction schemas, confidence policies, and review SOPs so operators can run the system without you
- Nice to have: LLM evaluation harnesses, vector search for document retrieval, AP/AR or ERP domain experience
Key Deliverables
Manual document handling measurably reduced or eliminatedExtraction pipelines running reliably with measured accuracy on golden document setsException review queues that human operators can actually workClear, well-documented document architecture and audit trails
Tech Stack & Skills
Google Document AIAWS TextractAzure Document IntelligenceTesseractUiPathPower AutomatePlaywright / browser RPALLM structured extraction (JSON Schema)PythonQueues & retriesRegex & parsingPII handling
Expectations
- Work asynchronously in a remote-first environment (Slack, email, documented reports)
- Operate highly independently - evaluated on proof-of-work and reliability of deliverables
- Collaborate during core hours, 10:00 AM – 4:00 PM Eastern Time
- Keep extraction schemas, pipelines, and review workflows clear, well-documented, and easy to navigate