You are viewing a preview of this job. Log in or register to view more details about this job.

Document AI Engineer

 

 

We are looking for a Document AI & RPA Engineer to own the document side of our automation practice - one of the core service categories on the REWORK platform. Client businesses process invoices, intake forms, contracts, and compliance paperwork by hand; you will replace that with engineered pipelines: multi-channel ingestion (mail parsing, SFTP drops, uploads), document classification, OCR/layout extraction, LLM-based structured extraction validated against JSON schemas, confidence-thresholded human-in-the-loop review, and idempotent, audit-logged writes into accounting, CRM, and ERP systems - including browser-level RPA where no API exists. This is a deeply technical role: you should care about extraction accuracy on golden document sets, dedup on re-ingest, and pipelines that fail loudly and recover cleanly.

Core Responsibilities

 

Ingestion & Extraction

 

  • Build multi-channel document intake: parsed mailboxes, SFTP/watched folders, webhook and portal uploads
  • Classify document types and route them through OCR/layout models plus LLM extraction constrained by JSON schemas
  • Score extraction confidence per field and design the thresholds that decide auto-post vs human review

RPA & System Writes

 

  • Automate legacy and no-API systems with UiPath, Power Automate, or Playwright using resilient selectors and recovery paths
  • Design idempotent, audit-logged writes into accounting/CRM/ERP targets - no double-posted invoices, ever

Reliability, Review & Evaluation

 

  • Build human-in-the-loop review queues with clear SOPs for the exceptions your thresholds surface
  • Engineer retries, dead-letter handling, and alerting; measure extraction accuracy and drift against golden document sets

Required Qualifications

 

Required Experience

 

  • Shipped document-processing or RPA systems that ran in production for real businesses
  • 1+ years hands-on with at least one Document AI/OCR stack AND one RPA toolchain, plus strong Python

System Design & Problem-Solving

 

  • Schema-first extraction design; sound judgment on accuracy vs cost tradeoffs across OCR, layout models, and LLM passes
  • Rigorous about idempotency, dedup on re-ingest, PII-safe handling, and retention - able to debug a mis-parsed field back to its pixel

Documentation & Nice-to-Haves

 

  • Documents extraction schemas, confidence policies, and review SOPs so operators can run the system without you
  • Nice to have: LLM evaluation harnesses, vector search for document retrieval, AP/AR or ERP domain experience

Key Deliverables

 

Manual document handling measurably reduced or eliminatedExtraction pipelines running reliably with measured accuracy on golden document setsException review queues that human operators can actually workClear, well-documented document architecture and audit trails

Tech Stack & Skills

 

Google Document AIAWS TextractAzure Document IntelligenceTesseractUiPathPower AutomatePlaywright / browser RPALLM structured extraction (JSON Schema)PythonQueues & retriesRegex & parsingPII handling

Expectations

 

  • Work asynchronously in a remote-first environment (Slack, email, documented reports)
  • Operate highly independently - evaluated on proof-of-work and reliability of deliverables
  • Collaborate during core hours, 10:00 AM – 4:00 PM Eastern Time
  • Keep extraction schemas, pipelines, and review workflows clear, well-documented, and easy to navigate