You are viewing a preview of this job. Log in or register to view more details about this job.

Prompt Engineering Lead

We are looking for a Prompt Engineering & AI Evals Lead to own prompt engineering - one of the core service categories on the REWORK platform - and, internally, the evaluation rigor behind every AI decision REWORK makes about a person's work. Proof-of-work submissions get auto-verified against rules and thresholds; skill assessments are AI-generated and AI-scored through SwitchAsses; matching and the Rework Score lean on model output. Bad prompts here mean a false auto-approval or an unfair rejection - real consequences for real professionals. You will build the eval harnesses, golden test sets, and prompt-version pipelines that keep these systems honest, and do the same prompt-design and red-teaming work as a client-facing deliverable for businesses that need their own AI systems evaluated and hardened.

Core Responsibilities

Eval Infrastructure

  • Build eval harnesses and golden test sets for auto-verification rules, matching/scoring prompts, and assessment generation
  • Version prompts and track accuracy, false-positive, and false-negative rates release over release
  • Design regression tests so a prompt tweak can't silently degrade auto-approval accuracy

Prompt Design & Hardening

  • Write and refine system prompts for auto-verification, SwitchAsses assessment scoring, and Rezy
  • Red-team prompts for jailbreaks, prompt injection, and edge cases before they reach production
  • Design prompts as a client deliverable - eval suites and hardening work for business AI systems

Measurement & Reporting

  • Quantify how auto-verification and scoring accuracy move with each prompt or model change
  • Flag drift when a model update changes behavior on the existing eval set
  • Document prompt versions, eval results, and hardening decisions so the reasoning is auditable later

Required Qualifications

Required Experience

  • Shipped production prompts with a measurable eval process behind them, not just prompt tuning by feel
  • 1+ years hands-on with the OpenAI or Anthropic APIs and at least one eval/testing framework

System Design & Problem-Solving

  • Thinks in false-positive vs false-negative tradeoffs and designs thresholds accordingly
  • Rigorous about reproducibility - same eval set, same scoring, comparable results across prompt versions

Documentation & Nice-to-Haves

  • Documents eval methodology and results so decisions are defensible, not just "it felt better"
  • Nice to have: statistics/ML background, red-teaming experience, trust & safety or content-moderation work

Key Deliverables

Measured, versioned eval suites for auto-verification, scoring, and assessment promptsDocumented accuracy and drift tracking across prompt and model changesHardened prompts resistant to jailbreaks and prompt injectionClear, well-documented eval methodology and prompt-version history

Tech Stack & Skills

System-prompt designFew-shot / Chain-of-ThoughtEval harnesses (Promptfoo or similar)Golden test-set designA/B testingConstitutional AIRed-teaming / jailbreak hardeningMulti-step tool-call promptsPythonOpenAI / Anthropic APIsStatistics basics

Expectations

  • Work asynchronously in a remote-first environment (Slack, email, documented reports)
  • Operate highly independently - evaluated on proof-of-work and reliability of deliverables
  • Collaborate during core hours, 10:00 AM – 4:00 PM Eastern Time
  • Keep eval methodology, prompt versions, and results clear, well-documented, and easy to navigate