← RAG Evaluation & RegressionAll roles
RAG eval · ContractRemote

RAG Evaluation Specialist

SeniorApply by 2026-12-31
RAG EvaluationFaithfulness MetricsCitation ChecksPythonGolden DatasetsRegression Packs
Apply now

The work

Build and maintain RAG regression packs: faithfulness, context precision/recall, and citation checks wired into CI. You work with client-chosen observability stacks — Halcer does not resell eval platforms.

You will

  • Curate golden datasets that reflect real user queries and document edge cases
  • Automate retrieval regression when indexes, chunking, or embed models change
  • Calibrate LLM-as-judge scorers and document known failure modes
  • Partner with retrieval engineers so eval failures map to actionable fixes

You likely have

  • 4+ years in data/ML or QA engineering with 2+ on production RAG systems
  • Hands-on experience with at least one eval toolchain (ecosystem tools or in-house harness)
  • Clear writing on what a score does and does not prove
  • Comfort saying “do not ship” when thresholds are not met

Candidate apply

Apply for RAG Evaluation Specialist

Six fields to start. Your profile is saved locally for faster re-apply across roles.

Resume

Upload PDF or Word, or paste your experience. One is enough.

Private POST only — applications are not published on this site. Files go to our hiring pipeline for screening.