← Back to Talent Bench
Ref #HL-EVAL-2202Senior Level · 6 Years Experience
Deployment Ready

RAG Evaluation & Regression Specialist

Specialization: Gen AI Evaluation & Safety · Industries: Enterprise SaaS, LegalTech, Customer Support Platforms · Vetting Rating: Interview-ready · RAG eval review

Contract EngagementContractor / Deployable
AvailabilityImmediate (Within 48 hrs)
Work ModelRemote worldwide
Benchmark Rate~$95 / hr
Verified Technical Stack:
RAG EvaluationFaithfulness MetricsCitation ChecksGolden DatasetsPythonCI GatesLLM-as-Judge CalibrationChunk Policy TestingDeepEval / PromptfooRetrieval Debugging

Professional Summary

Maintains RAG regression packs — faithfulness, citation, and context checks — wired into deploy pipelines. Implements with client-selected eval tooling.

Verified Impact & Project Highlights

  • Owns golden datasets tied to retrieval config versions and embed model upgrades
  • Partners with retrieval engineers to map eval failures to chunking and index changes
  • Documents judge drift when LLM-as-judge models are upgraded

Request an interview with HL-EVAL-2202

Hiring teams only. We reply within one business day with slots. Candidate names stay off this site.

Engineer looking at this profile? Do not fill the interview form.Apply to join the bench.