← Back to Talent Bench
Ref #HL-EVAL-2202Senior Level · 6 Years Experience
Deployment ReadyRAG Evaluation & Regression Specialist
Specialization: Gen AI Evaluation & Safety · Industries: Enterprise SaaS, LegalTech, Customer Support Platforms · Vetting Rating: Interview-ready · RAG eval review
Contract EngagementContractor / Deployable
AvailabilityImmediate (Within 48 hrs)
Work ModelRemote worldwide
Benchmark Rate~$95 / hr
Verified Technical Stack:
RAG EvaluationFaithfulness MetricsCitation ChecksGolden DatasetsPythonCI GatesLLM-as-Judge CalibrationChunk Policy TestingDeepEval / PromptfooRetrieval Debugging
Professional Summary
Maintains RAG regression packs — faithfulness, citation, and context checks — wired into deploy pipelines. Implements with client-selected eval tooling.
Verified Impact & Project Highlights
- Owns golden datasets tied to retrieval config versions and embed model upgrades
- Partners with retrieval engineers to map eval failures to chunking and index changes
- Documents judge drift when LLM-as-judge models are upgraded
Request an interview with HL-EVAL-2202
Hiring teams only. We reply within one business day with slots. Candidate names stay off this site.
Engineer looking at this profile? Do not fill the interview form.Apply to join the bench.