← Back to Talent Bench
Ref #HL-AI-905Senior Level · 7 Years Experience
Ready to InterviewSenior LLM Evaluation & Guardrails Engineer
Specialization: Gen AI Evaluation & Safety · Industries: Enterprise SaaS, Healthcare AI, Customer Support Platforms · Vetting Rating: Interview-ready · LLM evaluation review
Contract EngagementContractor / Deployable
AvailabilityAvailable in 1 week
Work ModelRemote worldwide
Benchmark Rate~$100 / hr
Verified Technical Stack:
LLM EvaluationGuardrailsPythonRAG QualityPrompt RegressionLangSmith / PhoenixSafety ClassifiersCI for MLPII Redaction Patterns
Professional Summary
Evaluation specialist who treats LLM quality as an engineering system: datasets, judges, guardrails, and regression gates. Works with client-chosen tracing stacks — not a platform vendor.
Verified Impact & Project Highlights
- Built groundedness and PII-leakage eval suites used as CI gates before model promotions
- Designed tool-call accuracy scoring for multi-step agents prior to production traffic
- Maintains judge calibration docs and version pins alongside application releases
Request an interview with HL-AI-905
Hiring teams only. We reply within one business day with slots. Candidate names stay off this site.
Engineer looking at this profile? Do not fill the interview form.Apply to join the bench.