← RAG Evaluation & RegressionAll roles
RAG eval · ContractRemote
GenAI Evaluation Engineer
SeniorApply by 2026-12-31
LLM EvaluationGuardrailsPythonPrompt RegressionRAG QualitySafety Policies
Apply nowThe work
Join a Halcer Gen AI delivery pod as the engineer who makes model behaviour measurable. You will build evaluation suites, prompt regression, and safety guardrails so agentic systems can ship with evidence — never with candidate or customer PII in public indexes.
You will
- Design offline and online evaluation sets for groundedness, toxicity, PII leakage, and tool-call accuracy
- Implement guardrail layers for retrieval, generation, and function calling
- Automate prompt regression in CI so model or index changes cannot silently degrade quality
- Document residual risk for delivery leads in plain language
You likely have
- 5+ years in ML/NLP engineering with at least 2 years on production LLM systems
- Hands-on evaluation frameworks (LLM-as-judge, human review sampling, or equivalent)
- Practical experience redacting or blocking PII in logs and traces
- Strong Python engineering and CI discipline
Candidate apply
Apply for GenAI Evaluation Engineer
Six fields to start. Your profile is saved locally for faster re-apply across roles.
GenAI Evaluation Engineer
Apply