← RAG Evaluation & RegressionAll roles
RAG eval · ContractRemote

GenAI Evaluation Engineer

SeniorApply by 2026-12-31
LLM EvaluationGuardrailsPythonPrompt RegressionRAG QualitySafety Policies
Apply now

The work

Join a Halcer Gen AI delivery pod as the engineer who makes model behaviour measurable. You will build evaluation suites, prompt regression, and safety guardrails so agentic systems can ship with evidence — never with candidate or customer PII in public indexes.

You will

  • Design offline and online evaluation sets for groundedness, toxicity, PII leakage, and tool-call accuracy
  • Implement guardrail layers for retrieval, generation, and function calling
  • Automate prompt regression in CI so model or index changes cannot silently degrade quality
  • Document residual risk for delivery leads in plain language

You likely have

  • 5+ years in ML/NLP engineering with at least 2 years on production LLM systems
  • Hands-on evaluation frameworks (LLM-as-judge, human review sampling, or equivalent)
  • Practical experience redacting or blocking PII in logs and traces
  • Strong Python engineering and CI discipline

Candidate apply

Apply for GenAI Evaluation Engineer

Six fields to start. Your profile is saved locally for faster re-apply across roles.

Resume

Upload PDF or Word, or paste your experience. One is enough.

Private POST only — applications are not published on this site. Files go to our hiring pipeline for screening.