← Back to Talent Bench
Ref #HL-GENAI-814Senior Level · 8 Years Experience
Deployment ReadySenior LLMOps & Semantic Search Engineer
Specialization: Gen AI & Machine Learning · Industries: E-Commerce & Retail, Customer Experience, Publishing · Vetting Rating: Interview-ready · LLM infrastructure review
Contract EngagementContractor / Deployable
AvailabilityAvailable in 1 week
Work ModelRemote worldwide
Benchmark Rate~$95 / hr
Verified Technical Stack:
Vector DatabasesCloudflare Workers AIEmbedding ModelsSemantic CachingPythonTypeScriptDockerKubernetesHugging Face TransformersTriton ServervLLM
Professional Summary
Production ML engineer focused on low-latency LLM inference, embedding pipelines, and semantic search at catalog scale.
Verified Impact & Project Highlights
- Built semantic caching layer that materially reduced LLM token spend in production
- Implemented automated prompt regression and evaluation suites using LLM-as-judge frameworks
- Migrated legacy lexical search to hybrid dense-sparse retrieval with tracked quality and conversion readouts
Request an interview with HL-GENAI-814
Hiring teams only. We reply within one business day with slots. Candidate names stay off this site.
Engineer looking at this profile? Do not fill the interview form.Apply to join the bench.