The LLM / GenAI Engineer will design, build, and deploy production AI systems, including retrieval-augmented generation pipelines, tool-using agents, model fine-tuning workflows, and evaluation infrastructure.
The role
spans experimentation and backend engineering, with a focus on reliability, latency, cost, and measurable model quality. Working with applied scientists, platform engineers, and product teams,
the role
will turn emerging foundation-model capabilities into secure, observable services used in real customer workflows. The position is remote, with a preference for candidates based in Austin, TX.
Key Responsibilities
Design and implement production RAG and agentic workflows using Python, LangChain, LlamaIndex, or equivalent frameworks
Build ingestion, chunking, embedding, reranking, and retrieval pipelines backed by vector stores such as pgvector, Pinecone, Weaviate, or Milvus
Develop LLM evaluation systems covering offline benchmarks, groundedness, factuality, safety, latency, cost, and regression testing
Fine-tune and optimize open-source language models using supervised fine-tuning, LoRA, QLoRA, quantization, and distributed training techniques
Deploy model-powered services through containerized APIs and cloud infrastructure using Docker, Kubernetes, AWS, GCP, or Azure
Instrument production systems with tracing, logging, feedback collection, and monitoring for quality degradation, drift, latency, and token usage
Partner with software engineers and applied scientists on architecture reviews, data strategy, experimentation, code quality, and production incident response
What We Are Looking For
3–8 years of experience in software engineering, machine learning engineering, or applied AI, including at least 1 year delivering LLM or GenAI systems to production
Advanced Python skills with experience building asynchronous services, REST or gRPC APIs, testing frameworks, and maintainable production code
Strong understanding of transformer-based language models, embeddings, tokenization, context windows, prompting, structured generation, and inference tradeoffs
Hands-on experience with RAG architecture, vector databases, hybrid search, reranking, document processing, and retrieval-quality measurement
Experience with at least one major cloud platform and production deployment tooling, including Docker, Kubernetes, CI/CD, and observability systems
Bachelor’s or master’s degree in computer science, engineering, mathematics, or a related technical field, or equivalent practical experience
Bonus: Experience with distributed GPU training or inference, open-source model serving, multimodal models, guardrails, MLflow, vLLM, TensorRT-LLM, or enterprise AI security