About The Role
The role is for someone who has moved beyond basic prompting and understands what it takes to build production-grade AI systems: robust RAG pipelines, agentic workflows, fine-tuning pipelines, and systematic evaluation frameworks.
The engineering team owns complex pieces of a high-scale AI platform, working directly with applied scientists and backend engineers to deploy performant LLM applications.
Key Responsibilities
- Design and implement production RAG pipelines using LangChain, LlamaIndex, or custom Python architectures
- Build and optimize vector database integrations such as Pinecone, Weaviate, or pgvector for low-latency semantic search
- Develop systematic LLM evaluation frameworks including benchmark suites, LLM-as-judge pipelines, and automated regression testing
- Execute parameter-efficient fine-tuning pipelines using LoRA and QLoRA on specialized domain datasets
- Write observable, tested, and well-documented Python code; participate actively in architecture reviews and deployment pipelines
What We Are Looking For
- 3 to 6 years of software engineering experience, with at least 2 years specifically focused on building LLM applications in production
- Deep familiarity with LLM orchestration frameworks, prompt engineering best practices, and context window optimization
- Strong proficiency in Python, async programming, REST API design, and cloud infrastructure integration
- Solid understanding of embedding models, vector spaces, tokenization, and model quantization techniques
- Bonus: Experience with custom CUDA kernels, open-source model deployment via vLLM or TGI, and published contributions to AI open-source projects