About The Role
The role is for someone who has moved beyond basic prompting and understands what it takes to build production-grade AI systems: robust RAG pipelines, agentic workflows, fine-tuning pipelines, and systematic evaluation frameworks.
The position will own complex pieces of the core AI platform and work directly with applied scientists, backend engineers, and enterprise stakeholders to deliver scalable generative AI applications.
Key Responsibilities
- Design and implement end-to-end RAG pipelines using LangChain, LlamaIndex, or custom retrieval architectures
- Build and optimize vector database integrations including Pinecone, Weaviate, or pgvector for semantic search at production scale
- Develop systematic LLM evaluation frameworks using benchmark suites, LLM-as-judge pipelines, and automated regression testing
- Execute instruction fine-tuning and parameter-efficient fine-tuning techniques such as LoRA and QLoRA on domain-specific datasets
- Write observable, tested, and well-documented Python code while participating actively in architecture and code reviews
What We Are Looking For
- 3 to 6 years of software engineering experience, with a minimum of 2 years specifically focused on LLMs and generative AI in production
- Deep familiarity with LLM orchestration frameworks like LangChain or LlamaIndex and hands-on experience with major foundation model APIs and open-source models
- Solid understanding of embedding models, vector databases, and semantic similarity optimization in distributed environments
- Strong Python development skills with proven comfort in asynchronous programming, RESTful API design, and cloud infrastructure
- Bonus: Experience publishing open-source AI tools, contributions to ML research, or background in distributed systems performance tuning