Page 1 of 4

AI Engineer

We build, optimize, and deploy production-grade AI systems and intelligent agent workflows for enterprise applications. As an AI Engineer Intern, you will work directly on our core model stack—building robust RAG pipelines, fine-tuning task-specific models, and optimizing high-throughput LLM inference for low-latency production environments.

What you'll achieve

• Design and implement advanced RAG (Retrieval-Augmented Generation) architectures using vector databases (Pinecone, Qdrant, or Weaviate), hybrid search, and semantic re-ranking.

• Fine-tune and evaluate open-source models (Llama, Mistral) using PEFT/LoRA techniques and curate specialized datasets for domain-specific downstream tasks.

• Build end-to-end LLM application logic and agentic workflows using Python, TypeScript, and frameworks like LangChain, LlamaIndex, or AutoGen.

• Establish automated evaluation frameworks (using tools like Ragas or TruLens) to continuously benchmark hallucination rates, retrieval accuracy, and token latency across model iterations.

What you bring

Strong Programming Fundamentals: Proficient in Python (required) and comfortable with TypeScript or Go. Familiarity with modern AI/ML libraries like PyTorch, Hugging Face transformers, and datasets.

Hands-On AI/LLM Experience: Proven project experience (via GitHub, research, or prior internships) building applications with LLMs, vector search, or custom embeddings.

Systems & Data Engineering Mindset: Understanding of modern API design (REST/gRPC), asynchronous execution, and standard database infrastructure (SQL, NoSQL, Vector DBs).

Analytical Approach: Demonstrated ability to systematically debug model performance issues, clean messy text datasets, and evaluate output quality programmatically.

Preferred Qualifications

• Experience with containerization (Docker) and cloud deployments (AWS, GCP, or Modal/RunPod).

• Familiarity with fine-tuning techniques (LoRA, QLoRA, SFT, DPO/RLHF).

• Contributions to open-source AI libraries or published work in machine learning/NLP.

What we offer

Direct Production Impact: Your code will ship straight into production environments handling live enterprise traffic — no throwaway research projects.

Mentorship: Daily pairing with senior AI and infrastructure engineers.

Tools & Hardware: Access to dedicated GPU infrastructure and standard cloud resource allocations.

Fast-Track Potential: Clear path to a full-time role based on performance during the internship.