
AI/ML Engineer
About Us
We are a fast-growing startup based in Pune, India, specializing in cutting-edge Data Science and Data Engineering solutions. Our team of dedicated professionals is committed to solving complex data challenges for companies worldwide.
Our Culture
We foster a vibrant startup culture that values:
Intellectual curiosity
Continuous learning
Positive work environment
Collaborative problem-solving
Role Overview
We are seeking a highly technical and proactive AI-ML Engineer to join our dynamic team. The ideal candidate will possess a strong blend of software engineering rigor and deep technical expertise in modern AI/ML architectures, infrastructure optimization, and generative AI systems. This role demands critical thinking, production-grade coding, and the ability to deploy, optimize, and scale robust AI models to solve complex, real-world problems. You should bring a deep curiosity for understanding model internals, and the adaptability required to navigate and implement rapidly evolving, state-of-the-art technologies.
Key Responsibilities
Deliver end-to-end AI engineering projects by designing, building, and deploying production-grade Machine Learning, Deep Learning, and Generative AI applications.
Develop high-quality software solutions in Python, collaborating with cross-functional engineering teams to integrate AI models into existing application codebases.
Implement advanced training strategies including mixed-precision training (FP16/BF16), gradient accumulation, and distributed training (Data/Model/Pipeline parallel) while profiling and maximizing GPU utilization.
Optimize, serialize (ONNX, TorchScript), and deploy models via high-throughput REST APIs (FastAPI) while managing latency vs. throughput trade-offs.
Implement robust MLOps workflows using DVC, Docker, and cloud platforms for model versioning, pipeline automation, and production monitoring.
Architect and optimize high-performance Retrieval-Augmented Generation (RAG) systems using hybrid search, semantic chunking, and cross-encoder reranking.
Design autonomous, multi-agent systems and orchestrated workflows utilizing function calling, the ReAct pattern, and advanced memory management architectures.
Implement state-of-the-art serving techniques (vLLM, speculative decoding, prompt caching) and quantization (INT8/INT4 via GPTQ/AWQ) to maximize inference throughput.
Move workflows from notebooks to production-grade pipelines, writing clean code and implementing unit/integration tests for ML (pytest, Great Expectations).
Actively diagnose production anomalies, including data/concept distribution shifts, silent model failures, and training bottlenecks (underfitting/overfitting).
Evaluate and benchmark LLM outputs using appropriate metrics and testing frameworks.
Design high-throughput data pipelines and optimized SQL/NoSQL queries for large-scale data processing and model feature injection.
Practice active listening to understand project requirements and team inputs.
Collaborate with stakeholders to translate complex business requirements into scalable AI/ML solutions and communicate technical trade-offs clearly.
Demonstrate strong technical communication skills, a high degree of ownership, and an action-biased approach to debugging and solving ambiguous engineering problems
Apply responsible AI principles, jailbreak awareness, and output validation guardrails to ensure ethical and safe model development.
Plan strategically and multitask efficiently to meet project deadlines.
Required Skills
Core Programming & ML
Strong Python programming skills with hands-on project experience, Git proficiency, and a basic understanding of CUDA
Expertise in Deep Learning architectures (Transformers, CNNs, RNNs) alongside a strong theoretical understanding of foundational ML algorithms (GBMs, Random Forests)
Deep proficiency in PyTorch or TensorFlow, with a strong emphasis on custom layer implementation and neural network training loops
Hands-on experience with modern training paradigms, including Self-supervised Learning, Contrastive Learning, and advanced Transfer Learning
Experience with NLP, Computer Vision, or Time Series Analysis
Proven experience writing clean, production-grade code, utilizing pytest and Great Expectations for data and model validation
Generative AI & LLMs
Hands-on experience with commercial LLM APIs (OpenAI, Anthropic, Groq) and hosting/deploying open-source foundational models (Llama, Mistral)
Proficiency with modern orchestration frameworks (LangChain, LlamaIndex) and building stateful multi-agent architectures using LangGraph or DSPy
Experience with vector databases (Pinecone, Weaviate, Chroma, pgvector), advanced chunking (fixed, semantic, recursive), and hybrid search implementation
Deep technical understanding of Parameter-Efficient Fine-Tuning (PEFT) mechanics—specifically LoRA/QLoRA low-rank decomposition—alongside instruction tuning and model alignment methodologies (DPO, RLHF)
Mastery of advanced prompt engineering, including structured output forcing, Chain-of-Thought, system prompt design, and building resilient multi-agent coordination systems
Hands-on experience with inference optimization and high-throughput serving, including quantization (GPTQ, AWQ), speculative decoding, prompt caching, and vLLM acceleration
MLOps & Deployment
Experience with MLOps practices, logging, and model registries (MLflow, Weights & Biases, DVC) along with model serving via Triton Inference Server or FastAPI
Production experience deploying, scaling, and monitoring models natively on cloud AI platforms, specifically AWS SageMaker or GCP Vertex AI
Experience building CI/CD pipelines for ML applications, with an emphasis on data and model versioning/lineage using DVC and Git
Data Engineering & Databases
Solid understanding of SQL, including advanced concepts like windowing functions and query optimization
Experience building and orchestrating data pipelines using Airflow or Prefect to feed specialized SQL and NoSQL vector/feature databases
Soft Skills & Professional Attributes
Strong critical thinking and problem-solving skills
Excellent written and verbal communication abilities
High degree of flexibility and adaptability to stay ahead of the rapid velocity of the open-source AI ecosystem (Hugging Face, vLLM, etc.)
Practical understanding of AI safety, compliance, guardrails (e.g., NeMo Guardrails or Llama Guard), and responsible AI practices.
Nice-to-Have
Contributions to open-source AI/ML repositories or GenAI orchestration frameworks.
Experience with real-time streaming data processing
Active participation in competitive ML spaces (Kaggle) or track record of reviewing/reproducing state-of-the-art AI research papers
Published research papers or conference presentations
Experience building and scaling Graph Databases or Knowledge Graphs for advanced RAG
Experience with multimodal architectures (Vision-Language models, Audio processing)
Qualifications
AI-ML Engineer: 2–5 years of hands-on experience engineering machine learning systems, optimizing infrastructure, and implementing LLM/GenAI workflows in production
Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, Statistics, or a related highly quantitative field
Demonstrated commitment to continuous learning through contributing to open-source, certifications, or self-study (especially in deep learning internals, MLOps, and modern GenAI frameworks)
What We Offer
Competitive salary commensurate with experience
Opportunity to work on diverse, cutting-edge AI/ML projects
Collaborative and innovation-driven work environment
Rapid growth and continuous learning opportunities
Exposure to latest AI technologies and industry best practices
.png)