top of page

AI/ML Engineer

About Us


We are a fast-growing startup based in Pune, India, specializing in cutting-edge Data Science and Data Engineering solutions. Our team of dedicated professionals is committed to solving complex data challenges for companies worldwide.


Our Culture


We foster a vibrant startup culture that values:

  • Intellectual curiosity

  • Continuous learning

  • Positive work environment

  • Collaborative problem-solving


Role Overview


We are seeking a highly technical and proactive AI-ML Engineer to join our dynamic team. The ideal candidate will possess a strong blend of software engineering rigor and deep technical expertise in modern AI/ML architectures, infrastructure optimization, and generative AI systems. This role demands critical thinking, production-grade coding, and the ability to deploy, optimize, and scale robust AI models to solve complex, real-world problems. You should bring a deep curiosity for understanding model internals, and the adaptability required to navigate and implement rapidly evolving, state-of-the-art technologies.


Key Responsibilities


  • Deliver end-to-end AI engineering projects by designing, building, and deploying production-grade Machine Learning, Deep Learning, and Generative AI applications.

  • Develop high-quality software solutions in Python, collaborating with cross-functional engineering teams to integrate AI models into existing application codebases.

  • Implement advanced training strategies including mixed-precision training (FP16/BF16), gradient accumulation, and distributed training (Data/Model/Pipeline parallel) while profiling and maximizing GPU utilization.

  • Optimize, serialize (ONNX, TorchScript), and deploy models via high-throughput REST APIs (FastAPI) while managing latency vs. throughput trade-offs.

  • Implement robust MLOps workflows using DVC, Docker, and cloud platforms for model versioning, pipeline automation, and production monitoring.

  • Architect and optimize high-performance Retrieval-Augmented Generation (RAG) systems using hybrid search, semantic chunking, and cross-encoder reranking.

  • Design autonomous, multi-agent systems and orchestrated workflows utilizing function calling, the ReAct pattern, and advanced memory management architectures.

  • Implement state-of-the-art serving techniques (vLLM, speculative decoding, prompt caching) and quantization (INT8/INT4 via GPTQ/AWQ) to maximize inference throughput.

  • Move workflows from notebooks to production-grade pipelines, writing clean code and implementing unit/integration tests for ML (pytest, Great Expectations).

  • Actively diagnose production anomalies, including data/concept distribution shifts, silent model failures, and training bottlenecks (underfitting/overfitting).

  • Evaluate and benchmark LLM outputs using appropriate metrics and testing frameworks.

  • Design high-throughput data pipelines and optimized SQL/NoSQL queries for large-scale data processing and model feature injection.

  • Practice active listening to understand project requirements and team inputs.

  • Collaborate with stakeholders to translate complex business requirements into scalable AI/ML solutions and communicate technical trade-offs clearly.

  • Demonstrate strong technical communication skills, a high degree of ownership, and an action-biased approach to debugging and solving ambiguous engineering problems

  • Apply responsible AI principles, jailbreak awareness, and output validation guardrails to ensure ethical and safe model development.

  • Plan strategically and multitask efficiently to meet project deadlines.


Required Skills


Core Programming & ML


  • Strong Python programming skills with hands-on project experience, Git proficiency, and a basic understanding of CUDA

  • Expertise in Deep Learning architectures (Transformers, CNNs, RNNs) alongside a strong theoretical understanding of foundational ML algorithms (GBMs, Random Forests)

  • Deep proficiency in PyTorch or TensorFlow, with a strong emphasis on custom layer implementation and neural network training loops

  • Hands-on experience with modern training paradigms, including Self-supervised Learning, Contrastive Learning, and advanced Transfer Learning

  • Experience with NLP, Computer Vision, or Time Series Analysis

  • Proven experience writing clean, production-grade code, utilizing pytest and Great Expectations for data and model validation


Generative AI & LLMs


  • Hands-on experience with commercial LLM APIs (OpenAI, Anthropic, Groq) and hosting/deploying open-source foundational models (Llama, Mistral)

  • Proficiency with modern orchestration frameworks (LangChain, LlamaIndex) and building stateful multi-agent architectures using LangGraph or DSPy

  • Experience with vector databases (Pinecone, Weaviate, Chroma, pgvector), advanced chunking (fixed, semantic, recursive), and hybrid search implementation

  • Deep technical understanding of Parameter-Efficient Fine-Tuning (PEFT) mechanics—specifically LoRA/QLoRA low-rank decomposition—alongside instruction tuning and model alignment methodologies (DPO, RLHF)

  • Mastery of advanced prompt engineering, including structured output forcing, Chain-of-Thought, system prompt design, and building resilient multi-agent coordination systems

  • Hands-on experience with inference optimization and high-throughput serving, including quantization (GPTQ, AWQ), speculative decoding, prompt caching, and vLLM acceleration


MLOps & Deployment


  • Experience with MLOps practices, logging, and model registries (MLflow, Weights & Biases, DVC) along with model serving via Triton Inference Server or FastAPI

  • Production experience deploying, scaling, and monitoring models natively on cloud AI platforms, specifically AWS SageMaker or GCP Vertex AI

  • Experience building CI/CD pipelines for ML applications, with an emphasis on data and model versioning/lineage using DVC and Git


Data Engineering & Databases


  • Solid understanding of SQL, including advanced concepts like windowing functions and query optimization

  • Experience building and orchestrating data pipelines using Airflow or Prefect to feed specialized SQL and NoSQL vector/feature databases


Soft Skills & Professional Attributes


  • Strong critical thinking and problem-solving skills

  • Excellent written and verbal communication abilities

  • High degree of flexibility and adaptability to stay ahead of the rapid velocity of the open-source AI ecosystem (Hugging Face, vLLM, etc.)

  • Practical understanding of AI safety, compliance, guardrails (e.g., NeMo Guardrails or Llama Guard), and responsible AI practices.


Nice-to-Have


  • Contributions to open-source AI/ML repositories or GenAI orchestration frameworks.

  • Experience with real-time streaming data processing

  • Active participation in competitive ML spaces (Kaggle) or track record of reviewing/reproducing state-of-the-art AI research papers

  • Published research papers or conference presentations

  • Experience building and scaling Graph Databases or Knowledge Graphs for advanced RAG

  • Experience with multimodal architectures (Vision-Language models, Audio processing)


Qualifications


  • AI-ML Engineer: 2–5 years of hands-on experience engineering machine learning systems, optimizing infrastructure, and implementing LLM/GenAI workflows in production

  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, Statistics, or a related highly quantitative field

  • Demonstrated commitment to continuous learning through contributing to open-source, certifications, or self-study (especially in deep learning internals, MLOps, and modern GenAI frameworks)


What We Offer


  • Competitive salary commensurate with experience

  • Opportunity to work on diverse, cutting-edge AI/ML projects

  • Collaborative and innovation-driven work environment

  • Rapid growth and continuous learning opportunities

  • Exposure to latest AI technologies and industry best practices

Apply Now To Get Your Job

bottom of page