AI/ML Engineer
Capsitech IT Services Private Limited
Location
🇮🇳 Jodhpur, India
Type
full_time
Salary
Undisclosed
Posted
3w ago
Job Description
AI / ML Engineer Experience: 2–4 Years | Domain: AI/ML, GenAI, Speech & Vision | Location: Jodhpur, India
YOUR ROLE
AND
RESPONSIBILITIES
As an AI/ML Engineer (2–4 years of experience), you will be responsible for designing, building, and deploying scalable, production-grade AI applications across Generative AI, Speech AI, Computer Vision, and Document Intelligence. You will bridge the gap between AI research and full-scale deployment by creating asynchronous microservices, multi-agent automation systems, and robust voice/image processing pipelines. Additionally, you will work closely with cloud infrastructure and MLOps tooling to containerize models, monitor system performance, and ensure enterprise-level reliability. YOUR PRIMARY
RESPONSIBILITIES
WILL INCLUDE • Voice, Audio & Image Processing: Build and maintain workflows for speech transcription (using Whisper), speaker diarization, voice embedding generation, speaker verification, computer vision tracking/ monitoring, and OCR-based document intelligence. • Generative AI & Agentic Systems: Architect production-ready Retrieval-Augmented Generation (RAG) platforms, Model Context Protocol (MCP) integrations, and multi-agent coordination pipelines for complex business automation. • API Development & Microservices: Design asynchronous RESTful APIs using FastAPI and Pydantic to expose ML model pipelines, incorporating secure authentication, webhook callbacks, and cloud storage integration. • Document Intelligence: Construct automated data extraction systems using tools like PaddleOCR and others to convert unstructured files and invoices into structured JSON formats. • Performance Optimization: Evaluate model outputs, refine prompt structures, and optimize pipeline latency to prevent hallucinations and maintain enterprise compliance. REQUIRED TECHNICAL AND PROFESSIONAL EXPERTISE • Domain Expertise: Demonstrated proficiency across Voice, Audio & Image Processing (speech transcription, speaker diarization, voice verification/embeddings, computer vision tracking, and OCR document intelligence). • Core Technical Stack: Strong programming skills in Python, SQL, FastAPI, Pydantic, Pandas, and NumPy. • Machine Learning & LLMs: Hands-on experience with PyTorch or TensorFlow, Scikit-learn, embeddings, vector search, RAG architectures, prompt engineering, and foundation models (GPT-4o, Gemini, Claude, Hugging Face). • Frameworks & Agents: Experience working with orchestration toolkits such as LangChain, LangGraph, CrewAI, or FastMCP. •