Founding Computer Vision & Edge AI Engineer- C++ / CUDA / TensorRT (India)
APEAF
Location
🇮🇳 India
Type
full_time
Salary
Undisclosed
Posted
2w ago
Job Description
Position Title: Real-Time Computer Vision & Edge AI Engineer (Founding Engineering Team / Core LLD) Reporting Structure: High-Level AI Architect (Principal ML Scientist, Google) Domain: Sub-16ms Edge AI, 3D Pose & Shape Estimation (SMPL-X), TensorRT C++ Inference, Zero-Copy Systems Performance Benchmark: Hard locked 60 FPS (<16.6 ms total frame budget) on dedicated RTX hardware 1. Position
Overview
& Architecture We are building a proprietary, ultra-low-latency spatial computing platform centered on high-fidelity 100% 3D Digital Twin architecture and real-time human digitization. In this role, you will serve as the Low-Level Design (LLD) Core AI Engineer, working directly alongside a Lead AI Scientist from Google. Your primary mandate is to solve complex surface occlusion and volumetric estimation challenges by building an ultra-fast C++ inference pipeline. This system must accurately regress a subject's true underlying 3D body shape and skeletal pose directly from a live camera feed. You will deploy models that extract parametric data (SMPL-X shape/pose parameters) and bridge these joint rotations seamlessly into our Vulkan graphics engine via shared GPU memory. System Architecture: Hardware Camera Ingestion: (V4L2 / GStreamer / CUDA) Raw RGB Frames (Zero CPU Copy) Edge AI Inference: (TensorRT / ONNX C++ API for 3D Pose Tracking, Kinematic Anchoring, SMPL-X Shape) 3D Skeletal Transforms & Shape Parameters Zero-Copy Shared Memory: (CUDA-Vulkan Bridge feeding directly into OpenRigLogic / MetaHuman Engine) 2.
Key Responsibilities
& Deliverables A. Real-Time 3D Pose & Shape Estimation Deploy and optimize state-of-the-art 3D human body reconstruction models (e.g., Shapy, SMPLify-X, CLIFF) to accurately regress the user's underlying skeletal structure and body volume, effectively bypassing unpredictable surface topologies and complex environmental occlusions. Extract mathematically stable shape parameters () and pose parameters () to drive the skeletal hierarchy of a high-fidelity digital avatar. B. Edge Inference Pipeline (TensorRT) Translate Python-based research models into production-grade C++ inference engines using NVIDIA TensorRT and ONNX Runtime. Implement INT8/FP16 quantization, layer fusion, and custom CUDA plugins to ensure the entire AI inference pass executes within a strict <10 ms budget per frame. C. Temporal Smoothing & Anti-Jitter Kinematics Implement highly optimized temporal filters (Kalman filters, One-Euro filters, optical flow tracking) in native C++ to eliminate all high-frequency jitter from the output joint rotations before they reach the graphics engine. Ensure kinematic constraints (e.g., fixed bone lengths) are strictly maintained to prevent the digital asset from stretching or warping dynamically. D. Zero-Copy Ingestion & Engine Synchronization Build hardware-accelerated video capture pipelines using V4L2 or GStreamer to ingest raw camera frames directly into GPU memory. Bridge the output coordinate data and transformation matrices to the graphics team using POSIX shared memory and CUDA-Vulkan interop (VK_KHR_external_memory_fd), eliminating CPU staging overhead. 3. Technical