Lead Data Science Engineer
Dassault Systèmes
Location
🇺🇸 NY, United States
Type
full_time
Salary
Undisclosed
Posted
1mo ago
Job Description
Location: New York, Hybrid Medidata follows a hybrid office policy in which employees who are hired for an in-person position are expected to work on site a certain number of days per week in accordance with Company policy. About our Company: Medidata is powering smarter treatments and healthier people through digital solutions to support clinical trials. Celebrating 25 years of ground-breaking technological innovation across more than 36,000 trials and 11 million patients, Medidata offers industry-leading expertise, analytics-powered insights, and one of the largest clinical trial data sets in the industry. More than 1 million users trust Medidata's seamless, end-to-end platform to improve patient experiences, accelerate clinical breakthroughs, and bring therapies to market faster. Discover more at www.medidata.com. Our Team: Medidata is looking for individuals who will help us tackle some of the most complex questions facing the industry today using our proprietary platform and advanced analytics. At Medidata, we never work alone. This role will partner heavily with all of the key stakeholder functions including product, delivery, data science, engineering, partnerships, and biostatistics. Successful Medidata AI candidates will be skilled in analytical/quantitative thinking, structured communication, and excited about building the next horizon of Medidata's mission to power smarter treatments and healthier people. You will be reporting to Director, Data Engineering.
Responsibilities
: - Apply advanced skills in data architecture, data science engineering, data modeling, and data quality using modern cloud-native technologies. - Develop ETL pipelines, working with vector databases, automation, and CI/CD using tools such as Python, SQL, and Git. - Established MCP governance standards for agent interactions, including authentication, authorization, audit logging, context management, and compliance controls. - Develop LLM applications using Retrieval-Augmented Generation (RAG) and support fine-tuning for domain-specific tasks. - Analyze and manipulate both structured and unstructured data sources, ensuring high data quality and readiness for downstream consumers. - Built agent-driven metadata management solutions to maintain data catalogs, business glossaries, lineage documentation, and governance policies. - Document and communicate technical work clearly to stakeholders at all levels, both technical and non-technical. - Collaborate effectively in Agile environments and cross-functional teams, building secure, scalable data pipelines into Snowflake from both on-premise and cloud-based sources.