Data Science Research Assistant / Data Lake Engineer (m/f/d)
SIAW-HSG | Universität St. Gallen
Location
🇮🇳 Ajmer, India
Type
full_time
Salary
Undisclosed
Posted
1mo ago
Job Description
Swiss Institute for International Economics and Applied Economic Research (SIAW-HSG) Data Science Research Assistant / Data Lake Engineer (m/f/d) Data Science Research Assistant / Data Lake Engineer (m/f/d) 70 % (limited, by 01.08.2026) Your tasks
Responsibilities
And Project The position supports the development of a research data lake for empirical work with large-scale financial, textual, licensed, and partly confidential datasets. The objective is to build a robust, well-documented, and reproducible data infrastructure that allows researchers to ingest, store, process, document, and analyze data efficiently and securely. Research And Infrastructure Tasks Will Include Data Engineering, Coding, Documentation, And Coordination With Researchers And IT/platform Providers. Core Tasks Include, Among Others Design and implementation of the research data lake • Support the design of a scalable data architecture for approximately 5 TB of research data • Structure data into raw, cleaned, and analysis-ready layers • Develop clear naming conventions, folder structures, access rules, and documentation standards • Ensure that the data lake supports long-term retention of raw and processed data Data ingestion and integration • Build automated workflows to import data from external providers, databases, APIs, file deliveries, and researcher-maintained sources • Integrate financial datasets, textual datasets, and other licensed research data into a consistent infrastructure • Implement validation checks, logging, error handling, and version control for data updates • Document data provenance, licenses, update frequencies, and usage restrictions Automation of research pipelines • Develop reproducible pipelines for cleaning, transforming, and preparing datasets for empirical research • Create reusable scripts and templates for recurring data tasks • Support researchers in converting manual data work into automated and documented workflows • Contribute to reproducible research practices through Git-based code management and clear pipeline documentation Data governance, confidentiality, and access management • Help implement procedures for handling licensed and confidential datasets • Support role-based access concepts, documentation of data permissions, and compliance with provider agreements • Prepare data inventories and metadata files to make datasets findable and usable by the research team • Coordinate with internal IT or external platform providers where needed Research support • Assist researchers with data preparation, quality checks, exploratory analysis, and technical troubleshooting • Provide documentation and short internal guides so that the infrastructure can be maintained beyond the initial project phase • Contribute to other data-intensive research projects at the Chair or Institute where appropriate The position is particularly suitable for a candidate who wants to combine data science, data engineering, and applied academic research.