Remote | Machine Learning & NLP Research Specialist
24-MAG
Location
πΊπΈ New York, United States
Type
full_time
Salary
Undisclosed
Posted
1mo ago
Job Description
We are sharing a specialised part-time consulting opportunity for US-based machine learning and natural language processing professionals with hands-on experience in Python, model training and evaluation, transformers, large language models, retrieval systems, and applied ML pipelines. This role supports an advanced AI initiative focused on identifying reasoning and capability gaps in frontier models. Selected professionals will design challenging real-world ML and NLP tasks, develop executable reference solutions, evaluate model performance, and analyse failures across language understanding, generation, retrieval, training workflows, and applied machine learning systems.
Key Responsibilities
ML & NLP Task Development β’ Design challenging machine learning and natural language processing problems based on practical research or industry experience β’ Create tasks involving model training, evaluation, language understanding, generation, retrieval, or applied ML pipelines β’ Target specific reasoning, implementation, and capability gaps in advanced AI models β’ Define clear specifications, expected behaviour, constraints, datasets, and evaluation criteria Reference Solutions & Python Development β’ Develop accurate reference solutions and supporting materials using Python β’ Integrate tasks into agent-based development and evaluation environments β’ Create executable tests, validation scripts, scoring logic, or model pipelines where appropriate β’ Ensure reference implementations are technically sound, reproducible, and appropriately challenging Model Evaluation & Failure Analysis β’ Evaluate model and agent performance across assigned ML and NLP tasks β’ Compare generated outputs with reference solutions and expected results β’ Identify tasks where models demonstrate meaningful limitations or inconsistent behaviour β’ Classify failures involving reasoning, implementation, retrieval, language understanding, generation, or instruction adherence β’ Document findings through clear and technically detailed written analysis Quality Calibration & Collaboration β’ Review tasks and evaluation methods developed by other machine learning specialists β’ Maintain consistent standards for difficulty, accuracy, realism, and technical quality β’ Participate in calibration and peer-review activities β’ Collaborate with other subject matter experts to improve evaluation coverage and reliability Ideal Profile Strong candidates may have: β’ Deep hands-on experience in machine learning, natural language processing, or both β’ Practical proficiency in Python demonstrated through professional, academic, or open-source work β’ Strong understanding of modern ML and NLP methods, including transformers and large language models β’ Experience with model training, fine-tuning, evaluation, retrieval, or production ML pipelines β’ Familiarity with PyTorch, TensorFlow, JAX, Hugging Face, or comparable frameworks and tooling β’ Ability to design realistic technical problems and develop complete reference solutions β’ Strong written communication and the ability to explain complex model behaviour clearly β’ Reliable availability for approximately 20 hours per week