Software Engineer III β AI/LLM Evaluation
Intelliswift - An LTTS Company
Location
πΊπΈ New York, United States
Type
contractor
Salary
$75β$85
Posted
1mo ago
Job Description
Job Title: Software Engineer III β AI/LLM Evaluation Location: New York, NY (100% Onsite) Job Type: Contract (12 Months, Extension Possible) We are seeking a Software Engineer III to join our clients cutting-edge AI research team focused on evaluating Large Language Models (LLMs) across the model development lifecycle. This role will support model assessment efforts spanning pretraining, mid-training, and post-training stages, helping researchers measure model quality, scalability, and performance. The ideal candidate has hands-on experience with LLMs, machine learning systems, Python, and PyTorch, along with a strong understanding of model evaluation methodologies.
Responsibilities
Execute and analyze evaluations for Large Language Models (LLMs) across different training stages. Track, organize, and maintain evaluation results across multiple models, benchmarks, and metrics. Port evaluation pipelines and benchmarks between frameworks while ensuring reproducibility and correctness. Develop and improve prompts used in model evaluation benchmarks. Analyze large volumes of evaluation data and provide actionable insights. Create reports, visualizations, and presentations to communicate evaluation findings. Collaborate with researchers and engineers to improve model assessment methodologies.
Required Qualifications
Bachelor's degree in Computer Science, Engineering, Mathematics, or a related technical field. 3β5 years of experience working with Machine Learning, Deep Learning, or Large Language Models. Hands-on experience with Python programming. Strong experience with PyTorch. Experience building and evaluating machine learning or deep learning systems. Understanding of model evaluation methodologies and benchmarking.
Preferred Qualifications
Experience with LLM pretraining and pretraining evaluation. Knowledge of perplexity metrics and their relationship to model performance. Familiarity with scaling laws and evaluating models across different parameter sizes. Experience working in AI research or advanced machine learning environments. Strong analytical and data visualization skills.