LLM Engineering Expert – AI Evaluation
vraify
Location
🇺🇸 United States
Type
full_time
Salary
Undisclosed
Posted
3d ago
Job Description
About the role
Vraify is hiring experienced LLM Engineering Experts to create and validate challenging, simulation-based engineering design problems for evaluating advanced AI agents. This remote contractor opportunity is open only to candidates residing in the United States or Canada. You will design multi-constraint tasks, configure open-source simulation tools, analyze agent execution logs, diagnose reasoning and tool-use failures, and build objective automated graders across electrical, mechanical, aerospace, control systems, systems engineering, and robotics. Important experience note This is a senior specialist role requiring at least 8 years of directly relevant engineering experience. To help candidates focus on opportunities aligned with their background, applications that do not meet this minimum will be automatically screened out. Freshers and early-career applicants are therefore not eligible for this role, and we warmly encourage them to consider opportunities better suited to their current experience level.
Responsibilities
- Author original, self-contained engineering design tasks with competing constraints, explicit optimization targets, validated reference solutions, and objective autograders.
- Build, run, and validate problem environments using open-source simulation tools and custom Python test benches.
- Evaluate coding-agent outputs and execution logs across repeated trials to identify systemic failure modes.
- Refine task difficulty using empirical model-performance data without introducing ambiguity or missing information.
- Collaborate with AI researchers, pod leads, and domain experts to integrate rigorous benchmarks into the model-evaluation pipeline.