Senior Software Engineer, Machine Learning Infrastructure & Automation
fal
Location
🇺🇸 United States
Type
full_time
Salary
Undisclosed
Posted
1d ago
Job Description
This a Full Remote job, the offer is available from: United States fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products. As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About this role
: Help fal's ML team move faster by building the automation, infrastructure, and developer tooling that makes developing, testing, and deploying generative AI models seamless. You'll own and improve the CI/CD systems supporting our rapidly growing collection of ML models and inference pipelines. Your focus will be on eliminating manual work, accelerating development cycles, and building reliable systems that allow ML engineers to ship new models and optimizations with confidence. This is a high-impact engineering role where you'll work closely with our Applied ML and ML Performance teams. You'll build everything from automated model validation and performance benchmarking to AI-powered development workflows that help engineers iterate faster. The ideal candidate thinks beyond traditional CI/CD and sees automation as a force multiplier for the entire engineering organization. What you’ll do: • Own ML CI/CD infrastructure Design, build, and maintain automated testing, validation, and deployment pipelines for our ML models and inference services. • Accelerate development cycles.Dramatically reduce CI execution times through intelligent parallelization, caching, test selection, and efficient use of compute resources. • Build automated model validation.Develop systems that test model outputs, detect quality regressions, and validate changes across different models, GPU architectures, and configurations. • Automate performance benchmarking. Build continuous performance testing that detects regressions in inference latency, throughput, GPU utilization, and cost. • Build automated pricing and deployment checks. Ensure model pricing, billing configurations, API schemas, and deployments are validated automatically before reaching production. • Extend our agentic engineering workflows. Develop AI-powered automation and agentic coding systems that automatically diagnose CI failures, identify regressions, propose fixes, and streamline engineering workflows. • Improve deployment reliability. Build automated safeguards, deployment verification, rollback mechanisms, and monitoring to ensure new model releases are reliable. • Eliminate engineering toil. Identify repetitive tasks across the ML team and build tools and systems that automate them, allowing engineers to focus on developing new models and improving performance.