Offensive AI Engineer – Frontier AI Security & Testing, VP
State Street
Location
🇺🇸 Berwyn, United States
Type
full_time
Salary
Undisclosed
Posted
1d ago
Job Description
Offensive AI Engineer – Frontier AI Security & Testing Who We Are Looking For State Street Global Cybersecurity (GCS) is looking for a hands-on Offensive AI Engineer to join Frontier AI Security & Testing (FAST). FAST is a new research and development team within Regulatory Assurance, Penetration Testing & Offensive Research. This role works at the intersection of frontier AI, cloud engineering, data engineering and offensive security. You will design, deploy and run the secure AWS-based platforms, evaluation harnesses and data pipelines that let State Street use frontier AI models safely as an offensive security capability. This is an engineering role, not an advisory or architecture-only one. You will build, configure, deploy, troubleshoot and automate. You will work directly with frontier models, agentic frameworks and enterprise APIs. The work also includes putting technical controls in place so that AI is used in an auditable, governed way in a highly regulated financial services environment. Why this role is important to us State Street is expanding its use of frontier AI across cybersecurity to strengthen resilience, speed up adversary emulation and improve how defenses are validated. FAST decides which models can be trusted, shows with evidence what they can do, builds the platform they run on, and turns model capability into tools that red team, penetration testing, purple team and threat hunt teams can use. As an Offensive AI Engineer, you will help build the execution, evaluation and control layer that lets these capabilities run safely at enterprise scale. What You Will Be Responsible For As an Offensive AI Engineer you will: AWS AI infrastructure and deployment • Design, engineer and deploy secure, isolated AWS environments to host and access frontier AI models. This includes Amazon Bedrock, SageMaker, private endpoint connectivity (PrivateLink/VPC endpoints), EKS/ECS, Lambda and API Gateway. • Build and maintain infrastructure-as-code (Terraform or CloudFormation) and CI/CD pipelines for AI workloads, with repeatable and auditable deployment patterns. • Engineer network segmentation, identity and access controls, secrets management and egress restrictions that create enforceable trust boundaries around agentic AI activity. Frontier model engineering and evaluation • Integrate with frontier model provider APIs and SDKs from major commercial vendors. Build model-agnostic abstraction layers that allow fast model onboarding and side-by-side comparison. • Design and build standardized evaluation harnesses and benchmarking frameworks. These should measure model capability, precision, false-positive rate, reliability, cost and failure modes across offensive security tasks in a repeatable way. • Develop agentic workflows, tool-use integrations and multi-agent orchestration patterns (e.g., MCP, LangGraph, Strands, or similar) that support reconnaissance, attack-path analysis and adversary simulation in controlled lab environments. • Stress-test models against known AI failure modes, including prompt injection, jailbreak susceptibility, data leakage, tool misuse, scope drift and unsafe emergent behavior. Data science and Databricks • Design and build Databricks pipelines (PySpark/SQL) to ingest, curate and analyze model telemetry, evaluation results and offensive security data sets. • Apply data science methods to evaluation results, including statistical comparison, scoring methodology, trend analysis and regression detection across model versions. • Build data products and dashboards that give leadership evidence of model performance, cost and risk. API integration • Build secure, scalable integrations between AI platforms and enterprise systems, security tools and data sources using REST APIs, event-driven architectures and tool-execution brokers. • Develop reusable API patterns and connectors for tool invocation, context injection and output validation. AI controls, logging and governance • Implement technical controls that limit agent autonomy where needed. These include guardrails, human-in-the-loop checkpoints, rate and scope limits, kill-switch mechanisms and least-privilege tool access. • Engineer full logging, monitoring and observability that capture agent inputs, outputs, tool calls and decision paths to support auditability and post-incident analysis. • Align AI use with enterprise Responsible AI, model risk management, legal and compliance