Senior Lead Software Engineer - ML Engineer for Agent Platform Be an integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch technology products.
As a Senior Lead Software Engineer - ML Engineer for Agent Platform at JPMorgan Chase within the Commercial and Investment Banking – Data Analytics Payments Team, you are a senior technical leader on the team that builds and runs NEO, the firm's agent runtime platform for Payments Technology. You set the architecture for how agents execute, communicate, remember, and get evaluated in a secure, stable, and scalable way. As a core technical contributor and technical direction-setter, you are responsible for the hardest technology decisions across multiple technical areas in support of the firm's business objectives, and for raising the engineering bar across the teams that build on NEO.
Job Responsibilities
Owns end-to-end architecture of the NEO agent runtime, including secure execution and isolation (micro-VMs such as Firecracker/Kata), agent-to-agent (A2A) communication, Model Context Protocol (MCP) tooling, the memory layer, and evaluation infrastructure
Executes creative software solutions, design, development, and technical troubleshooting with the ability to think beyond routine or conventional approaches to build solutions or break down technical problems
Develops secure and high-quality production code, and reviews and debugs code written by others; sets code and design standards adopted across teams building on NEO
Drives team adoption of enterprise-authorized AI-assisted engineering practices to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team
Applies and shapes the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation
Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of the platform and the agents running on it
Defines the permission-aware, auditable execution model for the runtime, including fine-grained authorization (OpenFGA) and runtime policy (OPA/Rego), so agents operate safely in a regulated environment
Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture
Leads communities of practice across Software Engineering to drive awareness and use of new and leading-edge technologies, and mentors Lead and senior engineers
Adds to team culture of diversity, opportunity, inclusion, and respect
Required Qualifications, Capabilities, and Skills
Formal training or certification on software engineering concepts and 5+ years applied experience
Hands-on practical experience delivering system design, application development, testing, and operational stability for platforms, runtimes, or distributed systems in production
Advanced in one or more programming language(s); strong Python plus a systems language (Go or Rust) for performance-sensitive runtime work
Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
Demonstrated experience building or operating LLM/agent systems in production, including tracing, evaluations, and guardrails
Proficient in all aspects of the Software Development Life Cycle
Advanced understanding of agile methodologies such as CI/CD, Application Resiliency, and Security
Demonstrated proficiency in software applications and technical processes within a technical discipline (e.g., cloud, artificial intelligence, machine learning)
In-depth knowledge of the financial services industry and their IT systems
Practical cloud native experience; production Kubernetes expected
Preferred Qualifications, Capabilities, and Skills
Experience architecting secure code execution and sandboxing with micro-VMs (Firecracker, Kata, gVisor) for multi-tenant isolation
Experience designing or implementing agent protocols (A2A, MCP) and multi-agent orchestration
Experience with agent or distributed memory systems — memory nodes, episodic/semantic memory, and graph-backed retrieval (Graph RAG)
Exposure to LLMs, RAG architectures, vector databases, and embedding-based retrieval systems
Fine-grained authorization (OpenFGA / Zanzibar-style) and policy engines (OPA/Rego)
Evaluation infrastructure for agents — offline/online evals, regression suites, and LLM-as-judge quality/safety gating in CI
Proficiency with Infrastructure as Code (Terraform) and containerized deployments (Docker, Kubernetes)