Disney Entertainment and ESPN Product & Technology is a global organization of engineers, product developers, designers, technologists, data scientists, and more – working to build and advance the technological backbone for Disney’s media business worldwide.
Product Engineering is a unified team responsible for the engineering of Disney Entertainment & ESPN digital and streaming products and platforms, including product engineering, media engineering, quality assurance, personalization, commerce, lifecycle, and identity.
The Observability & Insights group ensures Disney Streaming’s distributed systems are reliable, performant, and transparent. They build telemetry, dashboards, alerting, insights pipelines, and developer experience tooling that enable engineers to understand system health and act quickly.
Job Summary
As a Software Engineer II, you will design and build intelligent, AI‑driven systems that enhance the reliability and performance of Disney’s large‑scale streaming ecosystem. You will develop agentic systems, machine learning models, and real‑time pipelines that transform telemetry, logs, and user signals into automated detection, root‑cause analysis, and proactive insights.
In this role, you will contribute to autonomous agents capable of reasoning over complex system behavior, identifying real‑time issues, and driving faster detection and resolution across Disney+, Hulu, and ESPN. You will partner with engineering, product, and platform teams to embed intelligence directly into operational workflows, improving system resilience and customer experience at scale.
You will deliver high‑quality features end‑to‑end, contribute to system design and code reviews, and own components of production systems within a fast‑paced, AI‑native engineering environment.
Responsibilities
Design and operate production‑grade systems that leverage real‑time signals and AI‑driven detection to improve the health of streaming platforms, critical services, and customer experience.
Build and scale AI‑driven capabilities, including agentic AI systems powered by foundation models (e.g., Claude Opus/Sonnet, GPT‑4) to enable automated reasoning and decisioning, predictive modeling, and anomaly detection for real‑time system health and reliability.
Develop end‑to‑end data and decision‑making pipelines that transform telemetry, logs, and user signals into actionable insights, automated detection, and root‑cause analysis.
Create and deploy scalable APIs and services that deliver predictive signals, explainability, and insights to engineering teams, operational tools, and product stakeholders.
Partner cross‑functionally to embed intelligence into workflows (incident response, release validation, customer insights), reducing operational overhead and improving speed.
Drive innovation in observability, reliability, and developer productivity through applied AI and new approaches.
Basic Qualifications
Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
3+ years of backend development experience, including building AI‑powered or data‑driven applications and scalable APIs (e.g., FastAPI, Flask).
Practical experience in AI/ML engineering, with knowledge in at least one of the following areas:
Agentic Workflows: orchestrating foundation models (GPT‑4, Claude) using frameworks like LangChain or LangGraph.
Traditional ML: developing, training, or fine‑tuning models using PyTorch or TensorFlow.
Proficiency with AI‑assisted development tools (e.g., Cursor, Claude Code) to accelerate engineering velocity.
Experience with modern development practices, including version control (GitHub), containerization (Docker), and cloud‑native deployments (AWS/EKS).
Strong understanding of API design, microservices architecture, and standard SDLC workflows.
Strong analytical and technical skills to troubleshoot issues, perform rapid iteration, and generate viable solutions.
Excellent collaboration and communication skills, with the ability to work cross‑functionally and clearly explain complex technical concepts.
Preferred Qualifications
Experience with observability platforms (e.g., Datadog, Grafana, Conviva) and handling high‑volume telemetry data.
Familiarity with large‑scale data platforms and distributed data processing tools (e.g., PySpark, Pandas, Databricks, Snowflake).
Knowledge of prompt design, model evaluation, and fine‑tuning foundation models (GPT‑4, Claude).
Experience implementing production‑grade systems at scale within a fast‑paced, distributed environment.
The hiring range for this position in Glendale, CA is $117,500 to $157,500 per year and in New York, NY and Seattle, WA is $123,000 to $165,000 per year. The base pay will be determined by internal equity and may vary based on geographic region, job‑related knowledge, skills, and experience. A bonus and/or long‑term incentive units may be provided as part of the compensation package, in addition to full medical, financial, and other benefits depending on level and position offered.
#J-18808-Ljbffr