JPMorganChase Logo

JPMorganChase

Senior Lead Site Reliability Engineer - Cloud Data & Databricks

Posted 6 Days Ago
Be an Early Applicant
Hybrid
Hyderabad, Telangana, IND
Senior level
Hybrid
Hyderabad, Telangana, IND
Senior level
Leads SREs in designing, administering, and operating a managed AWS Databricks platform. Responsibilities include resilient multi-region architecture, observability, alerting, capacity planning, infrastructure automation, CI/CD, secure software development, incident response, troubleshooting, and postmortems. The role also establishes safe, auditable AI-assisted reliability workflows and collaborates with engineering, data, and vendor teams to improve platform scalability, performance, and operational excellence.
The summary above was generated by AI

Join us to shape the future of data and analytics, leveraging your expertise to deliver impactful technology solutions. Experience career growth and make a difference in a collaborative environment.


As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the AIML Data Platforms and Chief Data and Analytics Team, you will develop and deliver advanced technology products focused on data and analytics. You will tackle complex cloud data platform challenges, especially around Data Lake Tools, and work in an agile environment, collaborating with cross-functional teams. You will help drive the firm’s data and analytics journey, ensuring quality, integrity, and security of data, and leveraging AI/ML technologies to support commercial goals. You will contribute to a culture of innovation and operational excellence.

 

Job responsibilities

  • Lead a team of SREs to design, implement, and maintain a managed AWS Databricks platform, providing engineering and operational support to Application/Engineering teams
  • Perform platform design, set-up and configuration, workspace administration, and resource monitoring
  • Create multi-AZ, multi-region, and multi-cloud resiliency strategies for business-critical products and services
  • Lead evaluation sessions with external vendors, startups, and internal teams to assess architectural designs and technical credentials
  • Drive continuous improvement in system observability, alerting, and capacity planning
  • Collaborate with engineering and data teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence
  • Execute creative software solutions, design, development, and technical troubleshooting
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate reliability design and operational decisioning (e.g., incident/post-incident analysis and requirements traceability), validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., testing/validation automation and production readiness), ensuring traceability/auditability, resiliency, and security controls.
  • Develop secure, high-quality production code, and review/debug code written by others to ensure reliability and correctness.
  • Apply SRE best practices to improve reliability, scalability, and performance; eliminate or automate recurring issues; and maintain incident response procedures including root cause analysis and postmortems.

Required qualifications, capabilities and skills

  • Formal training or certification on software engineering concepts and 10+ years applied experience
  • Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and incident management
  • Experience with monitoring tools, automation frameworks, and CI/CD pipelines
  • Proficient in Python application program development with use of automated unit testing
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity.
  • Ability to set team practices for safe AI usage in operations (e.g., review/approval expectations and escalation paths) while maintaining resiliency, security, and auditability outcomes.
  • Experience with Terraform development and understanding of Terraform enterprise
  • Experience in delivering system design, application development, testing, and operational stability
  • Knowledge of Big Data distributed compute frameworks like Spark, Glue, MapReduce
  • Excellent troubleshooting, analytical, and communication skills

 

Preferred qualifications, capabilities and skills

  • Experience in Data pipelines using Spark
  • Exposure to AWS & Databricks Platform administration
  • Knowledge of containerization (Docker, Kubernetes) and orchestration
  • Familiarity with distributed systems and large-scale data processing
 

JPMorganChase Hyderabad, Telangana, IND Office

JP Morgan Tower, Salarpuria Sattva Knowledge City, HITEC City, Raidurgam, Hyderabad, Telangana, India, 500081

Similar Jobs

An Hour Ago
Hybrid
Hyderabad, Telangana, IND
Senior level
Senior level
Financial Services
Leads and coaches multiple software engineering teams while overseeing resources, budgets, technical execution, and operating practices. Drives enterprise adoption of AI-assisted engineering and SDLC automation with governance, security, quality, and measurable delivery outcomes. Guides Java backend development and AWS cloud-native distributed services, participates in architecture and design reviews, promotes testing, CI/CD, observability, and secure coding, and collaborates with product and business stakeholders.
Top Skills: Apache AirflowAurora PostgresqlAWSAws GlueCassandraCi/CdCloud FoundryDockerDomain-Driven DesignEc2Ecs/FargateEvent-Driven ArchitectureGithub CopilotJavaJava MessagingKafkaKubernetesMicroservicesNoSQLOracle RdbmsRestSoapSpringSpring BootSQLTest-Driven Development
An Hour Ago
Hybrid
Hyderabad, Telangana, IND
Entry level
Entry level
Fintech • Financial Services
Support trade services by processing and amending letters of credit, issuing and advising collections, and handling complex documentation. Collaborate with peers and managers to resolve issues, retain clients, identify growth referrals, and follow trade services policies and procedures.
Yesterday
In-Office
Hyderabad, Telangana, IND
Expert/Leader
Expert/Leader
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Own verification and end-to-end analysis of full-chip gate-level DRAM designs. Develop SystemVerilog and UVM testbenches, test cases, stimuli, test vectors, regressions, and coverage improvements. Simulate, debug, and analyze pre-silicon block and full-chip designs using digital, mixed-signal, and SPICE tools. Apply AI-based automation to accelerate verification while validating outputs against specifications, simulations, waveforms, assertions, and coverage. Guide newer engineers and collaborate globally on verification methodologies and environments.
Top Skills: AmsCmosDramFinsimHspicePerlPliPythonReal Number ModelsSimvisionSpiceSramSystemverilogUvmVerilogVirtuosoVsimWaveviewXcelium

What you need to know about the Hyderabad Tech Scene

Because of its proximity to leading research institutions and a government committed to the city's growth, Hyderabad's tech scene is booming. With plans to establish India's first "AI city," the city is on track to become one of the world's most anticipated tech hubs, with companies like TransUnion, Schrödinger and Freshworks, among others, already calling the city home.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account