Owns operational reliability for enterprise network services, including troubleshooting, incident and problem management, RCA closure, and major incident support. Builds production automation with Python, Shell, and Ansible; supports software-defined networking, routing, switching, firewalls, load balancers, and proxies. Improves observability, alert quality, runbooks, standardization, and resilience while partnering with development and platform teams. Uses validated, security-aware enterprise AI capabilities to accelerate incident triage and identify recurring reliability risks.
There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.
As a Site Reliability Engineer III at JPMorgan Chase within the Infrastructure Platforms team, you will solve complex and broad business problems with simple and straightforward solutions. Network SRE who owns troubleshooting and reliability improvements across network platforms. Leads problem management for recurring issues, drives automation-first operations, and partners with development teams to improve observability, alert quality, and resilience.
As a Site Reliability Engineer III at JPMorgan Chase within the Infrastructure Platforms team, you will solve complex and broad business problems with simple and straightforward solutions. Network SRE who owns troubleshooting and reliability improvements across network platforms. Leads problem management for recurring issues, drives automation-first operations, and partners with development teams to improve observability, alert quality, and resilience.
Job responsibilities
- Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate, supporting adoption of site reliability engineering best practices within your team
- Lead day-to-day operational ownership for network services, including complex troubleshooting and coordinated restoration.
- Drive incident and problem management by running structured investigations, producing high-quality RCAs, and ensuring corrective/preventive actions are delivered.
- Participate in major incident management, providing communications support, technical lead support, and mitigation execution.
- Design and implement production-grade automation using Python, Shell, and Ansible (e.g., drift detection, change validation, automated diagnostics, safe rollout helpers).
- Engineer and support software-defined networking capabilities, including SD-WAN, SDA, and broader SND.
- Engineer and support routing and switching across enterprise networks.
- Engineer and support security and L4–L7 network components, including firewalls, load balancers, and proxies.
- Improve reliability through standardization, guardrails, repeatable runbooks, continuous validation, and observability (dashboards, high-signal alerting, service health metrics) in partnership with developers/platform teams.
- Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
- Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 3+ years applied experience
- Demonstrated experience in incident response and problem management, including end-to-end ownership of RCAs through closure.
- Strong hands-on networking skills across enterprise routing/switching and security/L4–L7 components.
- Strong automation capability using Python, Shell, and Ansible in production operations.
- SRE mindset with practical understanding of reliability concepts, NFRs, and risk analysis approaches (including familiarity with FMEA or equivalent methods).
- Ability to work independently, prioritize effectively, and deliver with minimal oversight.
- Experience supporting software-defined networking environments (e.g., SD-WAN, SDA, and related tooling).
- Ability to build and operationalize monitoring/observability, including dashboards, alerting, and service health metrics.
- Strong communication and coordination skills during high-severity incidents and cross-team restoration efforts.
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
- Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
Preferred qualifications, capabilities, and skills
- Demonstrate experience with Cisco ACI / fabrics.
- Hold relevant certifications such as CCNP (preferred), CCNA, or other vendor certifications.
- Work effectively in a financial institution or other regulated environment.
JPMorganChase Hyderabad, Telangana, IND Office
JP Morgan Tower, Salarpuria Sattva Knowledge City, HITEC City, Raidurgam, Hyderabad, Telangana, India, 500081
Similar Jobs
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Designs, builds, operates, and evaluates AI-native agentic capabilities across the sales and order lifecycle. Responsibilities include creating autonomous workflows, authoring prompts and guardrails, testing probabilistic behavior, building conversational experiences, integrating APIs and business systems, ensuring production safety and reliability, and supervising AI coding agents. The role also requires full-stack delivery, operational ownership, cross-functional collaboration, mentoring, and accountability for quality, security, and customer-facing outcomes.
Top Skills:
Agent OrchestrationAutomated Test FrameworkCi/CdContainerizationEmbeddingsFlow DesignerFunction CallingGraphQLLarge Language Model ApisNow PlatformObservabilityRelational DatabasesRestRetrieval-Augmented GenerationTool CallingUi Builder
Fintech • Mobile • Payments • Software • Financial Services
Own end-to-end recruitment for AML teams in Hyderabad, scaling talent pipelines and delivering strong candidate and hiring-manager experiences. Build relationships with hiring teams, collaborate with sourcers, develop talent networks, and use hiring data to identify bottlenecks and optimize globally scalable recruitment processes. Partner with People teams on referrals, talent mapping, salary benchmarking, and related initiatives while supporting Wise’s inclusive culture.
Financial Services
Develops, debugs, and maintains secure ServiceNow HRSD software solutions within an agile corporate technology team. Responsibilities include configuring and scripting the ServiceNow platform, creating custom widgets, implementing business and client scripts, troubleshooting integrations, applying CI/CD and resiliency practices, and using authorized AI-assisted development tools while validating outputs for security and quality.
Top Skills:
Ci/CdHrsdNow AssistServicenowServicenow Paas
What you need to know about the Hyderabad Tech Scene
Because of its proximity to leading research institutions and a government committed to the city's growth, Hyderabad's tech scene is booming. With plans to establish India's first "AI city," the city is on track to become one of the world's most anticipated tech hubs, with companies like TransUnion, Schrödinger and Freshworks, among others, already calling the city home.


