Blue Yonder Logo

Blue Yonder

Sr Technical Consultant - Azure Cloud, DevOps, Sql & Site Reliability Engineering(SRE)

Posted 13 Days Ago
Be an Early Applicant
In-Office
Hyderabad, Telangana, IND
Senior level
In-Office
Hyderabad, Telangana, IND
Senior level
Own L2/L3 Azure cloud platform incidents, troubleshooting Kubernetes, containers, networking, observability, storage, databases, deployments, provisioning, and autoscaling. Develop automation, maintain operational documentation, support releases and disaster recovery, conduct root-cause analysis, and participate in production on-call rotations and incident reviews.
The summary above was generated by AI

Overview:

  • The BYX Platform provides shared cloud provisioning, compute, storage, observability and runtime capabilities for enterprise applications and implementation teams.
  • We are seeking an astute Cloud Platform with a strong technical foundation and hands-on experience in Microsoft Azure, Kubernetes, containers, observability, networking and production automation.

Our current technical environment:

  • Cloud and Compute: Microsoft Azure, Kubernetes, node pools, container workloads, Azure Container Registry, KEDA, Dask, autoscaling, CPU, memory and GPU capacity.
  • Observability: Logs, metrics and traces; collectors and exporters; Elastic/Elasticsearch, Logstash, Kibana, dashboards, alerts and correlation identifiers.
  • Data and Storage: MongoDB/Atlas, SQL, Redis, NFS, managed storage, replicas, regional capacity and service quotas.
  • Networking and Security: DNS, CIDR, private endpoints, firewall rules, allowlists, TLS/certificates, SPNs, API keys, tokens, image-pull credentials and Git-hosted secrets.
  • Operations and Automation: Linux, Git, CI/CD deployment workflows, infrastructure automation and scripting using Bash, Python or PowerShell.

What you’ll do:

  • Review and act on incidents, service requests, infrastructure requests and provisioning failures logged by implementation teams and platform users.
  • Own L2/L3 cloud-platform issues from initial triage through recovery, validation, communication, root-cause analysis and closure.
  • Troubleshoot Kubernetes pods, deployments, replicas, services, events, health checks, node pools, scheduling and resource constraints.
  • Diagnose container image-pull, startup, shutdown, registry-authentication, rollout and workload-reconciliation failures.
  • Investigate CPU, memory, GPU, quota, region, placement and capacity issues affecting platform workloads.
  • Troubleshoot workload and event-driven autoscaling using Kubernetes metrics and technologies such as KEDA.
  • Support Azure resource provisioning, provider operations, resource lifecycle workflows and reconciliation between desired and actual state.
  • Diagnose ACR, private endpoint, firewall, allowlist, CIDR, DNS, TLS and runtime-connectivity problems.
  • Trace logs, metrics and distributed telemetry from the workload through collectors, exporters and observability ingestion pipelines.
  • Support Elastic/Logstash/Kibana ingestion, index mappings, access, dashboards, alerts and environment filters.
  • Investigate operational issues involving MongoDB/Atlas, SQL, Redis, NFS and related managed data or storage services.
  • Support certificates, service principals, API keys, image-pull credentials and infrastructure credential rotation.
  • Support regional releases, environment configuration, disaster-recovery workflows and post-deployment validation.
  • Develop automation to improve platform reliability, reduce manual provisioning and shorten incident recovery time.
  • Maintain runbooks, dashboards, alerts, known-error records and diagnostic procedures for use by platform support teams.
  • Participate in capacity planning, incident reviews, change reviews, release readiness and an agreed production on-call rotation.

What we are looking for:

  • Minimum 5-10 years of relevant work experience in Azure cloud infrastructure, DevOps, site reliability engineering or platform operations.
  • This role includes Rotational Shifts(Night shifts of 2 months in Year)
  • Hands-on production experience with Microsoft Azure services and operational troubleshooting
  • Strong Kubernetes troubleshooting skills covering workloads, events, services, networking, scheduling, scaling and resource management.
  • Experience with a container registry; Azure Container Registry experience is highly relevant.
  • Practical observability experience across logs, metrics, traces, dashboards and alerts.
  • Experience with Elastic Stack components—Elasticsearch, Logstash and Kibana—or a closely comparable platform.
  • Working knowledge of DNS, TLS/certificates, CIDR, firewalls, proxies, load balancing and private networking.
  • Strong Linux administration, application log analysis and production incident-troubleshooting skills.
  • Scripting experience with Bash, Python or PowerShell for diagnostics and operational automation.
  • Experience troubleshooting CI/CD deployments, configuration changes and failed rollouts.
  • Working knowledge of Git, service identities and secrets-management fundamentals.
  • Working knowledge of MongoDB/Atlas, Redis, SQL or NFS operations is preferred.
  • Experience with infrastructure as code such as Terraform, Bicep or ARM templates and Kubernetes packaging such as Helm is preferred.
  • Experience with incident response, Root Cause Analysis, post-incident reviews and controlled production changes.
  • Strong collaboration and communication skills with the ability to work across application, security, networking and vendor teams.
  • Willingness to participate in a scheduled production on-call rotation and planned out-of-hours changes when required.

Our Values

If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success – and the success of our customers. Does your heart beat like ours? Find out here: Core Values

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.

Blue Yonder Hyderabad, Telangana, IND Office

14, 15 & 16th Floors Unit 3, Parcel 4, Octave Block Salarpuria Sattva Knowledge City Hyderabad , Hyderabad, India, 50008

Blue Yonder Hyderabad, Telangana, IND Office

14, 15 & 16th floors, Parcel 4, Unit 3, Octave Block, Salarpuria Sattva Knowledge City, Madhaphur, , Hyderabad, India, 500081

Similar Jobs

56 Minutes Ago
Hybrid
Hyderabad, Telangana, IND
Senior level
Senior level
Financial Services
Leads and coaches multiple software engineering teams while overseeing resources, budgets, technical execution, and operating practices. Drives enterprise adoption of AI-assisted engineering and SDLC automation with governance, security, quality, and measurable delivery outcomes. Guides Java backend development and AWS cloud-native distributed services, participates in architecture and design reviews, promotes testing, CI/CD, observability, and secure coding, and collaborates with product and business stakeholders.
Top Skills: Apache AirflowAurora PostgresqlAWSAws GlueCassandraCi/CdCloud FoundryDockerDomain-Driven DesignEc2Ecs/FargateEvent-Driven ArchitectureGithub CopilotJavaJava MessagingKafkaKubernetesMicroservicesNoSQLOracle RdbmsRestSoapSpringSpring BootSQLTest-Driven Development
An Hour Ago
Hybrid
Hyderabad, Telangana, IND
Entry level
Entry level
Fintech • Financial Services
Support trade services by processing and amending letters of credit, issuing and advising collections, and handling complex documentation. Collaborate with peers and managers to resolve issues, retain clients, identify growth referrals, and follow trade services policies and procedures.
Yesterday
In-Office
Hyderabad, Telangana, IND
Expert/Leader
Expert/Leader
Artificial Intelligence • Hardware • Information Technology • Machine Learning
Own verification and end-to-end analysis of full-chip gate-level DRAM designs. Develop SystemVerilog and UVM testbenches, test cases, stimuli, test vectors, regressions, and coverage improvements. Simulate, debug, and analyze pre-silicon block and full-chip designs using digital, mixed-signal, and SPICE tools. Apply AI-based automation to accelerate verification while validating outputs against specifications, simulations, waveforms, assertions, and coverage. Guide newer engineers and collaborate globally on verification methodologies and environments.
Top Skills: AmsCmosDramFinsimHspicePerlPliPythonReal Number ModelsSimvisionSpiceSramSystemverilogUvmVerilogVirtuosoVsimWaveviewXcelium

What you need to know about the Hyderabad Tech Scene

Because of its proximity to leading research institutions and a government committed to the city's growth, Hyderabad's tech scene is booming. With plans to establish India's first "AI city," the city is on track to become one of the world's most anticipated tech hubs, with companies like TransUnion, Schrödinger and Freshworks, among others, already calling the city home.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account