Techdome Logo

Techdome

Site Reliability Engineer

Reposted 4 Days Ago
Be an Early Applicant
In-Office
Hyderabad, Telangana, IND
Junior
In-Office
Hyderabad, Telangana, IND
Junior
Own production reliability across cloud environments by managing infrastructure, building zero-downtime CI/CD pipelines, codifying infrastructure with Terraform and Ansible, and driving observability. Define SLIs, SLOs, and error budgets; lead incident response, root-cause analysis, and post-incident reviews; optimize cloud costs and capacity; automate AI-powered operational workflows; and participate in on-call rotations.
The summary above was generated by AI
If it's down, it's on you. If it stays up, that's on you too.

At Techdome, you own production for real Healthcare, FinTech, AI, and SaaS products — not a ticket queue. You'll build pipelines, own incidents, ship zero-downtime releases, and be the person the team trusts at 2am.

What you'll do
  • Keep production available, reliable, and performant across every environment.
  • Manage and optimize cloud environments (AWS, Azure, or GCP).
  • Build zero-downtime CI/CD pipelines using Blue-Green, Rolling, and Canary strategies.
  • Codify infrastructure with Terraform and Ansible.
  • Drive observability with Prometheus, Grafana, ELK, Datadog, and OpenTelemetry.
  • Define and own SLIs, SLOs, and error budgets.
  • Lead incident response, RCA, and post-incident reviews.
  • Optimize cloud cost and plan capacity.
  • Automate operational workflows, including AI-powered alert triage and incident summarization.
  • Join the on-call rotation.

What you bring
  • 2+ years as an SRE, DevOps Engineer, Platform Engineer, or Cloud Engineer.
  • Hands-on production experience with AWS, Azure, or GCP.
  • Docker and Kubernetes experience under real load.
  • Infrastructure-as-Code expertise (Terraform, Ansible, or equivalent).
  • CI/CD pipelines built from scratch (Jenkins, GitHub Actions, GitLab CI, or similar).
  • Strong Linux, networking, and distributed-systems fundamentals.
  • Scripting skills in Python, Go, or Bash.
  • Experience with Blue-Green, Canary, and Rolling deployments in production.
  • Background in FinTech, Payments, Healthcare, or another high-availability domain.
  • Regular use of AI tools (Copilot, Claude, Cursor, ChatGPT, or similar) to move faster.

Extra credit: You've built AI-powered ops workflows (monitoring, alert triage, incident summarization), and you speak fluent SRE — SLOs, error budgets, chaos engineering.

Why Techdome?
Work across AI, Healthcare, Payments, and SaaS products. Own critical infrastructure from day one. Work directly with founders and senior engineering leadership. Real stakes, fast decisions, real growth.

Similar Jobs

4 Days Ago
Hybrid
Hyderabad, Telangana, IND
Senior level
Senior level
Financial Services
Lead reliability engineering for mission-critical network services, owning availability, performance, recoverability, incident response, root-cause analysis, and durable remediation. Architect automation using Python, Shell, and Ansible; provide technical leadership across SD-WAN, software-defined networking, routing, switching, firewalls, load balancers, and proxies. Improve observability, resilience testing, change safety, and operational readiness while mentoring engineers and applying secure, auditable AI-assisted workflows.
Top Skills: Ai-Assisted Reliability WorkflowsAnsibleCi/CdCisco Aci/FabricsFirewallsLoad BalancersProxiesPythonRoutingSd-WanSdaShellSoftware-Defined NetworkingSwitchingVmware Nsx
4 Days Ago
Hybrid
Hyderabad, Telangana, IND
Mid level
Mid level
Financial Services
Owns operational reliability for enterprise network services, including troubleshooting, incident and problem management, RCA closure, and major incident support. Builds production automation with Python, Shell, and Ansible; supports software-defined networking, routing, switching, firewalls, load balancers, and proxies. Improves observability, alert quality, runbooks, standardization, and resilience while partnering with development and platform teams. Uses validated, security-aware enterprise AI capabilities to accelerate incident triage and identify recurring reliability risks.
Top Skills: AlertingAnsibleCisco AciCisco SdaDashboardsEnterprise Routing And SwitchingFirewallsLoad BalancersNetwork ObservabilityProxiesPythonSd-WanShellSnd
7 Days Ago
Hybrid
Hyderabad, Telangana, IND
Mid level
Mid level
Financial Services
Design and maintain reliable, scalable, and secure AWS infrastructure using Terraform and automation scripts. Manage monitoring, logging, alerting, CI/CD, container orchestration, incident response, and performance optimization. Collaborate with development, operations, and security teams, participate in on-call rotations, enforce cloud security practices, document systems, and responsibly use AI-assisted engineering tools.
Top Skills: Ai-Assisted Software Development ToolsAlertingAWSAws LambdaCi/CdDockerInfrastructure As CodeJavaKubernetesLoggingMonitoringPythonRubyShellTerraform

What you need to know about the Hyderabad Tech Scene

Because of its proximity to leading research institutions and a government committed to the city's growth, Hyderabad's tech scene is booming. With plans to establish India's first "AI city," the city is on track to become one of the world's most anticipated tech hubs, with companies like TransUnion, Schrödinger and Freshworks, among others, already calling the city home.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account