Own production reliability by defining SLIs/SLOs, managing error budgets, building zero-downtime deployment pipelines, and codifying infrastructure with Terraform and Ansible. Monitor systems using observability platforms, lead incident response and root-cause analysis, optimize capacity and costs, automate operations with scripting and AI tools, and participate in shared on-call rotations. The role requires cloud, Kubernetes, CI/CD, Linux, networking, distributed-systems, and high-availability production experience.
Production doesn't page a ticket queue. It pages you.
Techdome runs live infrastructure for Healthcare, FinTech, AI, and SaaS products where downtime isn't an inconvenience — it's a compliance incident or a lost transaction. We need someone who treats uptime as a personal metric, not a team KPI.
The job, in outcomes:
- Availability is your scoreboard. You define SLIs/SLOs, own the error budget, and make the call on when to slow down shipping to protect it.
- Deploys don't cause incidents. You build CI/CD that ships Blue-Green, Canary, and Rolling releases with zero customer-facing downtime — and you built at least one of these pipelines from scratch before, not just configured someone else's.
- Infrastructure is code, not tribal knowledge. Terraform and Ansible define your environments; nothing gets clicked into existence in a console.
- You see problems before customers do. Prometheus, Grafana, ELK, Datadog, OpenTelemetry — instrumented well enough that alerts mean something and noise doesn't.
- Incidents end with a fix, not a Slack thread. You lead response, drive RCA, and turn postmortems into action items that actually close.
- Cost is engineered, not just monitored. You right-size and forecast capacity instead of reacting to the bill.
- You use AI to move faster, not to look modern. Alert triage, incident summarization, automated runbooks — if it can be scripted or delegated to a model, it should be.
- On-call is shared, not survived. You rotate in, and you're the person newer engineers want on the call at 2am.
What gets you in the door:
- 2+ years running production as an SRE, DevOps, Platform, or Cloud Engineer — real ownership, not observer status
- Production-grade AWS, Azure, or GCP experience
- Docker/Kubernetes under actual load, with the scars to prove it
- Terraform and Ansible (or equivalent IaC) in daily use
- Built CI/CD pipelines from zero — Jenkins, GitHub Actions, GitLab CI, or similar
- Solid Linux, networking, and distributed-systems fundamentals
- Python, Go, or Bash for scripting and automation
- Shipped Blue-Green, Canary, and Rolling deployments in production, not just in theory
- Domain background in FinTech, Payments, Healthcare, or another high-availability environment
- AI tools (Copilot, Claude, Cursor, ChatGPT) are already part of how you work, not a novelty
What sets you apart:
- You've shipped AI-powered ops tooling — triage bots, auto-summarized incidents, anomaly detection that actually fires correctly
- You talk SLOs, error budgets, and chaos engineering like it's your first language, because it is
Why this seat, specifically:
You're not the 15th hire on a platform team waiting for tickets. You report close to founders and senior engineering leadership, your infra decisions ship the same week you make them, and the systems you protect handle real healthcare data and real money — not staging environments.
Similar Jobs
Financial Services
Lead reliability engineering for mission-critical network services, owning availability, performance, recoverability, incident response, root-cause analysis, and durable remediation. Architect automation using Python, Shell, and Ansible; provide technical leadership across SD-WAN, software-defined networking, routing, switching, firewalls, load balancers, and proxies. Improve observability, resilience testing, change safety, and operational readiness while mentoring engineers and applying secure, auditable AI-assisted workflows.
Top Skills:
Ai-Assisted Reliability WorkflowsAnsibleCi/CdCisco Aci/FabricsFirewallsLoad BalancersProxiesPythonRoutingSd-WanSdaShellSoftware-Defined NetworkingSwitchingVmware Nsx
Financial Services
Owns operational reliability for enterprise network services, including troubleshooting, incident and problem management, RCA closure, and major incident support. Builds production automation with Python, Shell, and Ansible; supports software-defined networking, routing, switching, firewalls, load balancers, and proxies. Improves observability, alert quality, runbooks, standardization, and resilience while partnering with development and platform teams. Uses validated, security-aware enterprise AI capabilities to accelerate incident triage and identify recurring reliability risks.
Top Skills:
AlertingAnsibleCisco AciCisco SdaDashboardsEnterprise Routing And SwitchingFirewallsLoad BalancersNetwork ObservabilityProxiesPythonSd-WanShellSnd
Financial Services
Design and maintain reliable, scalable, and secure AWS infrastructure using Terraform and automation scripts. Manage monitoring, logging, alerting, CI/CD, container orchestration, incident response, and performance optimization. Collaborate with development, operations, and security teams, participate in on-call rotations, enforce cloud security practices, document systems, and responsibly use AI-assisted engineering tools.
Top Skills:
Ai-Assisted Software Development ToolsAlertingAWSAws LambdaCi/CdDockerInfrastructure As CodeJavaKubernetesLoggingMonitoringPythonRubyShellTerraform
What you need to know about the Hyderabad Tech Scene
Because of its proximity to leading research institutions and a government committed to the city's growth, Hyderabad's tech scene is booming. With plans to establish India's first "AI city," the city is on track to become one of the world's most anticipated tech hubs, with companies like TransUnion, Schrödinger and Freshworks, among others, already calling the city home.

