Own production reliability for high-availability healthcare, FinTech, AI, and SaaS systems. Responsibilities include defining SLOs and error budgets, building zero-downtime CI/CD pipelines, managing infrastructure with Terraform and Ansible, improving observability, leading incident response and postmortems, forecasting capacity, applying AI to operations, and participating in on-call rotations.
About the Role
Techdome runs live infrastructure for Healthcare, FinTech, AI, and SaaS products — environments where downtime isn't an inconvenience, it's a compliance incident or a lost transaction. We're hiring a Senior SRE who treats uptime as a personal metric, not a team KPI, and who wants direct, high-leverage ownership over production systems handling real money and real patient data — not staging environments. You'll report close to founders and senior engineering leadership, and your infrastructure decisions will ship the same week you make them.
Key Responsibilities
- Define SLIs/SLOs, own the error budget, and decide when to slow down shipping to protect it
- Build zero-downtime CI/CD pipelines supporting Blue-Green, Canary, and Rolling releases
- Manage all environments through Terraform and Ansible — no manual console changes
- Instrument systems with Prometheus, Grafana, ELK, Datadog, and OpenTelemetry so alerts are actionable, not noisy
- Lead incident response, drive root cause analysis, and ensure postmortem action items are closed
- Right-size infrastructure and forecast capacity proactively, rather than reacting to billing
- Apply AI to alert triage, incident summarization, and automated runbooks
- Participate in a shared on-call rotation as a dependable, trusted responder
Required Qualifications
- 2+ years of production ownership experience as an SRE, DevOps, Platform, or Cloud Engineer
- Hands-on, production-grade experience with AWS, Azure, or GCP
- Real-world Docker/Kubernetes experience under production load
- Daily use of Terraform and Ansible (or equivalent IaC tools)
- Experience building at least one CI/CD pipeline from scratch (Jenkins, GitHub Actions, GitLab CI, or similar)
- Strong fundamentals in Linux, networking, and distributed systems
- Proficiency in Python, Go, or Bash for automation and scripting
- Proven experience shipping Blue-Green, Canary, and Rolling deployments in production
- Prior work in FinTech, Payments, Healthcare, or another high-availability domain
- Practical, everyday use of AI tools such as Copilot, Claude, Cursor, or ChatGPT
Preferred Qualifications
- Experience building AI-powered operations tooling — triage bots, incident auto-summarization, or reliable anomaly detection
- Fluency in SLOs, error budgets, and chaos engineering as core practice, not theory
Why Join Techdome?
- Direct reporting line to founders and senior engineering leadership
- Infrastructure decisions ship the same week they're made — no bureaucratic delay
- You'll protect systems handling real healthcare data and real financial transactions
- A team that shares on-call and builds a culture where the SRE on call at 2am is someone people trust
Similar Jobs
Financial Services
Leads SREs in designing, administering, and operating a managed AWS Databricks platform. Responsibilities include resilient multi-region architecture, observability, alerting, capacity planning, infrastructure automation, CI/CD, secure software development, incident response, troubleshooting, and postmortems. The role also establishes safe, auditable AI-assisted reliability workflows and collaborates with engineering, data, and vendor teams to improve platform scalability, performance, and operational excellence.
Top Skills:
Ai/MlAutomation FrameworksAWSAws GlueCi/CdData LakeDatabricksDistributed SystemsDockerError BudgetsKubernetesMapreduceMonitoring ToolsPythonSlisSlosSparkSreTerraformTerraform Enterprise
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Establish SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems. Operate Kubernetes and GCP environments, implement Infrastructure as Code, optimize capacity and performance, and drive vulnerability remediation, penetration-testing support, and cloud platform security.
Top Skills:
Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonShellTerraform
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Supports reliability, availability, and performance of critical applications and platforms. Responsibilities include monitoring, incident response, observability, runbook maintenance, scripting and automation, root cause follow-up, postmortems, SLO and SLI adoption, alert improvement, and operational readiness. Collaborates with engineering, cloud, infrastructure, and application teams while using tools such as Kubernetes, Azure, dashboards, logs, and AI-assisted investigation systems.
Top Skills:
ApmAzure Application InsightsAzure DevopsAzure MonitorBashCi/CdDockerElastic/ElkGitGitGrafanaKubernetesLinuxAzurePowershellPrometheusPythonServicenowSplunkSQLTerraform
What you need to know about the Hyderabad Tech Scene
Because of its proximity to leading research institutions and a government committed to the city's growth, Hyderabad's tech scene is booming. With plans to establish India's first "AI city," the city is on track to become one of the world's most anticipated tech hubs, with companies like TransUnion, Schrödinger and Freshworks, among others, already calling the city home.



