Amtech Software Logo

Amtech Software

Site Reliability Engineer II

Posted 21 Days Ago
Remote
Hiring Remotely in India
Mid level
Remote
Hiring Remotely in India
Mid level
Build and operate monitoring, automation, and incident response systems for a multi-account AWS SaaS platform. Own services end-to-end, automate infrastructure with Terraform, maintain CI/CD in GitHub Actions, run containerized and serverless workloads (ECS Fargate, EKS, Lambda), improve SLO/SLI-driven reliability, participate in PagerDuty on-call, embed security and compliance (SOC 2, ISO 27001), and apply governed AI-assisted engineering workflows.
The summary above was generated by AI

About Vista Equity Partners
Vista Equity Partners is a leading global investment firm focused exclusively on enterprise software, data, and technology-enabled businesses. With over $100B in assets under management and a portfolio of 90+ software product companies worldwide, Vista accelerates growth through operational excellence, shared expertise, and long-term partnership. In India, Vista’s presence continues to expand with 45+ portfolio companies employing more than 17,000 professionals across technology, product, customer success, and operations — reinforcing India’s strategic role as a hub of innovation and talent within the Vista ecosystem.
Through its Agentic AI Factory, Vista is embedding Generative AI across its global portfolio — enabling companies to integrate intelligent, responsible AI into products, operations, and decision-making. This initiative is strengthened through portfolio-wide learning programs, leadership workshops, and AI hackathons that foster innovation, build fluency, and accelerate practical AI adoption across teams.

About Amtech
Amtech is scaling its Platform Engineering organization in India as our products move to a fully AWS-hosted, multi-tenant SaaS model. Our SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate.
As a senior individual contributor on the SRE track, you will own the reliability of significant production domains, design the observability and automation frameworks the team standardizes on, and act as incident commander for complex, cross-service incidents.
You will raise the bar for the India SRE pod: setting patterns, reviewing designs, and mentoring earlier-career engineers while remaining deeply hands-on. 
Our Employee Value Proposition
At Amtech, our people are our greatest differentiator. We create an environment where you can:
Purpose
Shape the future of manufacturing and supply chain operations by delivering mission-critical enterprise software used by industry-leading organizations.
Growth
Access continuous learning, leadership development, and cross-portfolio opportunities through Vista’s global network — accelerating both technical and managerial career paths.
Culture
Work in a collaborative, transparent, and people-first environment where values, accountability, and integrity guide every decision.
Innovation
Engage with cutting-edge technologies, including AI-driven automation, and contribute to modernizing financial systems and operational processes across the business.
Role Description
Amtech is scaling its Platform Engineering organization in India as our products move to a fully AWS-hosted, multi-tenant SaaS model. Our SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate.
As a Platform Engineer on the SRE track, you will build and operate the monitoring, automation, and incident response systems that keep Amtech's services running. You will own defined services end to end, automate away toil, and contribute to reliability standards used by the global team.
This is a hands-on technical role that combines operational discipline with software engineering to drive continuous reliability improvement.
Key Responsibilities
Reliability & Performance

  • Build and maintain monitoring, alerting, and reliability tooling on our OpenTelemetry-based stack (OpenObserve, CloudWatch) with PagerDuty for alert routing and escalation.
  • Analyze production performance, capacity, and error budgets to maintain agreed SLIs and SLOs for your services.
  • Implement automated health checks, scaling rules, and self-healing mechanisms to reduce manual intervention.
  • Contribute to root cause analysis and post-incident reviews, driving permanent fixes.

Automation & Operations

  • Build and maintain infrastructure automation in Terraform across Amtech's multi-account AWS organizations.
  • Develop and maintain CI/CD pipelines in GitHub Actions.
  • Operate containerized and serverless workloads on ECS Fargate, EKS, and Lambda, including RDS PostgreSQL-backed services.
  • Implement safe deployment patterns: automated rollbacks, blue/green, and canary releases.

Incident Response & On-Call

  • Participate in the 24/7 PagerDuty on-call rotation and lead response for incidents in your service area.
  • Reduce MTTD and MTTR through proactive automation and observability improvements.
  • Write and maintain runbooks used across the global SRE team.

Security & Compliance

  • Embed security into automation and deployments: IAM design, secrets management, least privilege.
  • Maintain systems in line with SOC 2 and ISO 27001 controls, producing audit evidence as part of normal operations.

AI Competency

  • Apply AI tools across IaC, pipeline, and automation work with plan review and non-production testing before promotion; understand the blast radius of AI-generated changes.
  • Validate AI output against specifications and standards, including AI-generated tests: review coverage and assertions, not just green results.
  • Use approved tools only and apply Amtech's data classification policy to every AI interaction.

Collaboration & Continuous Improvement

  • Partner with developers to design services for operability, scalability, and resilience.
  • Coordinate with U.S. and India peers to keep reliability practices consistent globally.

Skills & Qualifications

  • 2-4 years of hands-on experience in SRE, DevOps, or cloud engineering roles.
  • Demonstrated ability to operate production workloads on AWS (EC2, ECS/EKS, RDS, S3, IAM, VPC).
  • Working proficiency with Terraform and Git-based CI/CD (GitHub Actions or similar).
  • Solid scripting ability in Python or Bash applied to real automation problems.
  • Experience with modern observability practices (metrics, logs, traces) and tools such as OpenTelemetry, CloudWatch, Prometheus, or Grafana.
  • Understanding of SLO/SLI-driven operations and structured incident management.
  • Sound grasp of networking, DNS, and cloud security fundamentals.
  • Disciplined AI-assisted engineering practice: structured prompting, output validation, and awareness of where AI-generated code fails.
  • Bachelor's degree in Computer Science, Engineering, or a related discipline, or equivalent demonstrated skills.

Preferred Qualifications

  • AWS Certified SysOps Administrator, DevOps Engineer, CKA/CKAD, or Terraform Associate.
  • Experience supporting multi-tenant SaaS or account-per-customer AWS architectures.
  • Exposure to PagerDuty or equivalent incident management platforms.
  • Experience operating AI/LLM-backed services or building agentic automation under governance controls.

Why Join Amtech

At Amtech, you will drive meaningful financial impact in a growing enterprise software organization while benefiting from Vista’s world-class ecosystem. You’ll collaborate with talented peers, leverage cross-portfolio learning programs, and help shape the future of Amtech’s financial operations and systems. Build your career with Amtech — backed by the strength, scale, and innovation culture of Vista.

Similar Jobs

22 Days Ago
Remote
India
Mid level
Mid level
Security • Cybersecurity
Build, deploy, scale, and operate highly available distributed systems across regions. Maintain Kubernetes infrastructure, Helm charts, CI/CD pipelines, and automation. Implement observability, alerting, incident response, and runbooks. Troubleshoot production issues, perform root cause analysis, and collaborate with development teams to ensure production readiness and scalability.
Top Skills: Aws EcrAws EksAws IamAws S3Aws VpcBashCi/CdGitHelmKubernetesTerraform
23 Days Ago
Remote
India
Senior level
Senior level
Software
Ensure platform reliability and scalability by collaborating with development teams to design infrastructure and CI/CD pipelines, build monitoring/alerting, troubleshoot production incidents, participate in on-call rotations, automate configuration and deployments, maintain upgrades, improve security/compliance, mentor junior engineers, and drive process and organizational improvements.
Top Skills: AnsibleAWSAzureChefDockerGCPGoGrafanaJavaKubernetesNagiosPrometheusPuppetPythonRubyTerraform
15 Days Ago
Remote
India
Senior level
Senior level
Software
The Senior Site Reliability Engineer ensures the availability, performance, and security of Drivetrain's SaaS platform, managing multi-cloud infrastructure, optimizing CI/CD pipelines, and driving automation.
Top Skills: AWSAws CloudwatchEfkElkGCPGcp Operations SuiteGitGrafanaIstioJenkinsKubernetesKustomizeLinkerdPrometheusPythonTerraform

What you need to know about the Hyderabad Tech Scene

Because of its proximity to leading research institutions and a government committed to the city's growth, Hyderabad's tech scene is booming. With plans to establish India's first "AI city," the city is on track to become one of the world's most anticipated tech hubs, with companies like TransUnion, Schrödinger and Freshworks, among others, already calling the city home.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account