Amtech Software Logo

Amtech Software

Site Reliability Engineer II

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in India
Mid level
Remote
Hiring Remotely in India
Mid level
Build and operate reliable AWS-based SaaS platforms through observability, infrastructure automation, CI/CD, container and serverless operations, incident response, security controls, and self-healing systems. The role owns services end to end, improves SLOs and incident metrics, maintains Terraform and GitHub Actions automation, participates in 24/7 on-call, and mentors engineers while applying governed AI-assisted engineering practices.
The summary above was generated by AI

About Vista Equity Partners
Vista Equity Partners is a leading global investment firm focused exclusively on enterprise software, data, and technology-enabled businesses. With over $100B in assets under management and a portfolio of 90+ software product companies worldwide, Vista accelerates growth through operational excellence, shared expertise, and long-term partnership. In India, Vista’s presence continues to expand with 45+ portfolio companies employing more than 17,000 professionals across technology, product, customer success, and operations — reinforcing India’s strategic role as a hub of innovation and talent within the Vista ecosystem.
Through its Agentic AI Factory, Vista is embedding Generative AI across its global portfolio — enabling companies to integrate intelligent, responsible AI into products, operations, and decision-making. This initiative is strengthened through portfolio-wide learning programs, leadership workshops, and AI hackathons that foster innovation, build fluency, and accelerate practical AI adoption across teams.

About Amtech
Amtech is scaling its Platform Engineering organization in India as our products move to a fully AWS-hosted, multi-tenant SaaS model. Our SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate.
As a senior individual contributor on the SRE track, you will own the reliability of significant production domains, design the observability and automation frameworks the team standardizes on, and act as incident commander for complex, cross-service incidents.
You will raise the bar for the India SRE pod: setting patterns, reviewing designs, and mentoring earlier-career engineers while remaining deeply hands-on. 
Our Employee Value Proposition
At Amtech, our people are our greatest differentiator. We create an environment where you can:
Purpose
Shape the future of manufacturing and supply chain operations by delivering mission-critical enterprise software used by industry-leading organizations.
Growth
Access continuous learning, leadership development, and cross-portfolio opportunities through Vista’s global network — accelerating both technical and managerial career paths.
Culture
Work in a collaborative, transparent, and people-first environment where values, accountability, and integrity guide every decision.
Innovation
Engage with cutting-edge technologies, including AI-driven automation, and contribute to modernizing financial systems and operational processes across the business.
Role Description
Amtech is scaling its Platform Engineering organization in India as our products move to a fully AWS-hosted, multi-tenant SaaS model. Our SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate.
As a Platform Engineer on the SRE track, you will build and operate the monitoring, automation, and incident response systems that keep Amtech's services running. You will own defined services end to end, automate away toil, and contribute to reliability standards used by the global team.
This is a hands-on technical role that combines operational discipline with software engineering to drive continuous reliability improvement.
Key Responsibilities
Reliability & Performance

  • Build and maintain monitoring, alerting, and reliability tooling on our OpenTelemetry-based stack (OpenObserve, CloudWatch) with PagerDuty for alert routing and escalation.
  • Analyze production performance, capacity, and error budgets to maintain agreed SLIs and SLOs for your services.
  • Implement automated health checks, scaling rules, and self-healing mechanisms to reduce manual intervention.
  • Contribute to root cause analysis and post-incident reviews, driving permanent fixes.

Automation & Operations

  • Build and maintain infrastructure automation in Terraform across Amtech's multi-account AWS organizations.
  • Develop and maintain CI/CD pipelines in GitHub Actions.
  • Operate containerized and serverless workloads on ECS Fargate, EKS, and Lambda, including RDS PostgreSQL-backed services.
  • Implement safe deployment patterns: automated rollbacks, blue/green, and canary releases.

Incident Response & On-Call

  • Participate in the 24/7 PagerDuty on-call rotation and lead response for incidents in your service area.
  • Reduce MTTD and MTTR through proactive automation and observability improvements.
  • Write and maintain runbooks used across the global SRE team.

Security & Compliance

  • Embed security into automation and deployments: IAM design, secrets management, least privilege.
  • Maintain systems in line with SOC 2 and ISO 27001 controls, producing audit evidence as part of normal operations.

AI Competency

  • Apply AI tools across IaC, pipeline, and automation work with plan review and non-production testing before promotion; understand the blast radius of AI-generated changes.
  • Validate AI output against specifications and standards, including AI-generated tests: review coverage and assertions, not just green results.
  • Use approved tools only and apply Amtech's data classification policy to every AI interaction.

Collaboration & Continuous Improvement

  • Partner with developers to design services for operability, scalability, and resilience.
  • Coordinate with U.S. and India peers to keep reliability practices consistent globally.

Skills & Qualifications

  • 2-4 years of hands-on experience in SRE, DevOps, or cloud engineering roles.
  • Demonstrated ability to operate production workloads on AWS (EC2, ECS/EKS, RDS, S3, IAM, VPC).
  • Working proficiency with Terraform and Git-based CI/CD (GitHub Actions or similar).
  • Solid scripting ability in Python or Bash applied to real automation problems.
  • Experience with modern observability practices (metrics, logs, traces) and tools such as OpenTelemetry, CloudWatch, Prometheus, or Grafana.
  • Understanding of SLO/SLI-driven operations and structured incident management.
  • Sound grasp of networking, DNS, and cloud security fundamentals.
  • Disciplined AI-assisted engineering practice: structured prompting, output validation, and awareness of where AI-generated code fails.
  • Bachelor's degree in Computer Science, Engineering, or a related discipline, or equivalent demonstrated skills.

Preferred Qualifications

  • AWS Certified SysOps Administrator, DevOps Engineer, CKA/CKAD, or Terraform Associate.
  • Experience supporting multi-tenant SaaS or account-per-customer AWS architectures.
  • Exposure to PagerDuty or equivalent incident management platforms.
  • Experience operating AI/LLM-backed services or building agentic automation under governance controls.

Why Join Amtech

At Amtech, you will drive meaningful financial impact in a growing enterprise software organization while benefiting from Vista’s world-class ecosystem. You’ll collaborate with talented peers, leverage cross-portfolio learning programs, and help shape the future of Amtech’s financial operations and systems. Build your career with Amtech — backed by the strength, scale, and innovation culture of Vista.

Similar Jobs

2 Days Ago
Remote
India
Senior level
Senior level
Software
The Site Reliability Engineer will design scalable infrastructure, automated deployment pipelines, monitoring, and alerting systems. Responsibilities include troubleshooting production incidents, participating in on-call rotations, maintaining systems, improving security and compliance, optimizing reliability and efficiency, and mentoring junior engineers. The role requires collaboration with development and cross-functional teams, technical communication, infrastructure automation, and expertise in distributed systems, cloud platforms, and networking.
Top Skills: AnsibleAWSAzureChefDockerGCPGoGrafanaJavaKubernetesNagiosPrometheusPuppetPythonRubyTerraform
2 Days Ago
In-Office or Remote
India
Junior
Junior
Cloud • Security • Software • Cybersecurity
Build and improve reliable, scalable distributed content delivery systems. Define SLIs and SLOs, enhance monitoring and alerting, analyze performance data, resolve complex incidents, automate operational tasks, and participate in architecture reviews. The role requires scripting, Oracle SQL analysis, Unix/Linux expertise, and experience with observability tools including Prometheus, Grafana, ADBMS, and Datadog. Collaboration with product and cross-functional teams is central to ensuring high availability, performance, and resilience.
Top Skills: AdbmsBashCloud ComputingDatadogDevOpsGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux
4 Days Ago
Remote
India
Junior
Junior
Edtech • Software
Build and maintain scalable AWS cloud infrastructure, observability and monitoring systems, deployment pipelines, and automation. Participate in on-call rotations, incident response, post-mortems, and root-cause analysis. Support product engineering teams with infrastructure troubleshooting and operational best practices while applying security and compliance controls. The role requires experience with Terraform, CI/CD, Linux, shell scripting, cloud services, managed Kubernetes, databases, and debugging application code.
Top Skills: Amazon RedshiftAWSAws CodebuildAws CodepipelineEc2EksGithub ActionsGoJavaScriptJenkinsLinuxMongoDBOpensearchPythonS3Serverless FrameworksTerraformTypescriptUnix ShellVpc

What you need to know about the Hyderabad Tech Scene

Because of its proximity to leading research institutions and a government committed to the city's growth, Hyderabad's tech scene is booming. With plans to establish India's first "AI city," the city is on track to become one of the world's most anticipated tech hubs, with companies like TransUnion, Schrödinger and Freshworks, among others, already calling the city home.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account