About the Company
We are a global cloud technology and enterprise software company delivering mission-critical platforms that power digital transformation for Fortune 500 organizations across financial services, healthcare, cybersecurity, telecommunications, retail, and advanced manufacturing. Our cloud-native solutions support millions of users worldwide, providing secure, highly available, and scalable services that businesses depend on every day.
Reliability is the foundation of every product we build. As our platform ecosystem continues to expand, we are investing heavily in modern infrastructure, intelligent automation, and resilient distributed systems to ensure exceptional service availability and operational excellence. Our Site Reliability Engineering (SRE) organization partners closely with Software Engineering, Cloud Infrastructure, Security, and DevOps teams to create scalable platforms capable of supporting rapid innovation without compromising reliability or performance.
We are seeking an experienced Site Reliability Engineer to help architect, automate, and optimize enterprise infrastructure supporting large-scale cloud applications. This role blends software engineering with systems administration, focusing on improving reliability, scalability, observability, and operational efficiency through engineering excellence and automation.
The successful candidate will work with modern cloud platforms, Kubernetes, Infrastructure as Code (IaC), CI/CD pipelines, distributed systems, and observability technologies while driving initiatives that enhance platform resilience, reduce operational complexity, and improve customer experience.
If you’re passionate about building highly available cloud platforms, eliminating operational toil through automation, and solving complex infrastructure challenges at enterprise scale, we invite you to become part of our growing Site Reliability Engineering team.
Essential Duties and Responsibilities
- Design, implement, and maintain highly available, scalable, and secure cloud infrastructure supporting enterprise applications.
- Develop automation solutions that reduce manual operations and improve platform reliability.
- Build and maintain Infrastructure as Code using Terraform, CloudFormation, or equivalent technologies.
- Deploy, manage, and optimize Kubernetes clusters and containerized workloads across cloud environments.
- Design and improve CI/CD pipelines that support reliable software delivery and continuous deployment.
- Establish monitoring, logging, alerting, and observability standards using enterprise monitoring platforms.
- Improve system reliability through proactive capacity planning, performance tuning, and infrastructure optimization.
- Lead incident response activities, perform root cause analysis, and implement long-term corrective actions.
- Collaborate with Software Engineering teams to improve application resiliency and operational readiness.
- Implement disaster recovery, backup, and business continuity strategies for critical production systems.
- Strengthen infrastructure security by applying industry best practices and cloud security controls.
- Create technical documentation, operational runbooks, and engineering standards.
- Mentor engineers on reliability engineering principles, automation, and operational excellence.
- Evaluate emerging cloud and infrastructure technologies that improve platform stability and efficiency.
Job Qualifications and Requirements
- Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related discipline.
- Minimum of 6 years of professional experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Infrastructure Engineering.
- Strong expertise with AWS, Microsoft Azure, or Google Cloud Platform.
- Extensive experience with Kubernetes, Docker, and container orchestration.
- Strong knowledge of Infrastructure as Code using Terraform, CloudFormation, or Pulumi.
- Experience building CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI, or Azure DevOps.
- Proficiency with Linux administration, Bash, Python, or Go for infrastructure automation.
- Experience with observability platforms including Prometheus, Grafana, Datadog, Splunk, New Relic, or OpenTelemetry.
- Strong understanding of networking, load balancing, DNS, VPNs, and distributed systems architecture.
- AWS Certified DevOps Engineer, Kubernetes Administrator (CKA), Google Professional Cloud DevOps Engineer, or equivalent certification is highly preferred.
Personal Capabilities and Qualifications
The successful Site Reliability Engineer combines deep technical expertise with a systems-thinking mindset and a passion for automation. They proactively identify opportunities to improve platform reliability while collaborating effectively across engineering teams.
Ideal candidates possess:
- Strong systems architecture and infrastructure design capabilities.
- Excellent analytical and troubleshooting skills.
- Outstanding collaboration and communication abilities.
- Strong automation and scripting expertise.
- Close attention to detail and operational discipline.
- Ability to perform effectively during critical production incidents.
- Continuous learning mindset focused on emerging cloud technologies.
- Strong project management and organizational skills.
- Professional integrity and accountability.
- Customer-first approach to platform reliability and service excellence.
Strategic Support
As a member of the Site Reliability Engineering organization, the Site Reliability Engineer provides strategic support by:
- Improving platform availability and operational resilience.
- Driving automation that enhances engineering productivity.
- Supporting enterprise cloud transformation initiatives.
- Strengthening observability, monitoring, and incident response capabilities.
- Improving deployment reliability through modern DevOps practices.
- Optimizing cloud infrastructure performance and operational costs.
- Supporting business continuity and disaster recovery planning.
- Contributing to the long-term infrastructure and reliability engineering strategy.
Working Conditions
- Fully Remote position within the United States.
- Flexible work schedule supporting collaboration across multiple U.S. time zones.
- Participation in an on-call rotation for critical production systems.
- Company-provided development hardware, collaboration software, and home office technology support.
- Agile engineering environment focused on innovation, automation, and continuous improvement.
- Approximately 5–10% domestic travel for engineering summits, planning sessions, and technical workshops.
Job Function
The Site Reliability Engineer is responsible for designing, automating, and maintaining highly available cloud infrastructure that supports enterprise-scale applications and digital services. This role partners with software engineering, platform engineering, cybersecurity, and cloud operations teams to improve system reliability, scalability, observability, and operational efficiency. Success is measured through platform uptime, infrastructure performance, automation maturity, incident reduction, service reliability, and continuous engineering improvement.
Compensation & Benefits
Compensation Package
Base Salary: $285,000 – $365,000 USD annually
Compensation is determined based on site reliability engineering expertise, cloud platform experience, technical certifications, geographic location, and overall qualifications.
Eligible employees may also receive:
- Annual Performance Bonus
- Long-Term Equity Incentive Program (RSUs or Stock Options)
- 401(k) with Company Match
- Comprehensive Medical, Dental, and Vision Insurance
- Flexible Paid Time Off
- Paid Company Holidays
- Paid Parental and Family Leave
- Professional Cloud and Kubernetes Certification Reimbursement
- Annual Learning and Technology Conference Budget
- Home Office and Technology Stipend
- Employee Stock Purchase Program (where applicable)
- Wellness and Mental Health Benefits
- Life and Disability Insurance
- Tuition Assistance
- Leadership Development and Career Advancement Opportunities
Why Join Us
Our platform powers business-critical applications used by organizations around the world, and reliability is essential to everything we build. As a Site Reliability Engineer, you’ll help create resilient, scalable cloud infrastructure that enables continuous innovation while maintaining exceptional performance and availability.
You’ll work alongside world-class engineers using cutting-edge cloud technologies, Kubernetes, automation frameworks, and observability platforms to solve complex infrastructure challenges at scale. We foster a culture of ownership, collaboration, and continuous improvement where your expertise will directly influence the reliability and future evolution of our technology platforms.
If you’re passionate about automation, cloud infrastructure, distributed systems, and engineering reliable platforms that support millions of users, we encourage you to join our Site Reliability Engineering team and help shape the future of enterprise cloud services.