About the Company
We are a global cloud technology and digital transformation company delivering enterprise-scale cloud platforms, artificial intelligence, cybersecurity, software engineering, and managed infrastructure solutions to Fortune 500 organizations worldwide. Our cloud environments support mission-critical applications serving millions of users across healthcare, financial services, retail, manufacturing, telecommunications, government, and technology industries.
As organizations accelerate cloud adoption and distributed application development, observability has become essential to delivering reliable, secure, and high-performing digital experiences. Our Cloud Platform Engineering organization builds next-generation observability platforms that provide real-time visibility into infrastructure, applications, networks, and customer experiences. Leveraging cloud-native technologies, automation, AI-driven monitoring, and Site Reliability Engineering (SRE) principles, we empower engineering teams to proactively identify issues, optimize performance, and maintain exceptional service reliability.
We are seeking an accomplished Principal Cloud Engineer – Observability to lead the architecture, strategy, and implementation of enterprise observability platforms across multi-cloud environments. This highly influential technical leadership role is responsible for defining observability standards, driving platform modernization, establishing engineering best practices, and mentoring engineering teams responsible for monitoring, telemetry, distributed tracing, logging, and operational intelligence.
The Principal Cloud Engineer – Observability partners closely with Software Engineering, DevOps, Site Reliability Engineering, Cloud Infrastructure, Cybersecurity, Enterprise Architecture, Product Engineering, and Executive Leadership to build scalable observability solutions that improve platform resilience, operational efficiency, and customer experience.
This opportunity is ideal for a visionary cloud engineering leader with deep expertise in observability, cloud-native infrastructure, automation, and enterprise platform engineering.
Essential Duties and Responsibilities
- Define and execute the enterprise observability strategy across cloud infrastructure, applications, and distributed systems.
- Architect scalable monitoring, logging, distributed tracing, and telemetry platforms supporting mission-critical workloads.
- Lead implementation of enterprise observability solutions using OpenTelemetry, Prometheus, Grafana, Datadog, Splunk, New Relic, Dynatrace, Elastic, or equivalent platforms.
- Establish observability standards, engineering frameworks, governance policies, and operational best practices.
- Design proactive monitoring solutions supporting high availability, resiliency, and incident prevention.
- Partner with Site Reliability Engineering teams to improve service reliability, incident response, and operational excellence.
- Build Infrastructure as Code (IaC) and observability automation using Terraform, Kubernetes, Helm, Ansible, or comparable technologies.
- Optimize cloud platform performance, cost efficiency, scalability, and resource utilization through advanced telemetry and analytics.
- Develop executive dashboards, SLA/SLO/SLI reporting, and operational health metrics.
- Collaborate with Security Engineering to strengthen monitoring, threat detection, compliance, and audit capabilities.
- Mentor senior engineers while promoting cloud engineering, observability, and DevSecOps best practices.
- Evaluate emerging observability technologies and provide strategic technical recommendations.
- Lead architecture reviews, platform modernization initiatives, and enterprise cloud transformation programs.
- Support disaster recovery planning, capacity management, and business continuity initiatives.
Job Qualifications and Requirements
- Bachelor’s degree in Computer Science, Software Engineering, Information Technology, Computer Engineering, or a related technical discipline required.
- Master’s degree in Computer Science, Cloud Computing, or Engineering preferred.
- Minimum of 10 years of professional experience in cloud engineering, platform engineering, infrastructure engineering, or Site Reliability Engineering.
- Minimum of 5 years leading enterprise observability or cloud platform initiatives.
- Advanced expertise with AWS, Microsoft Azure, or Google Cloud Platform.
- Extensive experience with OpenTelemetry, Prometheus, Grafana, Splunk, Datadog, Dynatrace, Elastic Stack, New Relic, or comparable observability platforms.
- Strong knowledge of Kubernetes, Docker, microservices, service mesh technologies, and cloud-native architectures.
- Expertise in Infrastructure as Code using Terraform, CloudFormation, Pulumi, or similar automation frameworks.
- Experience with CI/CD, DevOps, GitOps, monitoring automation, and incident management processes.
- Professional certifications such as AWS Certified Solutions Architect Professional, Google Professional Cloud Architect, Microsoft Azure Solutions Architect Expert, CKA, or equivalent are highly preferred.
- Outstanding technical leadership, architecture, communication, and stakeholder management skills.
Personal Capabilities and Qualifications
The successful Principal Cloud Engineer – Observability is a highly respected technical leader who combines deep engineering expertise with strategic vision and operational excellence. They thrive in complex enterprise environments while mentoring teams and delivering innovative cloud platform solutions.
Ideal candidates possess:
- Enterprise cloud architecture and observability leadership expertise.
- Strong Site Reliability Engineering and platform engineering knowledge.
- Exceptional analytical and systems-thinking capabilities.
- Outstanding communication and executive presentation skills.
- Strong mentoring and technical leadership abilities.
- Advanced troubleshooting and incident management expertise.
- High attention to detail and engineering discipline.
- Adaptability within rapidly evolving cloud technologies.
- Professional integrity and ownership.
- Passion for innovation, automation, and continuous improvement.
Strategic Support
As a senior technical leader within the Cloud Platform Engineering organization, the Principal Cloud Engineer – Observability provides strategic support by:
- Defining enterprise observability architecture and engineering standards.
- Improving application reliability, availability, and customer experience.
- Accelerating cloud modernization and digital transformation initiatives.
- Strengthening operational intelligence and proactive monitoring capabilities.
- Supporting DevSecOps, automation, and platform engineering initiatives.
- Optimizing cloud infrastructure performance and operational costs.
- Developing engineering talent through mentorship and technical leadership.
- Contributing to long-term enterprise cloud strategy and technology innovation.
Working Conditions
- Fully Remote position within the United States.
- Flexible executive-level schedule supporting global engineering organizations across multiple time zones.
- Approximately 10–20% domestic and occasional international travel for engineering summits, customer engagements, executive planning sessions, and technology conferences.
- Company-provided high-performance workstation, cloud development resources, enterprise software, and home office technology support.
- Collaborative engineering environment focused on innovation, technical excellence, and continuous learning.
Job Function
The Principal Cloud Engineer – Observability is responsible for leading enterprise observability architecture, cloud platform engineering, monitoring strategy, and operational intelligence initiatives. This role partners with executive leadership and engineering organizations to build scalable, secure, and highly available cloud platforms that improve service reliability, operational visibility, and customer experience. Success is measured through platform availability, system performance, observability maturity, engineering productivity, operational efficiency, and strategic business impact.
Compensation & Benefits
Compensation Package
Base Salary: $372,000 – $424,000 USD annually
Compensation is determined based on cloud architecture expertise, observability leadership experience, enterprise platform engineering knowledge, technical certifications, geographic location, and overall qualifications.
Eligible employees may also receive:
- Annual Executive Performance Bonus
- Long-Term Equity Incentive Program (Performance Stock Awards & RSUs)
- Executive 401(k) with Company Match
- Comprehensive Medical, Dental, and Vision Insurance
- Flexible Paid Time Off
- Paid Company Holidays
- Paid Parental and Family Leave
- Professional Certifications and Continuing Education Reimbursement
- Executive Leadership Development Programs
- Annual Technology Conference and Innovation Budget
- Home Office and Executive Technology Stipend
- Employee Stock Purchase Program (where applicable)
- Wellness and Mental Health Benefits
- Life and Disability Insurance
- Career Advancement into Distinguished Engineer, Director of Cloud Engineering, Vice President of Platform Engineering, or Chief Technology Officer pathways
Why Join Us
Modern cloud platforms depend on world-class observability to deliver exceptional customer experiences. As our Principal Cloud Engineer – Observability, you’ll define the technology strategy behind enterprise monitoring, telemetry, and operational intelligence while shaping the future of cloud engineering across a global organization.
You’ll collaborate with leading cloud architects, Site Reliability Engineers, software developers, and executive technology leaders while working with cutting-edge observability platforms and cloud-native technologies. We foster a culture of innovation, technical excellence, and continuous learning where your expertise will directly influence enterprise-scale engineering and digital transformation.
If you’re passionate about cloud architecture, observability, platform engineering, and building resilient technology ecosystems, we encourage you to apply and become part of our Cloud Platform Engineering & Site Reliability Engineering organization.