Site Reliability Engineer
EPAM Systems
We are looking for a Site Reliability Engineer to keep cloud services reliable, observable, and automated across multi-tenant Kubernetes environments on Azure. You will own production health, reduce toil through automation, and partner with developers to improve resilience. Operate Kubernetes clusters and containerized workloads running on Azure Troubleshoot production incidents end-to-end across network, OS, platform, and application layers Automate repetitive operational tasks with Python, Bash, or PowerShell to eliminate toil Define and track SLIs/SLOs and drive improvements to meet reliability targets Build and tune monitoring and alerting to detect issues before clients are impacted Improve platform reliability through capacity, performance, and failure-mode analysis Partner with development teams to harden services and improve operability standards Document runbooks and operational procedures to speed up diagnosis and recovery Perform root cause analysis and implement preventive actions after incidents 2+ years of experience in Site Reliability Engineering or DevOps for production systems 2+ years of experience operating Kubernetes and containerized workloads Hands-on experience with Microsoft Azure services for running workloads SLA/SLO adherence experience including defining and tracking SLIs/SLOs Infrastructure fundamentals in networking and operating systems Strong Linux administration skills Strong scripting skills in Python, Bash, or PowerShell Incident response skills with calm, structured troubleshooting under pressure Clear communication skills for cross-team collaboration during incidents and reviews Collaborative mindset to work effectively with development teams English proficiency level B2 (Upper-Intermediate) Argo CD experience for GitOps-based deployments Elastic Stack experience for observability workflows Istio experience for service mesh traffic management Windows Administration experience including Windows Server operations
- ...Capital One Technology Labs Mexico is building a Site Reliability Engineering center in Mexico City and is hiring a Manager-level Backend Engineer to own the reliability of settlement platforms. You will work across on-prem data centers and AWS, alongside UK engineers,...Sugerido
- ...discover how valued you’ll be. We are currently seeking an experienced professional to join our team in the role of Site Reliability Engineer Ensure that HSBC’s technology platforms and services are stable, resilient, secure, scalable and highly available, both...SugeridoEmpleo permanenteTrabajo híbridoHorario flexible
- ...Site Reliability Engineer (SRE) Responsibilities Design, implement, and maintain scalable and highly available infrastructures. Monitor and ensure the performance and reliability of production systems. Implement automation for recurring tasks and operational...SugeridoRemotoHorario flexible
- ...HSBC in Mexico seeks an experienced Site Reliability Engineer to ensure technology platforms are stable, secure and highly available. You will own services end to end, drive automation, and promote DevSecOps practices across teams. The role emphasizes incident prevention...SugeridoTrabajo híbrido
- ...efficiency and maximize self-service availability for financial institutions and retailers across the globe. We are looking for Site Reliability Engineer (SRE) to join our team, with an initial focus on production operations (AppOps). This role is ideal for early-career...Sugerido
- ...SEMICONDUCTORES Y SISTEMAS AVANZADOS DE BAJA CALIFORNIA Job Area Engineering Group, Engineering Group Software Engineering General... ...for: Scalability High availability Performance Reliability Cost efficiency Implement redundancy, failover, and...
- ...Cognizant seeks an experienced Application Reliability Engineer to lead reliability initiatives, analyze production issues, and drive improvements in stability, observability, and operational readiness. You will partner with engineering, operations, platform, and product...
- We are looking for a Junior Site Reliability Engineer to help keep cloud services reliable, observable, and automated across Kubernetes and Microsoft Azure environments. You will support incident response, improve monitoring, and reduce toil through scripting while collaborating...
- ...job involves: We are seeking a skilled and experienced Site Reliability Engineer 2 to join our dynamic team. As a Senior SRE, you will be responsible... ...for leading the implementation and maintenance of reliable, scalable infrastructure and deployment pipelines using...Inicio inmediato
- We are looking for a Site Reliability Engineer to strengthen reliability, observability, and platform operations across cloud and Kubernetes environments. In this role, you will improve service health through automation, infrastructure as code and CI/CD practices. Apply...
- ...Levi Strauss & Co. is seeking a Site Reliability Engineer for their Data & AI Platform Engineering team in Mexico City. You will monitor production systems, respond to incidents, and improve infrastructure automation on platforms including GCP. This role offers an exciting...
- ...WeWork Reforma Latino (97001), Mexico, Ciudad de Mexico, Ciudad de Mexico Principal Associate (Site Reliability Engineering) We’re building a Site Reliability Engineering center in Mexico City and hiring Principal Associate SREs to join one of our founding teams....PrácticaTrabajo híbrido
- ...Mastercard Inc. in Mexico City is seeking a Site Reliability Engineer II to join the Business Operations team. You will help ensure reliability, scalability, and performance of critical applications while mentoring teammates and championing automation. You will work...
- We are looking for a hands-on Senior Site Reliability Engineer to help maintain, enhance, and support a Java services ecosystem in close collaboration with an SRE peer and a backend engineering team. You will strengthen reliability, observability, and operational readiness...
- ...NCR Atleos in Mexico City is seeking a Site Reliability Engineer (SRE) to join our production operations team. You will support cloud platforms, incident management, automation, and CI/CD pipelines to improve reliability and performance. The role emphasizes observability...
- ...Jones Lang LaSalle (JLL) seeks a Senior Site Reliability Engineer to design and maintain scalable infrastructure across cloud accounts, using Terraform and IaC practices. You will lead incident response, define SLOs/SLIs, and drive AI-powered reliability improvements....
- Lead Site Reliability Engineer page is loaded## Lead Site Reliability Engineerlocations: Mexico City, Mexicotime type: Full timeposted on: Posted Yesterdayjob requisition id: R1000681WeWork Reforma Latino (97001), Mexico, Ciudad de Mexico, Ciudad de MexicoLead Site Reliability...PrácticaTrabajo híbrido
- We are looking for a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java backend services ecosystem alongside another SRE and a backend engineering team. You will strengthen reliability, observability, and incident response practices.ResponsibilitiesProvide...
- ...We are looking for a Junior Site Reliability Engineer to help keep cloud services reliable, observable, and automated across Kubernetes and Microsoft Azure environments. You will support incident response, improve monitoring, and reduce toil through scripting while collaborating...
$149,800 - $241,500 por año
...something genuinely rare on your CV, keep reading. About the role: We’re looking for a Senior Site Reliability Engineer who’s passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You’ll design and maintain...Tiempo completoDesde casaRemotoHorario flexible- ...practices and network services used by engineering teams. Implement and improve monitoring... ...repetitive work and improve system reliability and maintainability. Collaborate with... ...Experience automating manual work and building reliable, maintainable systems. Ability to...Tiempo completoDesde casaRemotoHorario flexible
- ...projects within an international network of expertise, this is where you belong. We are currently searching for a Senior Site Reliability Engineer: The Challenge (Responsibilities) Engage in and improve the whole lifecycle of services—from inception and design,...Trabajo híbrido
$7,000 - $12,000
...the value of advanced skills and experience. About the Role: We are seeking a talented and motivated best-in-class Senior Site Reliability Engineer. This role presents an exciting opportunity to thrive in a dynamic, fast-paced environment within a rapidly growing team,...Tiempo completo- .../ PostgreSQL / GCP Para laborar en Guadalajara Esquema Hibrido Inglés conversacional requerido Buscamos un Ingeniero SRE (Site Reliability Engineer) con experiencia en .NET, PostgreSQL y Google Cloud Platform (GCP) para integrarse a un equipo especializado en garantizar...PrácticaTrabajo híbrido
- ...IO Connect Services, empresa especializada en soluciones técnicas en la nube, busca un Site Reliability Engineer Senior para trabajar en un entorno remoto en México. Diseñar, construir y escalar servicios de producción y clústeres en múltiples data centers. Se requiere...Remoto
- # Site Reliability Engineer SeniorNuevoPublicado el 5 oct 2026Ubicación: MéxicoModalidad: RemotoContrato: Tiempo completoNivel: SeniorIdioma: Inglés## Acerca del rolOferta para un Site Reliability Engineer senior en IO Connect Services, empresa especializada en soluciones...Tiempo completoEmpleo permanenteContratoRemoto
- ...and do not endorse products or services of GitLab. An overview of this role We’re looking for a Principal Engineer with deep expertise in Site Reliability, Backend, or Platform Engineering to help shape the next phase of GitLab Dedicated , our fully managed...Tiempo completoRemoto
- ...Responsibilities The Opportunity We are looking for a self-driven, software engineering mindset SRE engineer to: • Drive new shift left activities critical to apply Site Reliability Engineering (SRE) and quality assurance principles within the application design...PrácticaTrabajo por turnos
- ...getting started. How You'll Make an Impact: As a Sr. Site Reliability Engineer, you'll be the guardian of our platform's reliability and... ...What You Bring to the Team: Design and implement reliable, scalable, and efficient cloud infrastructure to support our...Trabajo remotoTrabajar en la oficinaDesde casaFin de semana
- ¿A quien buscamos? Buscamos un Site Reliability Engineer que asegure la confiabilidad, estabilidad, eficiencia y seguridad de las plataformas productivas de Redarbor mediante la implementación de prácticas SRE, automatización, observabilidad avanzada y mejora continua...PrácticaTemporal
¿Desea recibir más vacantes?
Suscríbase y reciba vacantes similares a Site Reliability Engineer. ¡Sea el primero en aplicar!


