Senior Site Reliability Engineer, Production Engineering New Lisbon, Portugal
2 semanas atrás
.Senior Site Reliability Engineer, Production Engineering Lisbon, Portugal Please note that we have a hybrid approach to work and would like to find someone who can come into our offices in Lagoas Park once a week. Who We Are Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences across every network—even those beyond their ownership. Leveraging AI and an unparalleled set of cloud, internet, and enterprise network telemetry data, ThousandEyes enables IT teams to proactively detect, diagnose, and resolve issues before they impact end-user experiences. About The Role We are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform. Key Responsibilities Identify and provide solutions to common obstacles hindering operational excellence across engineering teams. Partner with application developers using cloud-native tools to address novel challenges around scale, performance, and reliability. Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform. Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability. Use cloud-native observability and reliability tools such as Prometheus, Istio, and ArgoCD. Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code. What You'll Do Collaborate with software engineers to ensure architecture and services are optimized for availability, latency, and performance. Design and implement scalable operations tooling to support platform growth and scaling across multiple regions. Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure. Participate in and improve our 24x7 incident response and on-call rotation. Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability. Automate production operations to provide guardrails and continuous platform operation. Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing. Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform. Required Qualifications Expert-level knowledge of Kubernetes and its ecosystem. Proficiency in software development with languages such as Python or Go. In-depth knowledge of cloud providers, preferably AWS
-
Lisboa, Portugal Tbwa ChiatDay Inc Tempo inteiroSenior Site Reliability Engineer, Production Engineering Lisbon, PortugalPlease note that we have a hybrid approach to work and would like to find someone who can come into our offices in Lagoas Park once a week. Who We Are Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital...
-
Lisboa, Portugal Tbwa ChiatDay Inc Tempo inteiro.Senior Site Reliability Engineer, Production EngineeringLisbon, PortugalPlease note that we have a hybrid approach to work and would like to find someone who can come into our offices in Lagoas Park once a week.Who We AreCisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences...
-
Lisboa, Portugal Tbwa ChiatDay Inc Tempo inteiroSenior Site Reliability Engineer, Production EngineeringLisbon, PortugalPlease note that we have a hybrid approach to work and would like to find someone who can come into our offices in Lagoas Park once a week.Who We AreCisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences...
-
Lisboa, Portugal Tbwa ChiatDay Inc Tempo inteiroSenior Site Reliability Engineer, Production EngineeringLisbon, PortugalPlease note that we have a hybrid approach to work and would like to find someone who can come into our offices in Lagoas Park once a week.Who We AreCisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences...
-
Senior Software Engineer
Há 2 dias
Lisboa, Portugal Datadog Tempo inteiroSenior Software Engineer - Site Reliability (Lisbon) Lisbon, Portugal The Site Reliability teams at Datadog are responsible for ensuring that our high-volume, low-latency environments continue to perform around the clock. These teams collaborate closely with our product engineers to ensure that Datadog can monitor millions of servers and containers, ensuring...
-
Senior Software Engineer
3 semanas atrás
Lisboa, Portugal Datadog Tempo inteiroSenior Software Engineer - Site Reliability (Lisbon) Lisbon, PortugalThe Site Reliability teams at Datadog are responsible for ensuring that our high-volume, low-latency environments continue to perform around the clock. These teams collaborate closely with our product engineers to ensure that Datadog can monitor millions of servers and containers, ensuring...
-
Senior Software Engineer
Há 3 dias
Lisboa, Portugal Datadog Tempo inteiroSenior Software Engineer - Site Reliability (Lisbon)Lisbon, PortugalThe Site Reliability teams at Datadog are responsible for ensuring that our high-volume, low-latency environments continue to perform around the clock. These teams collaborate closely with our product engineers to ensure that Datadog can monitor millions of servers and containers, ensuring...
-
Senior Software Engineer
Há 4 dias
Lisboa, Portugal Datadog Tempo inteiroSenior Software Engineer - Site Reliability (Lisbon)Lisbon, PortugalThe Site Reliability teams at Datadog are responsible for ensuring that our high-volume, low-latency environments continue to perform around the clock. These teams collaborate closely with our product engineers to ensure that Datadog can monitor millions of servers and containers, ensuring...
-
Site Reliability Engineering Infrastructure Specialist
2 semanas atrás
Lisboa, Lisboa, Portugal Buscojobs Portugal Tempo inteiroAbout UsBuscojobs Portugal is a leading company in the field of language operations, providing fast, efficient, and high-quality translations through its advanced AI-powered platform.Job Title: Site Reliability Engineering InternSalary Range: €30,000 - €45,000 per annum (dependent on experience)Location: Lisbon, PortugalJob Description:We are seeking a...
-
Senior Site Reliability Engineer
Há 1 mês
Lisboa, Lisboa, Portugal Iownit Tempo inteiroAbout the Role: We are seeking a talented Senior Site Reliability Engineer to join our Infrastructure team at Iownit. As a key member, you will be responsible for ensuring the scalability, reliability, and performance of our production environment. Your expertise in AWS infrastructure and container orchestration systems will be instrumental in driving...
-
Site Reliability Engineering Intern
3 semanas atrás
Lisboa, Portugal Buscojobs Portugal Tempo inteiroWe are seeking a motivated and talented Site Reliability Engineer (SRE) Intern to join our team. As an SRE Intern, you will have the opportunity to work alongside experienced engineers to ensure the reliability, availability, and performance of our systems and services. You will gain hands-on experience in maintaining and improving the infrastructure that...
-
Site Reliability Engineering Intern
3 semanas atrás
Lisboa, Portugal Buscojobs Portugal Tempo inteiroWe are seeking a motivated and talented Site Reliability Engineer (SRE) Intern to join our team. As an SRE Intern, you will have the opportunity to work alongside experienced engineers to ensure the reliability, availability, and performance of our systems and services. You will gain hands-on experience in maintaining and improving the infrastructure that...
-
Site Reliability Engineering Intern
3 semanas atrás
Lisboa, Portugal Buscojobs Portugal Tempo inteiroWe are seeking a motivated and talented Site Reliability Engineer (SRE) Intern to join our team. As an SRE Intern, you will have the opportunity to work alongside experienced engineers to ensure the reliability, availability, and performance of our systems and services. You will gain hands-on experience in maintaining and improving the infrastructure that...
-
Sr. Site Reliability Engineer
5 meses atrás
Lisboa, Portugal Kudzu Interactive, Inc. Tempo inteiroSr. Site Reliability Engineer (SRE) Remote but must be located in Portugal High-level... The Sr. Site Reliability Engineer (SRE) is responsible for the availability, performance, monitoring, release engineering, and incident response, among other things, of the platforms and services the company runs and owns. SRE ensures that enterprise services have...
-
Site Reliability Engineer
2 semanas atrás
Lisboa, Portugal Equadis Sa Tempo inteiroYour missions as a Site Reliability Engineer! You will focus on understanding requirements, designing and deploying solutions using Google Cloud and, as a key member of our engineering practice solve tough problems and ensure success in designing and building complex world-class applications on public cloud. You will be responsible for ensuring the...
-
Site Reliability Engineer
4 semanas atrás
Lisboa, Portugal Equadis Sa Tempo inteiroYour missions as a Site Reliability Engineer!You will focus on understanding requirements, designing and deploying solutions using Google Cloud and, as a key member of our engineering practice solve tough problems and ensure success in designing and building complex world-class applications on public cloud.You will be responsible for ensuring the...
-
Sr. Site Reliability Engineer
4 meses atrás
Lisboa, Portugal Kudzu Interactive, Inc. Tempo inteiroSr. Site Reliability Engineer (SRE)Remote but must be located in PortugalHigh-level...The Sr. Site Reliability Engineer (SRE) is responsible for the availability, performance, monitoring, release engineering, and incident response, among other things, of the platforms and services the company runs and owns. SRE ensures that enterprise services have...
-
Site Reliability Engineer
4 semanas atrás
Lisboa, Portugal Pertemps Erp Tempo inteiroAre you excited by the idea of working within a global organization, helping to power a suite of innovative, cloud-native applications? Do you thrive in environments centered on cloud technology, microservices, and machine learning? If so, we may have the perfect opportunity for you.About the Role As a Senior Site Reliability Engineer, you'll be a key part...
-
Site Reliability Engineer
4 semanas atrás
Lisboa, Portugal Pertemps Erp Tempo inteiroAre you excited by the idea of working within a global organization, helping to power a suite of innovative, cloud-native applications? Do you thrive in environments centered on cloud technology, microservices, and machine learning? If so, we may have the perfect opportunity for you.About the Role As a Senior Site Reliability Engineer, you'll be a key part...
-
Site Reliability Engineer
1 semana atrás
Lisboa, Portugal Pertemps ERP Tempo inteiroAre you excited by the idea of working within a global organization, helping to power a suite of innovative, cloud-native applications? Do you thrive in environments centered on cloud technology, microservices, and machine learning? If so, we may have the perfect opportunity for you.About the RoleAs a Senior Site Reliability Engineer, you’ll be a key part...