Senior Platform Engineer
Salve esta vaga e mantenha sua pesquisa organizada
Crie uma conta gratuita para salvar vagas, criar alertas e retornar a esta listagem a partir do seu painel.
It's more than a job
In this IT role, you will manage, develop, and support the technology infrastructure that powers Kuehne+Nagel’s global systems and operations. This contribution is vital to making both special and ordinary moments in the lives of people a reality. In fact, your work enables our teams to deliver fresh fruit, time-sensitive healthcare products and much more to people all around the world. Working in IT at Kuehne+Nagel contributes to more than we imagine.
We are seeking a hands‑on Senior IT Platform Engineer to support and operate a Kafka-based event streaming platform. This role focuses on day‑to‑day operations, incident resolution, and platform reliability, ensuring seamless service for internal users. You will work closely with engineering teams, platform owners, and internal customers to maintain system stability, improve performance, and enhance self‑service capabilities.
Platform Operations & Reliability
- Operate, monitor, and maintain a Kafka-based messaging platform hosted on AWS (EKS).
- Ensure high availability, performance, and resilience of the platform.
- Monitor system health using logs, metrics, and observability tools (e.g., Grafana, Loki).
- Support scaling of infrastructure to handle increasing workloads and integrations.
Incident Management & Troubleshooting
- Investigate and resolve platform incidents across Kafka components (brokers, producers, consumers).
- Analyze logs and metrics to identify root causes and implement fixes.
- Escalate complex issues to engineering teams when needed.
- Participate in on‑call (L3 support) rotations.
Runbook Execution & Operational Excellence
- Execute operational tasks following runbooks and SOPs.
- Perform Kafka configuration tasks (topics, ACLs, schemas) using defined processes.
- Continuously improve operational procedures and documentation.
- Contribute to post‑incident reviews and preventive improvements.
Customer Support & Collaboration
- Act as a primary contact for internal platform users.
- Support users via collaboration tools (Slack, Teams).
- Provide technical guidance on platform usage and best practices.
- Translate user issues into actionable technical requirements.
Automation, CI/CD & Platform Improvements
- Support CI/CD pipelines (GitLab) for infrastructure and configuration changes.
- Assist in automating infrastructure using GitOps (ArgoCD) and Helm.
- Improve platform observability, performance, and operational efficiency.
- Enable self‑service capabilities for teams (authentication, access, topic management).
What we would like you to bring
Core Technical Skills
- Experience with Apache Kafka and its ecosystem (e.g., Kafka Connect, Schema Registry)
- Strong understanding of distributed systems (partitioning, replication, scaling, fault tolerance)
- Hands‑on experience with Kubernetes and container orchestration (EKS preferred)
- Experience with cloud platforms, preferably AWS (VPC, EKS, S3, IAM, networking, private links)
- Familiarity with CI/CD pipelines (GitLab) and version control using Git
- Understanding of Infrastructure as Code (IaC) and GitOps practices (Terraform, Helm, ArgoCD)
Operations & Support Skills
- Proven experience in IT operations, platform support, or production support environments
- Strong troubleshooting skills using logs, metrics, and monitoring tools
- Experience with incident management processes, escalation, and root cause analysis
- Familiarity with runbooks, SOPs, and structured operational workflows
- Experience supporting high‑availability, distributed production systems
Observability & Monitoring
- Experience with observability and monitoring tools such as: Grafana, Prometheus, Loki, Mimir, Alloy (LGTM stack)
- Ability to analyze system behaviour, performance, latency, and throughput
Automation & Scripting
- Scripting skills in Bash, Python, or Go for automating operational tasks
- Experience automating infrastructure provisioning and configuration
Messaging & Data Platform Knowledge
- Experience managing Kafka clusters, topics, ACLs, Schema registries, and data contracts
- Understanding of asynchronous communication patterns and event‑driven architecture
Additional Technical Exposure
- Basic understanding of the Java ecosystem
- Experience with container tooling and cloud‑native platforms
- Exposure to data integration platforms and large‑scale messaging systems
Soft Skills & Mindset
- Strong problem‑solving a