Senior Infrastructure Engineer (Network, Datacenter)
<div class="show-more-less-html__markup show-more-less-html__markup--clamp-after-5 relative overflow-hidden"> <p><strong>Your journey:</strong></p><p>We are seeking an experienced Senior Infrastructure/Systems Engineer to help operate, maintain, and improve our on-premise infrastructure. You will work hands-on with networking equipment, Linux servers, Kubernetes, Redis.Ensuring that our platforms are secure, scalable, and highly available.</p><p>In this role, you will be responsible for the availability, scalability, and security of our systems and networks, ensuring smooth day-to-day operations, reliable deployments, and secure access for our teams.</p><p>This is a hands-on technical role, ideal for a technically driven engineer who thrives in complex, production-grade infrastructure environments and enjoys keeping systems stable, secure, and performant.</p><p>This position offers a hybrid work mode, combining on-site collaboration with flexible remote work.</p><p><br/></p><p><strong>Key Responsibilities:</strong></p><ul><li>Administer and maintain network infrastructure, including switches, routers, VPNs, and firewalls, ensuring reliable and secure connectivity across the datacenter.</li><li>Manage routing, switching, and secure remote access for internal teams and services.</li><li>Perform datacenter maintenance, including hardware installation, cabling, capacity planning, and coordination with datacenter and connectivity providers.</li><li>Operate, maintain, and harden Linux-based servers in our on-premise datacenter.</li><li>Manage server lifecycle activities, including provisioning, patching, upgrades, and decommissioning.</li><li>Manage and enhance Kubernetes clusters, including configuration, upgrades, scaling, and deployment automation using Helm and Docker.</li><li>Manage, monitor, and troubleshoot RabbitMQ clusters, ensuring message delivery reliability, scalability, and fault tolerance.</li><li>Administer and optimize Redis, Elasticsearch, and MariaDB/MySQL for performance, stability, and data integrity.</li><li>Support and execute database migration and infrastructure modernization projects.</li><li>Implement and maintain infrastructure-as-code practices using Terraform, Ansible, GitLab CI/CD, and Puppet.</li><li>Maintain and improve monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, Zabbix).</li><li>Contribute to incident response, root cause analysis, and on-call rotations, and help establish incident response strategies and procedures.</li><li>Collaborate with development and product teams to support deployment pipelines and performance optimization.</li><li>Ensure compliance with internal security and operational policies (e.g., GDPR, data protection).</li><li>Evaluate and maintain third-party tools, hardware, and services.</li></ul><p><br/></p><p><strong>We look for:</strong></p><ul><li>5+ years of hands-on experience in network, system, or infrastructure engineering, focused on on-premise environments.</li><li>Strong networking skills (TCP/IP, routing, switching, firewalls, VPNs, and secure access).</li><li>Strong expertise in Linux server administration (Debian/Ubuntu preferred.</li><li>Hands-on experience with datacenter operations and hardware maintenance.</li><li>Solid practical knowledge of Kubernetes, Docker, and Helm, including cluster management, deployments, upgrades, and troubleshooting.</li><li>Deep understanding of RabbitMQ, including clustering, high availability, performance tuning, and troubleshooting.</li><li>Experience with infrastructure automation and configuration management using Terraform, Ansible, GitLab CI/CD, and Puppet.</li><li>Experience with monitoring and observability tools (Prometheus, Grafana, Zabbix).</li><li>Excellent problem-solving and analytical skills, with attention to performance, reliability, and maintainability.</li></ul><p><br/></p><p><strong>Bonus Points / Nice to have</strong></p><ul><li>Certified Kubernetes Administrator (CKA) and/or CCNP certification.</li><li>Experience managing RabbitMQ, Redis, Elasticsearch, and MySQL/MariaDB at scale in production.</li><li>Experience with high-availability and distributed on-prem systems.</li><li>Familiarity with Rancher for Kubernetes cluster management, and Foreman for provisioning.</li><li>Knowledge of service meshes (e.g., Istio).</li><li>Familiarity with container security and system hardening best practices.</li><li>Knowledge of disaster recovery and business continuity planning for on-prem environments.</li></ul><p><br/></p><p><strong>Tech Stack</strong></p><p><em>Operating Systems</em></p><ul><li>Ubuntu</li></ul><p><em>Databases & Queues</em></p><ul><li>MariaDB</li><li>Elasticsearch</li><li>Redis</li><li>RabbitMQ</li></ul><p><em>Web & Proxy Services</em></p><ul><li>Nginx</li></ul><p><em>Orchestration & Automation</em></p><ul><li>Kubernetes</li><li>Terraform</li><li>Ansible, AWX</li><li>RKE (Rancher Kubernetes Engine)</li><li>Rancher</li><li>GitLab CI/CD</li><li>Puppet</li><li>Mesos / Marathon</li></ul><p><em>Monitoring & Observability</em></p><ul><li>Zabbix</li><li>Prometheus</li><li>Grafana</li><li>New Relic</li></ul><p><strong>What We Offer</strong></p><ul><li>Everything you need to do a great job (MacBook, etc.).</li><li>Free weekly German classes to help you settle into Berlin life.</li><li>Wellpass gym access and many other perks.</li><li>Flexible hours with hybrid working between our great offices and home.</li><li>A friendly, diverse group of colleagues from all different nationalities, genders, and orientations!</li></ul> </div>