Curriculum vitae
Platform engineer with a background in scientific research computing. Rotkreuz, Switzerland · benjamin@rottler.io
Skills
Experience
- Lead a team of three engineers responsible for operating the company's Linux environment and driving platform-engineering initiatives
- Drive the renewal of legacy Linux systems, rebuilding servers on the automated, Ansible-based infrastructure
- Deliver a Linux-based development environment for application developers (Linux VMs, Podman, Development Containers)
- Own the Flowable workflow platform (Terraform, Helm, AKS), running in production for a product team
- Set up and evaluate Backstage as the company's internal developer portal, starting with the software catalog and templates for bootstrapping new applications
- Co-managed 120+ Linux VMs across production environments, including load balancers, web, database, FTP, and SMTP servers
- Designed and implemented a modular Ansible role structure, replacing manual server administration with fully automated, version-controlled deployments
- Migrated monitoring from Nagios to the Grafana LGTM stack: Alloy for centralized log and metrics collection, self-service dashboards for development teams, and synthetic testing of all company SaaS applications via the Prometheus Blackbox exporter
- Developed Terraform modules and Helm charts to deploy Flowable workflow engine instances on Kubernetes (AKS), with CI/CD pipelines for automated rollout of new versions
- Designed and built an on-prem hyperconverged Proxmox/Ceph cluster to host a new Linux-VM-based development environment for application developers
- Migrated legacy load balancers to HAProxy and hardened Rocky Linux 9 and Apache based on CIS benchmarks
- Supported development and customer-support teams and took part in the IT infrastructure on-call rotation
The ATLAS-BFG cluster is a high-throughput cluster (HTC) that allows users from local particle physics research groups and global members of the ATLAS collaboration to analyse collision data from the ATLAS experiment. The core of the cluster is a batch system with 3400 CPU cores and a distributed object storage system with 4 PB of storage space. The cluster is part of the World Wide LHC Computing Grid (WLCG).
- Participation in the operation and maintenance of the cluster
- Integration of new and existing services into the cluster's IaC solution.
- Maintenance and further development of the cluster's monitoring infrastructure
- Support for local users
ATLAS-BFG cluster · ATLAS experiment
StackPuppet, Packer, Docker, Prometheus, Grafana, Nagios, Slurm, CentOS 7, Alma 9
AUDITOR (AccoUnting DatahandlIng Toolbox for Opportunistic Resources) is a flexible and extensible framework for the accounting of shared and opportunistic resources.
- Planning of new features and technical management of contributing developers
- Release management: Publishing new releases, communication with users
- Operation of multiple AUDITOR instances within the ATLAS-BFG cluster
StackRust, Python, PostgreSQL, Docker, GitHub Actions
HammerCloud is a monitoring and testing framework, that is developed and used by the ATLAS collaboration to monitor the availability and performance of the global grid infrastructure of the WLCG.
- Migrating from Python 2/Django 1 to Python 3/Django 4
- Migration to CERN's new OIDC-based single sign-on solution
- Redesign of the architecture to a container-based setup
StackPython, Django, Docker, Docker Compose
- Measurement of the cross section and signal strength of non-resonant Higgs-boson pair-production
- Measurement of the trilinear Higgs-boson self-coupling
- Analysis of ~100TB of data
- Improving the event selection with neural networks
- Statistical analysis (confidence intervals, hypothesis testing)
- Improved sensitivity by a factor of three compared to previous analysis
- Six-month research stay at CERN
- Correction of weekly exercises and final exams
- Discussion of the exercises with the students in weekly tutorials
Assisting the administrators of the ATLAS-BFG cluster