Full vacancy
About the job
Join Telavox as a Systems Engineer for Storage Are you a hands-on systems engineer who thrives in operational depth and wants to own the storage that the rest of engineering runs on? At Telavox, we run our own: our own Ceph clusters, on-prem Kubernetes, and a fleet of physical servers across six countries. We're providers of infrastructure, not consumers of cloud - a deliberate choice, and what our product and economics demand. Storage is where we're investing next. Ceph specifically is not a requirement. If you've run distributed or enterprise storage in production - Ceph, ZFS, Gluster, Lustre, MinIO, a vendor array - the instincts are what we're after, and we'll teach you ours. About the job This is a hands-on operations role. You'd own how our storage runs day-to-day - health, capacity, upgrades, expansion, and the operational foundations underneath it. You'll have real input into architecture, and you'll set the long-term direction with our storage lead rather than on your own. The emphasis here is deliberately operational - if making storage actually work, and being trusted to do it unsupervised, is the part you want, there's a lot of room in this role. Ceph (block and S3) is our main estate - real production storage, growing across regions, with Sweden the largest footprint. Your day-to-day would be mostly spent on things like:
Keeping the clusters healthy - issues, monitoring, capacity planning, upgrades, and expansion across regions
Getting the foundations right: stability, security, predictable scaling.
Making the work repeatable - runbooks, upgrade procedures that hold across regions, automation for what's manual today
You won't work in isolation from the rest of the platform: on-prem Kubernetes runs most of our workloads and has clear impact on the storage domain, and our Ansible-driven server and OS fleet spans six countries
How we work together: You’ll join our Infrastructure Engineering team of around 15 people, working from our Malmö HQ, Gjuteriet, with international colleagues working remotely. English is our working language, and Swedish is a bonus. There’s no mandatory on-call rotation, although Compute and Storage have an informal secondary list. You’ll have real ownership from day one, with plenty of room to identify, propose, and implement improvements. We also use Claude Code across Engineering, and you’ll help explore how it can support our platform work. About you You treat automation as the direction and repetition as its trigger, you factor in risk from the start rather than discovering it at the end, and you're pragmatic enough to do the manual work until automation catches up. We're looking for:
Real operational depth in distributed or enterprise storage - a system you've run in production, been on the hook for, and debugged under pressure. Which system matters far less than how deeply
