Senior Site Reliability Engineer (Data Platform)
In short
Senior Site Reliability Engineer (Data Platform) needed for a global leader in cloud and content delivery. Focus on reliability of massive data pipelines, distributed databases (Cassandra, MongoDB), observability (Prometheus, Grafana), and automation (Python, Go, Java). Requires 5+ years of SRE experience and strong Linux/Unix skills.
AI-written summary based on the listing content.
On behalf of our client, a global leader in cloud and content delivery, we are seeking a highly skilled Site Reliability Engineer to join a critical team at the heart of their global infrastructure.
This team is responsible for a massive data pipeline that collects telemetry from hundreds of thousands of servers and delivers this vital data to customers for analytics and reporting. This is a role for a true systems engineer who understands that reliability at this scale is fundamentally a data problem.
We are not looking for a typical DevOps engineer. We need a deep-thinking SRE who is passionate about the reliability of large-scale data systems, observability, and solving complex problems at the intersection of software, network, and infrastructure.
What You Will Do (Your Impact):
Engineer World-Class Data Systems: Your core focus will be the reliability and performance of massive data platforms. You will work extensively with large-scale distributed databases to ensure data integrity, availability, and low-latency performance. While the environment heavily utilizes Cassandra and MongoDB, your expertise with other modern NoSQL or distributed data systems will be highly valued.
Become the Ultimate Troubleshooter: You will be the highest technical escalation point for the most complex reliability and performance issues, leading investigations that span the entire global stack.
Build Insightful Observability: You will design and build the observability fabric that allows us to understand the health of the platform in real-time, using tools like Prometheus and Grafana to create actionable insights, not just noise.
Solve Problems with Code: You will write high-quality automation and internal tooling using Python, Go, or Java. Your code will streamline operations, automate diagnostics (increasingly with AI assistance), and empower other teams with self-service workflows.
Drive Long-Term Reliability: You will partner closely with Engineering, Product, and Network teams to influence architecture, identify systemic weaknesses, and deliver scalable, long-term solutions that prevent future incidents.
Who We're Looking For (Your Profile):
You have deep, hands-on experience engineering and operating large-scale, distributed data systems. You are comfortable working with data at scale, analyzing metrics, and troubleshooting complex data integrity issues.
Practical experience with modern NoSQL databases is essential. We have a strong preference for candidates who have worked with Cassandra or MongoDB, but we are open to experts in other similar technologies.
You are a seasoned SRE or Systems/Infrastructure Engineer with at least 5 years of experience managing mission-critical distributed systems.
You are a proficient programmer with experience building automation and tooling in languages like Python, Go, or Java.
You have a strong foundation in Linux/Unix administration and low-level system troubleshooting.
You have a relentless drive to find the root cause of complex problems and deliver robust, production-grade solutions.
| Published | 2026-09-04 |
| Expires | 2026-11-29 |
| Source |
|
Hexjobs App
Tools tailored to this listing.
Hexjobs App
Tools tailored to this listing.
Similar offers
Solutions Architect with AI SDLC
Future Processing
KrakówTest Automation Engineer with Playwright
Billennium
KrakówFullstack Developer (Go & Node.js/Python)
Clurgo
KrakówProgramista PHP / Laravel Developer
CStore
KrakówSenior Mobile Software Engineer (iOS) - Consumer Experience
Allegro
Kraków