Senior Site Reliability Engineer
Kurzfassung
Senior SRE at Akamai Technologies in Kraków. Responsibilities include managing AI platform reliability, automation, and incident response. Requires SRE/infrastructure expertise, Kubernetes, observability tools (Prometheus, Grafana), and Python/Go.
Schlüsselwörter
Von KI erstellte Kurzfassung des Anzeigentextes.
Do you enjoy solving complex reliability challenges for cutting-edge technology?
Do you have a passion for automation and building systems that scale?
Join the Akamai Inference Cloud Team
The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications with unmatched performance, compliance, and economics.
Partner with the bestAs a Senior SRE, responsibilities include owning reliability workstreams for Akamai's serverless inference platform, building automation and tooling, and contributing to architecture and operational decisions. Opportunities exist to take ownership of critical reliability problems end-to-end, partner with product engineering teams, and develop expertise in GPU infrastructure, Kubernetes at scale, and AI inference workloads.
As a Site Reliability Engineer, you will be responsible for:
Building and maintaining observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and driving improvements when targets are missed
- Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response
- Integrating AI workloads into Akamai's existing incident management processes, building runbooks, participating in on-call rotations, and conducting blameless post-mortems
- Building and maintaining CI/CD integrations, deployment safety checks, and rollback automation
- Collaborating with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases
- Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructureDo what you loveTo be successful in this role you will:
Demonstrate expertise in SRE, infrastructure, or platform engineering, managing large-scale distributed systems with extensive operational experience.
Demonstrate expertise in Kubernetes and large-scale containerization systems.
Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.
Demonstrate proficiency in Python or Go for automation, CI/CD pipelines, deployment safety, and infrastructure-as-code like Terraform.
Interest in or experience with AI/ML infrastructure, model serving, or GPU workloads
Resolve issues independently while maintaining accountability throughout the process.
Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.
Build your career at Akamai
Our ability to shape digital life today relies on developing exceptional people like you. The kind that can turn impossible into possible. We’re doing everything we can to make Akamai a great place to work. A place where you can learn, grow and have a meaningful impact.
With our company moving so fast, it’s important that you’re able to build new skills, explore new roles, and try out different opportunities. There are so many different ways to build your career at Akamai, and we want to support you as much as possible. We have all kinds of development opportunities available, from programs such as GROW and Mentoring, to internal events like the APEX Expo and tools such as Linkedin Learning, all to help you expand your knowledge and experience here.
Learn moreNot sure if this job is the right match for you or want to learn more about the job before you apply? Schedule a 15-minute exploratory call with the Recruiter and they would be happy to share more details.
| Veröffentlicht | 2026-08-27 |
| Quelle |
|
Hexjobs App
Auf diese Anzeige zugeschnittene Tools.
Hexjobs App
Auf diese Anzeige zugeschnittene Tools.
Ähnliche Stellen
Solutions Architect with AI SDLC
Future Processing
KrakówTest Automation Engineer with Playwright
Billennium
KrakówFullstack Developer (Go & Node.js/Python)
Clurgo
KrakówProgramista PHP / Laravel Developer
CStore
KrakówSenior Mobile Software Engineer (iOS) - Consumer Experience
Allegro
Kraków