Перейти до вмісту
DataArt

Senior Site Reliability Engineer

Lublin
Warszawa
Łódź
Kraków
+1
1 місяць тому

Коротко

Senior SRE role in Lublin, PL. Enhance platform observability, automate operations, improve CI/CD. Requires Python, Linux, cloud (AWS/GCP/Azure), Docker, Kubernetes, Terraform, CI/CD, and monitoring tools. On-call rotation included.

Стислий виклад підготував ШІ на основі тексту оголошення.

ClientOur client is developing a reliable, scalable, and user-friendly ticketing and streaming platform for high school sports. Their goal is to create a solution that allows parents, students, and fans to purchase tickets and stream live events effortlessly, ensuring accessibility from any device, anywhere, at any time.

Working Schedule: 8:00–17:00 Eastern Standard Time (EST). A 4-hour overlap with EST is required for effective collaboration.

Project overview

The project focuses on building and maintaining a highly available, scalable cloud platform that supports business-critical services. Engineering teams continuously improve system reliability, performance, and operational excellence through automation, observability, and modern software delivery practices.

Position overviewAs a Senior Site Reliability Engineer, you will work at the intersection of software engineering and operations, partnering with application, DevOps, and QA teams. You'll enhance observability, automate operational processes, improve CI/CD workflows, and drive reliability initiatives that enable teams to deliver resilient, high-performing software at scale.

Responsibilities

Enhance platform observability by designing and maintaining metrics, alerts, dashboards, and monitoring capabilities that improve system visibility and reduce incident resolution time.

Build and maintain automation, operational tooling, and monitoring solutions that increase service reliability and uptime.

Work closely with software development and QA teams to embed reliability best practices into software delivery, release processes, and testing strategies.

Promote operational excellence by driving preventive measures, facilitating blameless post-incident reviews, and supporting capacity and scalability planning.

Take part in an on-call rotation, ensuring timely investigation and resolution of production incidents affecting critical services.

Requirements

Strong hands-on experience with Python, particularly for scripting, automation, and operational tooling.

Proficiency in at least one of the following programming languages: Java, C++, or Go.

Solid knowledge of Linux environments, cloud platforms (AWS, GCP, or Azure), and containerized infrastructure using technologies such as Docker, Kubernetes, and Terraform.

Experience designing and maintaining CI/CD pipelines, working with version control systems, and implementing automated testing practices.

Practical experience with observability and monitoring platforms (such as Prometheus, Grafana, ELK, Datadog, or similar), including troubleshooting through log and metric analysis.

Experience identifying and documenting Critical User Journeys and translating them into measurable SLA/SLO objectives that support automation and operational excellence.

Strong collaboration and communication skills, with the ability to work effectively across multidisciplinary engineering teams, especially during critical production events.

A reliability-first mindset with the belief that system stability is a shared responsibility across engineering teams.

Familiarity with AI-assisted engineering tools (such as Claude and Codex) and their use within modern software development workflows.

Nice to have

Experience developing or maintaining end-to-end and integration tests for distributed or microservices-based systems.

Knowledge of performance optimization, capacity management, or chaos engineering practices.

Experience contributing to internal developer platforms, automation tools, or reliability engineering initiatives.

Understanding of production security, compliance requirements, or change management processes.

Relevant industry certifications.

What We Offer:

Vacation days: Up to 26 business days per year.10 illness/special days off per year (fully paid, no medical papers needed) for all contract types

Health and life insurance (Luxmed)

MyBenefit platform with Multisport option

Internal psychological support service

English language classes from the first working dayAccess to external learning platforms: O’

Reilly, LinkedIn Learning, Udemy, and a wide catalog of diverse internal training

  • Flexible workplace: work from the office, from home, or choose a hybrid optionTech Skills Mentoring Program
  • Opportunities to develop as a public speaker, mentor, or technical interviewer
  • Fully paid idle (bench) when not involved in a project
  • Certification reimbursement (AWS, GCP, Microsoft, etc.)
Опубліковано 2026-08-07
Діє до 2026-10-27
Джерело