Zum Inhalt springen
Craftware

Site Reliability Engineer

Warszawa
vor 5 Tagen

Kurzfassung

Craftware seeks an experienced Site Reliability Engineer (SRE) in Warsaw for a full-time, remote role. Responsibilities include ensuring system reliability, monitoring business processes, leading incident response, and driving automation. Experience with Salesforce platforms is required.

Von KI erstellte Kurzfassung des Anzeigentextes.

Craftware is a technology company of over 500 experts, empowering large organizations to solve complex business challenges with modern IT solutions - from sales systems and automation to data platforms and AI. We operate where technology must be reliable, secure, and scalable. We deliver end-to-end projects: from analysis and architecture through implementation to development and maintenance. We are a trusted partner of industry leaders such as Salesforce, Veeva, UiPath, and Databricks.

Model: Remote Engagement: Full-time (B2B)

We are looking for an experienced Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational continuity of a complex enterprise ecosystem consisting of multiple highly integrated systems. The SRE will be responsible for the technical stability of individual applications as well as the end-to-end reliability of critical business processes spanning multiple platforms and services, working closely with DevOps teams, application owners, architects, support teams, and Single Points of Contact (SPOCs) responsible for integrated and dependent systems mostly based on Salesforce platform with Sales Cloud, Service Cloud and Experience Cloud.

Responsibilities

  • Ensure high availability, reliability, performance, and operational continuity of business-critical systems and services
  • Monitor and analyze end-to-end business processes across multiple applications, APIs, middleware components, messaging platforms, and external systems
  • Identify system dependencies and assess their impact on business service availability
  • Establish and maintain monitoring, observability, alerting, and dashboards covering both technical and business-process metrics
  • Define and track SLIs, SLOs, availability targets, and other reliability metrics
  • Coordinate major incidents involving multiple systems and technical teamsLead root cause analysis for production incidents, integration failures, performance degradation, and service interruptions
  • Coordinate troubleshooting and communication with SPOCs, system owners, vendors, infrastructure teams, and external providers
  • Manage and prioritize the work of the DevOps team responsible for deployment, monitoring, automation, infrastructure, and operational support
  • Drive automation of operational activities, deployments, health checks, recovery procedures, and system maintenance
  • Maintain operational runbooks, troubleshooting guides, escalation paths, and recovery procedures
  • Ensure appropriate backup, disaster recovery, failover, and business continuity mechanisms are implemented and validated
  • Support release planning, production readiness, risk assessment, dependency analysis, and rollback strategies
  • Proactively identify reliability risks, performance bottlenecks, single points of failure, and architectural weaknesses
  • Work with development and architecture teams to improve resilience, scalability, fault tolerance, retry mechanisms, and graceful degradation
  • Lead post-incident reviews and ensure corrective and preventive actions are implemented
  • RequirementsA minimum of intermediate proficiency with the Salesforce platform is required. Strong experience in Site Reliability Engineering, DevOps, Production Engineering, Application Operations, or a similar role
  • Experience working with complex, highly integrated enterprise architectures
  • Strong understanding of end-to-end business process monitoring and dependency management
  • Experience with incident management, root cause analysis, problem management, and service restoration
  • Practical knowledge of monitoring, logging, alerting, and observability platforms
  • Good understanding of SLI, SLO, SLA, availability, latency, throughput, and reliability concepts
  • Experience with CI/CD, release management, infrastructure automation, and deployment processes
  • Understanding of high availability, disaster recovery, failover, and resilience patterns
  • Ability to coordinate technical activities across multiple teams and system owners
  • Experience managing or coordinating a DevOps or operations-focused engineering team
  • Strong analytical, troubleshooting, and communication skillsNice to have
  • Experience with cloud platforms (AWS, Azure, or GCP) in a production operations context
  • Experience with containerization and orchestration (Docker, Kubernetes)

Scripting/automation skills (Python, Bash, or similar)

Familiarity with APM/observability tooling (e.g., Datadog, Dynatrace, New Relic, Grafana, Prometheus)

Experience with messaging/integration middleware (e.g., Kafka, MQ, ESB platforms)ITIL or similar IT service management framework knowledgeWe offerB2B contract (rate up to 190 PLN net/h + VAT)

Fully remote service delivery

  • Broad range of projects (internal, international) - genuine variety of clients and tasks
  • Budget for skills development and certifications as part of the collaboration
  • Regular collaboration reviews and discussion of project scope
  • Additional benefits available as part of the collaboration
  • Networking and team-building events for project teams
Veröffentlicht 2026-09-14
Läuft ab 2026-11-22
Quelle