Senior Site Reliability Engineer – DevOps / 24x7 Production
21 000 - 25 200
PLN
(B2B)
16 000 - 18 000
PLN
brutto (Employment Contract)
Keep production unstoppable — design, automate, and safeguard 24x7 services with SRE excellence.
Location & work model Krakow-based opportunity with a hybrid work model (up to 3 remote days per week).
As a Senior Site Reliability Engineer – DevOps / 24x7 Production, you will be working for our client as part of a global DevOps team, supporting highly available (24x7) production services. You will implement SRE best practices to strengthen service availability, performance, security, and transparency—while continuously improving the reliability of critical systems.
Your main responsibilities:
-
Support 24x7 production services by applying Site Reliability Engineering best practices to improve stability and uptime.
-
Resolve production incidents, perform root cause analysis, and facilitate post-incident reviews to prevent recurrence.
-
Contribute to software architecture design in order to build reliable and scalable solutions.
-
Drive end-to-end SDLC activities across requirements gathering, design, development, testing, deployment, and ongoing maintenance.
-
Define application SLIs and SLOs, and build & maintain observability to continuously monitor operational performance.
-
Create and execute action plans when failures occur to restore service health and meet target reliability metrics.
-
Plan and execute application/infrastructure migrations, disaster recovery exercises, and product upgrades.
-
Enhance automation and develop self-service capabilities to reduce manual effort and improve the user experience.
You're ideal for this role if you have:
-
Minimum 5 years of professional experience in Production Application Support or Site Reliability Engineering, including strong troubleshooting and issue prevention in high-pressure environments.
-
Proficiency with automation and build/monitoring tools such as Ansible, Jenkins, Prometheus, and Grafana.
-
Strong analytical and troubleshooting skills with a reliability-first mindset.
-
Solid engineering background with experience in one or more of: Java, Python, NodeJS, and SQL.
-
Knowledge of the Software Development Life Cycle (SDLC) and practical application of its principles.
-
Good communication skills and the ability to collaborate effectively with globally dispersed, cross-functional teams and vendors.
-
Experience supporting large-scale services, including providing on-call support as part of a rotation to respond rapidly to critical incidents.
It is a strong plus if you have: (optional)
- Prior experience supporting large Atlassian Jira and Confluence Data Centre instances.
- Solid knowledge of observability and monitoring practices using Grafana and Prometheus.
Language Required for the role :
- Fluent English (good command required for collaboration in a global team).
Eligibility for the role :
- Only candidates with an existing legal right to work in the European Union will be considered for this role.
#MAKEYourCareerBETTER
-
Interested? Apply now and include your CV (preferably in English) along with a statement confirming your consent to the processing and storage of your personal data.
Internal number #9408
Dodatki
ITDS Clubs
Access to medical insurance
Access to Multisport
Access to Pluralsight
Integrational Events
Ambassador Program
Twoje kryteria