Describe your role

Job Offers For You

Back to All Offers

How to use the job offers’ basket?

Checkout

New

Senior Site Reliability Engineer – DevOps / 24x7 Production

Krakow
Hybrid

21 000 - 25 200

PLN

(B2B)

16 000 - 18 000

PLN

gross (Employment Contract)

English
Senior
Banking
Devops

Keep production unstoppable — design, automate, and safeguard 24x7 services with SRE excellence.
Location & work model Krakow-based opportunity with a hybrid work model (up to 3 remote days per week).

As a Senior Site Reliability Engineer – DevOps / 24x7 Production, you will be working for our client as part of a global DevOps team, supporting highly available (24x7) production services. You will implement SRE best practices to strengthen service availability, performance, security, and transparency—while continuously improving the reliability of critical systems.

Your main responsibilities:

  • Support 24x7 production services by applying Site Reliability Engineering best practices to improve stability and uptime.

  • Resolve production incidents, perform root cause analysis, and facilitate post-incident reviews to prevent recurrence.

  • Contribute to software architecture design in order to build reliable and scalable solutions.

  • Drive end-to-end SDLC activities across requirements gathering, design, development, testing, deployment, and ongoing maintenance.

  • Define application SLIs and SLOs, and build & maintain observability to continuously monitor operational performance.

  • Create and execute action plans when failures occur to restore service health and meet target reliability metrics.

  • Plan and execute application/infrastructure migrations, disaster recovery exercises, and product upgrades.

  • Enhance automation and develop self-service capabilities to reduce manual effort and improve the user experience.

You're ideal for this role if you have:

  • Minimum 5 years of professional experience in Production Application Support or Site Reliability Engineering, including strong troubleshooting and issue prevention in high-pressure environments.

  • Proficiency with automation and build/monitoring tools such as Ansible, Jenkins, Prometheus, and Grafana.

  • Strong analytical and troubleshooting skills with a reliability-first mindset.

  • Solid engineering background with experience in one or more of: Java, Python, NodeJS, and SQL.

  • Knowledge of the Software Development Life Cycle (SDLC) and practical application of its principles.

  • Good communication skills and the ability to collaborate effectively with globally dispersed, cross-functional teams and vendors.

  • Experience supporting large-scale services, including providing on-call support as part of a rotation to respond rapidly to critical incidents.

It is a strong plus if you have: (optional)

  • Prior experience supporting large Atlassian Jira and Confluence Data Centre instances.
  • Solid knowledge of observability and monitoring practices using Grafana and Prometheus.

Language Required for the role :

  • Fluent English (good command required for collaboration in a global team).

Eligibility for the role :

  • Only candidates with an existing legal right to work in the European Union will be considered for this role.

#MAKEYourCareerBETTER

  • Interested? Apply now and include your CV (preferably in English) along with a statement confirming your consent to the processing and storage of your personal data.

Internal number #9408

Benefits

ITDS Clubs

Access to medical insurance

Access to Multisport

Access to Pluralsight

Integrational Events

Ambassador Program

Your criteria

Want to get notified about new job openings
that match the above criteria?

By providing e-mail address and clicking „Notify Me” you agree that ITDS will send e-mail notification that match criteria you have provided.