US Jobs US Jobs     UK Jobs UK Jobs     EU Jobs EU Jobs


Site Reliability Engineer

The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, performance, and scalability of both internal (in-house) and customer environments.

This role acts as a critical bridge between the product platform engineering, development and customer operation team, ensuring seamless deployment, operation, and support of our solutions across many environments.

DUTIES AND RESPONSIBILITIES Maintains and optimizes availability, performance, and resilience of in-house and on-premises productive environments.

Monitors system health, troubleshoots incidents, and implements proactive reliability solutions.

Manages deployments, upgrades, and patching platform level changes of multiple customers.

Demonstrates a customer-focused approach and ownership over production systems.

Serves as technical liaison to customers (internal and external), explaining TGWs network, software and hardware platform architecture along with delivery and deployment approach.

Performs hands-on standing up of customer environments, and ongoing operations.

Develops and tests high-availability and disaster recovery plans for customers.

Engages with project management and customers on infrastructure and platform topics.

Collaborates with Platform Engineering and Global Infrastructure teams to ensure TGWs standard deliverable is continuously improved with customer real world feedback.

Partners with solutions architects to ensure successful delivery of customer projects.

Documents issues and best practices for incident prevention and resolution.

Improves observability (logging, metrics, alerting) across customer environments.

Drives root cause analysis and collaborates with respective teams to implement preventative measures.





Share Job