Senior Site Reliability Engineer (SRE)
Location: Pittsburgh, PA / Cleveland, OH / Dallas, TX
FTE
Position Overview
We are seeking an experienced Senior Site Reliability Engineer (SRE) to support production operations, application reliability, performance management, and continuous improvement initiatives.
The selected candidate will work closely with production support and engineering teams to ensure critical internal and external applications maintain appropriate levels of availability, reliability, and uptime.
This role requires strong experience in production support, incident management, monitoring, troubleshooting, log analysis, automation identification, infrastructure technologies, databases, and application servers. The SRE will also provide technical leadership and collaborate with geographically distributed teams.
Key Skills
- Site Reliability Engineering (SRE)
- Production Support / Application Support
- Incident & Problem Management
- Linux
- Windows Server
- Oracle / PL/SQL / DB2
- Dynatrace / DT Managed
- GlassBox / ITCAM / TrueSight / OEM
- Tomcat / Apache / WebSphere (WAS) / IIS
- REST & SOAP Web Services
- Log Analysis & Troubleshooting
- AIOps / NLP
- Monitoring & Performance Management
- Automation
- Root Cause Analysis
- Business Analytics
- Agile
- Technical Leadership
- Client-Facing Production Support
- Monitor distributed systems and proactively identify potential production issues.
- Support troubleshooting and participate in on-call activities.
- Manage, track, and coordinate production incidents and application outages.
- Lead incident-analysis and problem-management meetings.
- Identify opportunities for operational and production-support automation.
- Monitor applications and related infrastructure to maintain system reliability.
- Coordinate follow-up activities through incident resolution and closure.
- Troubleshoot complex application issues using system and application logs.
- Participate in critical incident calls and contribute technical expertise toward resolution.
- Perform root cause analysis and recommend corrective actions.
- Research and reproduce user issues to validate solutions.
- Resolve technical problems that cannot be handled by junior team members.
- Provide technical guidance and solutions to the production-support team.
- Introduce process improvements and innovative solutions for operational challenges.
- Develop and maintain SOPs, operational procedures, and knowledge documentation.
- Collaborate with offshore and geographically distributed teams.
- Work with client technical teams, SMEs, and leadership.
- Support extended or weekend hours when required during critical production events.
- Participate in overlapping business-hour shifts for critical meetings and activities.
- 5+ years of overall IT experience.
- 2–3 years of business analytics and technical leadership experience.
- Strong experience with production/application support in a client-facing environment.
- Strong understanding of Site Reliability Engineering and production operations.
- Hands-on experience troubleshooting production applications and analyzing log files.
- Strong knowledge of system-management, monitoring, and support analytics tools.
- Experience with incident management, root cause analysis, and problem resolution.
- Strong understanding of AIOps and NLP concepts.
- Experience identifying opportunities for automation and process improvement.
- Strong problem-solving and analytical capabilities.
- Ability to recommend efficient and cost-effective technical solutions.
- Experience working with geographically distributed/onshore-offshore teams.
- Excellent client-facing verbal and written communication skills.
Strong knowledge of:
- Oracle
- PL/SQL
- DB2
Experience developing and consuming:
- REST APIs
- SOAP Web Services
Application Servers / Web Servers
Strong knowledge of:
- Tomcat
- Apache
- WebSphere (WAS)
- IIS
- Extensive experience with Linux
- Good understanding of Windows Server
- Linux and Windows server configuration and troubleshooting
Experience with monitoring tools such as:
- Dynatrace
- Dynatrace Managed / DT Managed
- GlassBox
- ITCAM / ITCAMS
- TrueSight
- Oracle Enterprise Manager (OEM)
- Agile methodology
- SOP and technical documentation
- Performance management
- System reliability and availability
- Production incident coordination
- Technical research and solution evaluation
- Process improvement
- Automation opportunity identification
- Strong stakeholder and client communication
#M1
#DI-CB2
#L1 - KB1
Ref: #404-IT Pittsburgh

