Job Description
We are looking for an experienced Site Reliability Engineer (SRE) with strong Production Support experience and hands-on expertise in AI/LLM technologies, Kubernetes, and Docker. The ideal candidate will be responsible for maintaining highly available production environments, troubleshooting critical issues, improving system reliability, and supporting AI/LLM-based applications and platforms.
Key Responsibilities
- Provide 24x7 production support for critical applications and services, including incident response, troubleshooting, and root cause analysis.
- Monitor system health, application performance, availability, and reliability across production environments.
- Deploy, manage, and troubleshoot applications running on Kubernetes and Docker environments.
- Support AI/ML and LLM-based applications, APIs, and services in production.
- Troubleshoot application, infrastructure, container, networking, and performance-related issues.
- Participate in incident, problem, and change management processes.
- Perform Root Cause Analysis (RCA) and implement permanent solutions to recurring production issues.
- Develop and maintain automation scripts using Python, Bash, or Shell to reduce manual operational activities.
- Work with cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Implement monitoring, logging, alerting, and observability solutions for production workloads.
- Collaborate with Development, DevOps, AI/ML, and Infrastructure teams to improve system reliability and deployment processes.
- Support CI/CD pipelines and automated deployment processes.
- Participate in on-call rotations and handle critical production incidents within defined SLAs.
- Continuously improve system scalability, performance, availability, and operational efficiency.
Required Skills
- 5+ years of experience in SRE, Production Support, DevOps, or Site Reliability Engineering.
- Strong hands-on experience with Kubernetes/K8s and Docker.
- Experience supporting production applications and critical incidents.
- Hands-on exposure to LLM, Generative AI, AI/ML applications, or AI platforms.
- Strong Linux/Unix administration and troubleshooting skills.
- Experience with Python, Bash, or Shell scripting.
- Experience with at least one cloud platform: AWS, Azure, or Google Cloud Platform.
- Knowledge of CI/CD, Git, monitoring, logging, and observability.
- Strong understanding of application performance, system reliability, scalability, and high availability.
- Excellent troubleshooting, analytical, and communication skills.
Preferred Skills
- Experience supporting LLM/GenAI applications in production.
- Knowledge of LLM APIs, model serving, inference, or AI platforms.
- Experience with Prometheus, Grafana, ELK/EFK, Splunk, Datadog, or similar monitoring tools.
- Experience with Terraform or Ansible.
- Knowledge of microservices and REST APIs.
- Experience with Helm and Kubernetes deployments.
- Understanding of networking, DNS, TCP/IP, load balancing, and cloud infrastructure.

