Site Reliability Engineer - Dundalk
We are looking for a Site Reliability Engineer (SRE) to ensure the reliability, scalability, and security of our cloud-based ERP solutions. You will collaborate with software engineers, DevOps teams, and product teams to improve system performance, automate operations, and enhance observability.
Responsibilities:
- Infrastructure Management: Design and maintain scalable, cloud-based architectures.
- Automation: Develop scripts using Python, Unix Shell, and PowerShell to optimize deployments and monitoring.
- System Observability: Implement logging, monitoring, alerting, and performance tuning best practices.
- Containerization & Orchestration: Manage applications using Docker, Kubernetes, and related tools.
- Incident Response: Diagnose and resolve infrastructure and network issues proactively.
- Security & Reliability: Work with development teams to integrate best practices into the software lifecycle.
- On-Call Support: Participate in rotations and improve system resilience.
Requirements:
- Bachelor's/Master?s in Computer Science or related field.
- 3+ years of experience with cloud technologies, especially AWS.
- Strong expertise in system administration, automation, and scripting (Python, Unix Shell, PowerShell).
- Solid knowledge of Linux, Windows, networking (DNS, firewalls, load balancing).
- Hands-on experience with Docker, Kubernetes, and container orchestration.
- Familiarity with Elasticsearch, SRE principles, DevOps, and DevSecOps.
- Problem-solving mindset and ability to design scalable solutions.