Cloud Engineer 2, Site Reliability Engineering
KinaxisAbout Kinaxis: About Kinaxis Are you looking to join an innovative, market-leading company where you can truly elevate your career? At Kinaxis we are serious about culture, we are serious about technology, we are serious about customers, and we are serious about not taking ourselves too seriously. If you are looking to be part of an incredible growth story, then we might just be the place for you! In 1984, we started out as a team of three engineers. Today, we have grown to become a global organization with over 2000 employees around the world, 6 global office and a best-in-class HQ in Ottawa, Canada. As winners of several Top Employer awards globally, we are proud to work with our customers and employees towards solving some of the biggest challenges facing supply chains today. Kinaxis is a global leader in modern supply chain orchestration, powering complex global supply chains, and supporting the people who manage them. Our powerful, AI infused platform provides full transparency and visibility across end-to-end supply chains, enabling our customers to make faster, better decisions. We are trusted by renowned global brands to provide the agility and predictability needed to navigate today’s volatility and disruption. With more than 40,000 users in over 100 countries, we are expanding our team as we continue to innovate and revolutionize how we support our customers. About the team: Location Ottawa and Toronto, Canada - Hybrid Other Canadian locations - Remote About the team The Site Reliability Engineering team owns the delivery, operation, and monitoring of Kinaxis products and cloud infrastructure in production. We are focused on keeping services reliable, performant, and available for our global customers, 24x7. The team builds and operates the tooling and automation needed to support production at scale. This includes infrastructure-as-code, deployment pipelines, and platform-level automation using tools such as GitHub Actions, Terraform, ArgoCD, and Ansible. About the role We are seeking a Cloud Engineer II within the Site Reliability Engineering team to support the reliability, automation, and operational efficiency of the Kinaxis SaaS platform. In this role, you will contribute to building, operating, and maintaining systems that ensure our services remain available and stable in production. You will work closely with Product, Security, Platform, and Support teams to help deploy and operate services in a way that meets performance and reliability expectations. You will contribute to automation efforts and assist in translating technical designs into practical implementations. This is a hands-on technical role where you will develop your skills in cloud infrastructure, automation, and production operations. You will participate in incident response, troubleshooting, and continuous improvement efforts to enhance the reliability and operability of the platform About the role: Vacancy Status This is an existing job vacancy What you will do Support the operation of production systems to meet SLA targets and maintain service availability for customer workloads. Contribute to automation across multiple environments (Azure, GCP, and datacenters), including infrastructure as code (Terraform), deployment pipelines (GitHub Actions, ArgoCD), and scripting (Python, Bash, PowerShell, Ansible). Assist in managing the lifecycle of production and Hands-on Lab (HOL) environments, including deployment, resource management, and troubleshooting. Participate in an on-call rotation, working with the team to investigate and resolve incidents. Troubleshoot infrastructure and application issues, escalating and collaborating as needed to resolve problems effectively. Work closely with Product, Platform, and Support teams to support the operation and stability of services in production. Learn from and collaborate with senior engineers, contributing to team knowledge sharing and continuous improvement of practices. What we are looking for …