About this role
Careers that change lives start here. Medtronic is a global leader in healthcare technology with a Mission to alleviate pain, restore health, and extend life. Our 95,000 employees work across more than 150 countries to put patients first — developing innovative medical technologies that improve the lives of 72+ million patients each year.
Your unique talents will help shape the future of healthcare while building a career grounded in purpose, growth, and impact. A Day in the Life This position as Platform Engineer focuses on continuing to innovate, extend, and deploy our internally developed platforms based on opensource standards used for deploying and managing Gitlab, Github, ADO and JFrog Artifactory services; these services are currently hosted in both our internal datacenters and in their respective vendor SaaS platforms. Your assignment is to meet the technology and streamlined platform needs for our Medtronic’s business units source code management or binary release management.
Our platform is built and managed using programmatic pipelines developed and managed using a variety of software code and scripting languages, including but limited to, Terraform, Golang, CloudFormation, Jenkins and GitLab CI/CD, Lambda, some Python, and Java scripting Responsibilities may include the following and other duties may be assigned. Start of the day: Review application and infrastructure health, overnight alerts, incidents, SLOs, and critical service dashboards. Monitor & analyze: Analyze metrics, logs, traces, and events across Kubernetes, cloud, applications, APIs, and databases.
Incident response: Investigate alerts and production issues, correlate telemetry, identify root causes, and drive rapid resolution. Observability engineering: Build and enhance dashboards, alerts, SLOs, distributed tracing, log pipelines, and application monitoring. Automation: Automate monitoring onboarding, alert configuration, dashboard creation, and operational tasks using Terraform, APIs, Python, Helm, and GitOps.
Developer enablement: Partner with application teams to instrument services and establish standardized observability patterns. Reliability improvement: Identify recurring failures, performance bottlenecks, alert noise, and capacity risks; implement preventive solutions. Platform optimization: Continuously improve Prometheus, Grafana, Dynatrace, OpenTelemetry, and logging platforms for scalability and reliability.
Collaboration: Work closely with DevOps, Platform, Application, Cloud, Network, and Security teams during incidents and engineering initiatives. End of day: Review reliability trends, document RCA/action items, track SLOs, and identify opportunities for further automation and proactive monitoring. Required Knowledge and Expertise: Strong experience in SRE, Observability, Production Engineering, or Platform Engineering .
Hands-on experience with Prometheus, Grafana, Solarwinds, Dynatrace, OpenTelemetry, ELK/OpenSearch , or similar platforms. Strong understanding of Kubernetes and cloud-native architectures . Experience with metrics, logs, traces, events, APM, distributed tracing, and telemetry pipelines.
Strong knowledge of SLI, SLO, SLA, error budgets, availability, latency, saturation, and reliability engineering . Experience developing monitoring and alerting for microservices, APIs, containers, databases, and infrastructure. Proficiency in Python, Bash, or PowerShell for automation.
Experience with Terraform, Ansible, Helm, GitLab/GitHub, CI/CD, and GitOps . Strong understanding of networking, Linux, DNS, load balancing, HTTP/HTTPS, and cloud infrastructure. Experience with incident management, RCA, problem management, and production support.
Strong analytical, troubleshooting, communication, and stakeholder-management skills. Nice To Have: 2+ years of IT experience with a Bachelor's Degree in Engineering, MCA, or MSc. AWS/Azure/GCP observability experience.
Kubernetes platforms such as EKS, AKS, OpenShift, or Nutanix Kubernetes . OpenTelemetry Collector and telemetry pipelines. PromQL, LogQL, KQL, or equivalent query languages.
Dynatrace DQL, dashboards, workflows, and automation . Experience with Grafana Loki, Tempo, Jaeger, Fluent Bit/Fluentd . Experience implementing AI/ML-assisted observability and AIOps .
Knowledge of chaos engineering and resilience testing. Experience building enterprise-scale Observability-as-a-Service platforms. Physical Job Requirements The above statements are intended to describe the general nature and level of work being performed by employees assigned to this position, but they are not an exhaustive list of all the required responsibilities and skills of this position. Recruitment Fraud Alert We are aware of phishing scams targeting job seekers.
Please keep the following in mind: Apply only through official Medtronic channels. All legitimate Medtronic recruiting communications come from approved Medtronic platforms and official @medtronic.com email addresses. Medtronic will never ask for payment or sensitive personal information (such as bank account or Social Security details) during early stages of the hiring process.
Any such requests are not legitimate. If you receive a suspicious message claiming to be from Medtronic, do not respond, click links, or open attachments. If you have any questions, concerns regarding the authenticity of a communication alleged to have been made by or on behalf of Medtronic, please contact us immediately at AskHR@medtronic.com .
Benefits & Compensation Medtronic offers a competitive Salary and flexible Benefits Package A commitment to our employees lives at the core of our values. We recognize their contributions. They share in the success they help to create.
We offer a wide range of benefits, resources, and competitive compensation plans designed to support you at every career and life stage. This position is eligible for a short-term incentive called the Medtronic Incentive Plan (MIP).
