About this role
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve.
This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. Associate Director — Reliability Engineering & Operations Job description About Lilly At Lilly, everything we do starts with patients.
We unite caring with discovery to make life better for people around the world. Headquartered in Indianapolis, Indiana, our global team of over 50,000 employees work with urgency and purpose to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. We bring our best to this work because people depend on it.
If you’re driven by purpose and determined to make a meaningful difference for patients, we invite you to bring your skill and your commitment to Lilly. About Tech@Lilly At Lilly, technology is not a support function. It is how a global medicine company operates, innovates, and delivers.
Lilly in Hyderabad builds the capabilities that make this possible, cloud platforms, AI systems, and automation at enterprise scale, all in service of a purpose that makes this technology work genuinely distinctive, from advancing drug discovery to enabling connected clinical trials to keeping a global medicine company running at the standard patients deserve. Experience 10+ years Location Hyderabad (Onsite) Employment Type Full-time Job Family M1 — Engineering Management / Site Reliability / Production Engineering About the technology organization Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable.
Within Tech@Lilly, the Digital Core organization applies a product, platform, and reliability-first mindset, ensuring that operational capabilities scale sustainably across the enterprise. About the Team Tech@Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose, creating medicines that make life better for people around the world, including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations.
We hire the best technology professionals from a variety of backgrounds, so they can bring an assortment of knowledge, skills, and diverse thinking to deliver creative solutions in every area of our business. The Digital Core team leads Lilly's transformation into the Digital and AI era. They inspire digitally empowered teams to new ways of working and accelerate innovation and agility.
This team powers and advances the entire company by building and maintaining world-class technology capabilities and platforms. Role summary As Associate Director of Reliability Engineering, you lead the engineering team accountable for the stability, observability, and operational quality of a multi-application production estate. You manage a team of senior reliability and production engineers — site reliability leads, production operations leads, principal production engineers, and quality engineers — and you own the operating rhythm that turns production signal into durable engineering work.
This is a first-line engineering management role with real production weight. You are the named owner for the reliability of your estate, the escalation point above shift leads during major incidents, and the manager who decides what your team stops doing manually and starts doing through code. You spend your time on people, priorities, and the engineering decisions that make reliability stick — not on tickets, status reports, or escalation theatre.
What you'll be doing 1) Lead the reliability engineering team Manage, coach, and develop a team of senior reliability and production engineers across site reliability, production operations, and quality engineering. Own hiring, performance, career development, and succession for the team; build a strong bench and credible engineering pathways for both engineers and team leads. Run the operating cadence — standups, on-call reviews, incident reviews, weekly reliability reviews — that keeps the team focused on outcomes, not noise.
Protect engineering time: ensure operational load is bounded, on-call is sustainable, and durable engineering work is funded, scheduled, and shipped. 2) Own reliability outcomes for the estate Hold the team accountable to service-level objectives, error-budget policy, and the engineering standards that make them real, not decorative. Govern on-call quality: page volume, time-to-acknowledge, time-to-recover, repeat-offender rate, and the ratio of toil to engineering work — and act on the numbers.
Be the named-owner escalation point above shift leads during major incidents. Run command discipline, drive communications to senior stakeholders, and ensure root-cause analyses produce engineering work — not narrative. Own the operational risk posture of the supported estate: what's fragile, what's improving, what needs investment, what should be decommissioned or returned to vendor — and carry that view into the portfolio conversation.
3) Drive the transition from manual operations to engineering Partner with the platform and automation engineering team on which production patterns become agent-assisted or fully automated remediations, and which require platform-side investment. Identify, prioritize, and sponsor the engineering work that retires toil at the source — instrumentation, self-healing runbooks, infrastructure-as-code coverage, deployment hardening. Hold a clear, defended point of view on what your team should stop doing manually in the next 6–12 months, what that requires in platform investment, and what it unlocks in reduced cost-to-serve.
Track and report the steady-state-ops curve: are we measurably moving from human-executed to engineering-executed work, quarter over quarter? 4) Compliance, security, and audit readiness Set and enforce the team's engineering standards for observability, alerting, change management, and incident response, and hold the line on them under delivery pressure. Ensure operational practices comply with applicable regulatory requirements (GxP, SOX, or equivalent), and produce credible audit evidence as a by-product of normal engineering work — not as a separate scramble.
Partner with security, validation, and change governance teams on controls that protect production without slowing engineering down. 5) Stakeholder partnership and global delivery Partner with application owners, product managers, and platform leaders on the reliability of their services — what the SLOs say, what the error budget allows, and what needs to change. Be willing and able to say no when the bar isn't met, and to show how to get there.
Coordinate with global reliability counterparts in other regions to deliver a seamless follow-the-sun operating model — including shift coverage, handoff quality, and consistent incident command across geographies. Represent the team's work to senior technology leaders with clear, data-backed narratives — what improved, what didn't, what's next, what it cost. How you will succeed At the first-line engineering management level, success is defined by team outcomes and the quality of the engineering operation you run: Major incidents trend down, time-to-recovery improves, and recurring failures get fixed at the root.
Your team operates a credible SLO and error-budget regime — not a dashboard, an operating discipline. Toil drops measurably as the team automates and engineers out the patterns it used to handle by hand. Your engineers grow — senior ICs sharpen, strong performers get stretched, the bench deepens.
Peer engineering leaders trust your team's judgment, your incident calls, and the engineering standards you hold. What you should bring Required 16+ years of progressive engineering experience, with at least 5 years managing engineering teams in Site Reliability, Production Engineering, DevOps, or Infrastructure Engineering. Hands-on management of a multi-application production estate — not a single product team — including on-call ownership, named-owner escalation for high-severity incidents, and accountability for measurable reliability outcomes.
Demonstrated experience running a 24×7 on-call practice across geographies, including shift coverage design, handoff discipline, escalation paths, and on-call sustainability. Working fluency with the SRE technical stack: service-level objectives and error budgets; observability and telemetry (OpenTelemetry, and at least one of Splunk, Datadog, New Relic, or Grafana/Prometheus); infrastructure-as-code (Terraform or equivalent); CI/CD and deployment hardening; container platforms; and at least one major public cloud (AWS, Azure, or GCP). Working technical understanding of full-stack application architecture — front-end, back-end services, APIs, databases, and messaging or integration layers — sufficient to reason about failure modes, performance bottlenecks, and remediation across the whole stack, not just the infrastructure beneath it.
Hands-on operational experience with SaaS and vendor-hosted applications — integration and API/webhook reliability, vendor SLAs and escalation paths, configuration drift, and the observability gaps that come with systems you don't fully control. Familiarity with automation and orchestration platforms — workflow and job orchestration, RPA or low-code automation, and scripting (Python, Go, or equivalent) — used to engineer out manual operational work and build self-healing remediation. Demonstrated ability to convert incident learning into durable engineering work — running a blameless postmortem culture and tracking that findings actually ship as code.
Track record of building and developing engineers — hiring senior ICs, growing junior engineers, and running performance and career conversations that move people forward. Working comfort in regulated or audited environments (GxP, SOX, HIPAA, PCI, or equivalent), including change control, audit evidence, and validated-system constraints. Bachelor's degree or higher in Computer Science, Information Technology, or a closely related engineering field.
Preferred Experience standing up a Site Reliability or Production Engineering practice from a small founding team — including hiring the first cohort, writing the first set of standards, and earning the right to set the engineering bar. Experience integrating AIOps or agent-assisted operations into a live production estate — including the guardrails, confidence thresholds, and human-in-the-loop boundaries that make automated remediation safe. Track record of driving an ops-to-engineering transition — measurably reducing manual operational load by investing in automation, instrumentation, and platform work, with the numbers to show it.
Experience leading or partnering across geographies in a follow-the-sun delivery model. Experience in pharma, healthcare, financial services, or other regulated industries. Master's degree in Computer Science, Engineering, or a related field.
Leadership expectations Leads engineers first and processes second — earns credibility by being technically present, not by managing from a distance. Combines engineering rigor with operational pragmatism, and knows when each is the right answer. Holds the line on reliability commitments and engineering quality, even under delivery pressure.
Makes the unglamorous calls — pausing a deploy, telling a product owner the bar isn't met, investing in toil reduction over a flashier feature — and makes them stick. Leads through the engineering bar they set and the people they grow, not through org-chart authority. Treats team capability growth as a first-class engineering outcome.
Additional information Availability to work flexible work hours is/may be required. This team supports continuous operations and may require non-standard work hours, including some work on weekends and holidays. Appropriate adjustments in benefits will be provided for employees working non-standard hours where applicable.
This is an onsite role based in Hyderabad. Candidates should be open to working different shifts when required to align with global delivery partners. At Lilly, caring is not only what we do for patients.
It is how we work. We believe the people who dedicate themselves to making medicines better deserve an environment that makes their lives better too, one where they are supported, respected, and given the space to do their best work. This is not just a policy.
It is who we are. Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form for further assistance.
Lilly does not discriminate on the basis of age, race, color, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability, or any other legally protected status. Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form ( https://careers.lilly.com/us/en/workplace-accommodation ) for further assistance.
Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response. Lilly does not discriminate on the basis of age, race, color, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability or any other legally protected status. #WeAreLilly
