About this role
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve.
This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. Position Summary The Senior Scientific Data Curator will lead the systematic discovery, assessment, harmonization, and quality assurance of Lilly’s scientific datasets across the full breadth of TuneLab’s modeling domains—small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo pharmacokinetics and toxicology, antibody and biologics developability, and clinical PK/PD—in support of a strategic, cross-modality data unification initiative.
This role sits at the intersection of biological and pharmacological domain expertise and data science, translating decades of fragmented, heterogeneous datasets spanning discovery through the clinic into a unified, AI-ready data infrastructure. The curator will partner closely with computational scientists, DMPK scientists, pharmacometricians, antibody engineers, and external consortium collaborators to ensure that the data substrate underpinning TuneLab’s federated AI/ML models is comprehensive, well-documented, and scientifically sound. Core Responsibilities Data Discovery & Assessment Conduct comprehensive inventory of historical and ongoing datasets across TuneLab’s modeling domains—small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo PK and toxicology, antibody and biologics developability, and clinical PK/PD—spanning therapeutic areas (oncology, immunology, metabolic diseases, neuroscience, etc.) and 20+ years of discovery, preclinical, and clinical data Assess and score data quality, completeness, and integration feasibility for each dataset, accounting for the distinct data structures of each domain, including assay and dose–response measurements, concentration–time profiles and dosing regimens, in vivo study readouts, sequence- and structure-derived features for biologics, and biomarker and clinical covariate data Map metadata gaps across legacy systems and source platforms, documenting study contexts, assay and protocol methods, protocol deviations, data quality flags, and provenance information Develop automated pipelines (including LLM-assisted extraction where appropriate) to identify and extract domain-relevant data from internal documents, assay databases, and study reports into standardized, model-ready formats Produce a prioritized data assessment report recommending which domains, therapeutic areas, and indications to integrate first, based on data volume, complexity, portfolio relevance, and model feasibility Data Harmonization & Integration Design and implement standardized, extensible schemas for the integrated multi-domain database, working with computational partners to ensure AI/ML readiness across small-molecule and biologics modalities Build and maintain data harmonization pipelines: label normalization, unit and assay-condition standardization across studies and sources, time-point and dose alignment, sequence and structure normalization for biologics, and covariate encoding Apply domain-driven quality control practices—sequence validation, hidden duplicate detection, cross-source discrepancy resolution, and cross-species dataset integration using allometric scaling where applicable Develop and execute outlier detection protocols, flagging and adjudicating anomalous values in collaboration with clinical pharmacologists, DMPK scientists, toxicologists, and antibody engineers as appropriate to the domain Create reproducible data quality assurance workflows with documented acceptance criteria and audit trails Curate and enrich metadata to enable cross-study and cross-domain querying—linking compound and molecule identifiers, sequence and construct identifiers, assay methods, formulation details, and study design parameters Cross-functional Partnership Serve as the primary data domain expert for external consortium partners working within Lilly’s controlled cloud environment Collaborate with pharmacometricians, DMPK scientists, toxicologists, and antibody engineers to validate harmonized datasets against legacy models and established analyses (e.g., NONMEM/Monolix outputs for the clinical PK/PD domain) Work with the TuneLab ML team to ensure curated datasets meet the input specifications for the platform’s multi-task ML models, representation and foundation-model embeddings, and mechanistic/hybrid PK/PD frameworks (e.g., Neural ODE, SINDy) Contribute to platform deployment by supporting the development of data dictionaries, user documentation, and training materials for internal and consortium end users Required Qualifications M.S. or Ph.D. in Computational Data Science, Pharmacometrics, Pharmaceutical Sciences, Computational Biology, Cheminformatics, Biomedical Informatics, or a related quantitative discipline 1+ year of hands-on experience curating, harmonizing, or building analysis-ready datasets from biological, chemical, or clinical data sources Demonstrated skill in scientific dataset construction with domain-driven QC: sequence validation, duplicate detection, cross-source discrepancy resolution, or equivalent rigor applied to noisy real-world data Proficiency in Python and/or R for data wrangling, transformation, and quality checks at scale Working knowledge of pharmacological or chemical data structures across one or more TuneLab domains—for example, ADME/ADMET assay data, in vivo PK and toxicology readouts, antibody and biologics developability measurements, or clinical concentration–time and covariate data Track record of producing clear data documentation, quality reports, and data dictionaries Preferred Qualifications Experience integrating cross-species datasets (e.g., allometric scaling) or multi-source public/internal data to expand training sets for ML models Familiarity with cheminformatics and computational biology tooling, including biologics-specific tools (e.g., ANARCI, protein language model embeddings, molecular operating environment software) Exposure to LLM-assisted workflows for information extraction, document parsing, or automated data-pipeline development Experience with the data conventions of one or more TuneLab domains—population PK/PD modeling tools (NONMEM, Monolix, nlmixr) or CDISC standards (SDTM, ADaM) for clinical PK/PD; ADMET/DMPK assay conventions for small molecules; or developability assays for biologics Familiarity with cloud-based data infrastructure (AWS, Azure, or GCP) and version-controlled, reproducible analysis environments (Git, Docker, Conda) Prior experience providing curated data to federated learning or collaborative ML initiatives Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions.
If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form ( https://careers.lilly.com/us/en/workplace-accommodation ) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response. Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.
Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL). Actual compensation will depend on a candidate’s education, experience, skills, and geographic location.
The anticipated wage for this position is $132,000 - $244,200 Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees. #WeAreLilly
