
Career Category Engineering Job Description Role Description We are seeking an experienced Senior Data Engineer to lead the design, development, and delivery of scalable enterprise data solutions. This role will build and optimize batch and real-time data pipelines, reusable integration frameworks, and governed data platforms that support analytics, AI, and self-service data access.
The ideal candidate has deep expertise in Databricks, Apache Spark, AWS, data modeling, governance, and production operations. The role also provides technical leadership, defines engineering standards, mentors other engineers, and partners with architecture, business, analytics, data science, and DevOps teams.
Experience in manufacturing, biotechnology, pharmaceutical, life sciences, or another regulated industry is preferred. Roles and Responsibilities Lead the design and development of scalable batch and real-time ETL/ELT pipelines.
Own complex data solutions from requirements and design through deployment and production support. Build cloud-based lakehouse and data-platform solutions using Databricks and AWS.
Develop reusable, metadata-driven data integration frameworks. Integrate structured, semi-structured, and unstructured data from enterprise, manufacturing, API, and third-party sources.
Optimize Spark workloads, Databricks compute, SQL queries, partitioning, storage, and caching for performance and cost. Implement workflow orchestration, monitoring, alerting, data-quality controls, and recovery processes.
Develop CI/CD pipelines and automated testing for data solutions. Implement metadata management, lineage, cataloging, governance, RBAC, and data-security controls.
Define data models, data contracts, integration patterns, and reusable engineering standards. Lead architecture reviews, code reviews, troubleshooting, and root-cause analysis.
Mentor engineers and provide technical guidance across delivery teams. Collaborate with architects, analysts, data scientists, product teams, and DevOps teams.
Support estimation, sprint planning, technical roadmaps, and delivery-risk management. Evaluate emerging technologies and recommend solutions based on scalability, security, maintainability, and business value.
Participate in operational support, including occasional off-hours support. Basic Qualifications One of the following: Master’s degree in Computer Science , Engineering, Information Technology, Data Science, or a related field and at least 7 years of relevant experience.
OR Bachelor’s degree in a related field and at least 9 years of relevant experience. Must-Have Skills Advanced hands-on experience with Databricks, Apache Spark, PySpark , Spark SQL, Delta Lake, Python, and SQL.
Experience designing and operating production-grade batch and streaming pipelines. Strong understanding of distributed computing, lakehouse architecture, data warehousing, and data integration.
Experience with Databricks Workflows or comparable orchestration tools. Strong experience with AWS data, compute, storage, security, and monitoring services.
Experience with Spark performance tuning, cluster optimization, partitioning, and cost management. Experience with Git, CI/CD, automated testing, monitoring, and production deployment.
Experience implementing data quality, metadata management, lineage, governance, and access controls. Strong understanding of RBAC, least privilege, encryption, auditability, and regulated-data requirements.
Experience leading technical design, code reviews, and complex production implementations. Ability to define reusable patterns, engineering standards, and development best practices.
Strong communication, collaboration, mentoring, and problem-solving skills. Experience working in Agile or Scaled Agile delivery environments.
Preferred Qualifications Experience in biotechnology, pharmaceutical, life sciences, manufacturing, or another regulated industry. Experience with Unity Catalog, data products, Data Fabric, Data Mesh, or similar enterprise data architectures.
Experience building APIs and secure data services. Experience with relational, NoSQL, analytical, operational, or vector databases.
Experience with OLAP and OLTP data modeling and performance tuning. Experience with Kafka, Kinesis, or other streaming technologies.
Experience supporting AI and Generative AI solutions, including RAG, embeddings, vector search, and governed enterprise data access. Familiarity with AI-assisted development tools such as GitHub Copilot, OpenAI Codex, or equivalent platforms.
Preferred Certifications AWS Certified Data Engineer or another relevant AWS certification. Databricks Certified Data Engineer Associate or Professional.
Scaled Agile Framework certification. Soft Skills Excellent analytical and troubleshooting skills.
Strong written, verbal, presentation, and stakeholder-communication skills. High degree of ownership, initiative, and self-motivation.
Ability to manage multiple priorities in a fast-paced environment. Ability to work effectively with global and virtual teams.
Strong attention to detail and commitment to engineering quality. Ability to influence technical decisions and drive work to completion. .
Originally posted by Amgen. View original posting
Similar roles at other companies
BioCareerAI is an independent job platform and is not affiliated with or endorsed by Amgen.
