Lead Data Engineer (GenAI / LLM Applications)

By Admin · Data Science, Analytics & AI

BANGALORE (IN)HYBRIDPOSTED: 9/8/2026
TECHFull-time
HIRING ORGANIZATIONCLARIOClinical Research and Healthcare Technology
Clario

ROLE DEFINITION & RESPONSIBILITIES

Core Responsibilities

Design, build, and maintain scalable software architectures, data platforms, and data pipelines to support clinical trial operations and business intelligence.
Write clean, reusable Python code utilizing frameworks such as Flask and libraries like pandas, NumPy, Plotly, and ag-Grid.
Build and integrate GenAI/LLM-powered solutions, including Retrieval-Augmented Generation (RAG) pipelines, intelligent agents, and automated workflows using AWS Bedrock, LangChain, and GitHub Copilot.
Develop and optimize complex SQL queries, procedures, views, and analytical window functions across Oracle, MS SQL Server, PostgreSQL, and Snowflake.
Design, execute, and monitor ETL pipelines using Snowflake and orchestrate workflows with Apache Airflow or similar scheduling frameworks.
Establish data quality frameworks, data modeling standards, versioning, and governance practices across structured and unstructured data sources.
Deploy, manage, and troubleshoot end-to-end cloud solutions on AWS infrastructure (S3, EC2, Lambda, Secrets Manager, Bedrock) within Git-based CI/CD environments.
Partner with product managers, data scientists, and analysts to produce data-driven insights and interactive dashboards using Plotly and Power BI.

Technical & Domain Standards

Demonstrate expertise in GenAI data engineering workflows, RAG architectures, LLM integration (AWS Bedrock, LangChain), Python, advanced SQL, and Snowflake ETL pipelines.
Adhere to healthcare and clinical trial compliance, data governance, source-to-target mapping, and cloud deployment standards.

Leadership & Business Impact

Lead data engineering deliverables across the full Software Development Lifecycle (SDLC), from requirements gathering to production support.
Drive continuous improvement in platform reliability, code performance, automated testing, and developer productivity using AI-assisted tools.

Benefits & Total Rewards

Strategic data leadership position at the intersection of clinical research, global trial data, and generative AI innovation.
Competitive salary aligned with local market standards, comprehensive health benefits, and remote/hybrid flexibility in India.

REQUIRED VERIFIED SKILLS

Python
Flask
SQL
Snowflake
AWS
LangChain
Generative AI
RAG Pipelines
Apache Airflow
PostgreSQL
Oracle
GitHub Copilot

COMPENSATION & REWARDS

Salary: Competitive / Disclosed upon candidate shortlisting
Salaries for Lead Data Engineers with 5+ years of experience specializing in GenAI and cloud data infrastructure in Bangalore/Remote India offer highly competitive market packages, complemented by corporate performance bonuses and comprehensive healthcare coverage.

ORGANIZATION CONTEXT & CULTURE

Clario is actively hiring for this position.

INDUSTRY: Clinical Research and Healthcare Technology

EXPECTED CAREER ROADMAP

This Lead Data Engineer position provides a high-impact progression pathway toward Principal AI/Data Architect, Director of Data Science & Engineering, or Global Head of Data Operations. Professionals gain hands-on experience deploying modern LLM/RAG solutions in the clinical technology sector, opening senior technical leadership opportunities globally across biopharmaceutical and data science industries.

APPLICATION ADVISORY

Highlight hands-on experience building RAG architectures, working with LLM frameworks (LangChain, AWS Bedrock), advanced SQL performance tuning, and Snowflake ETL pipelines. Be prepared to present concrete examples of end-to-end Python microservice deployments and workflow orchestration during technical interviews.

READY TO SUBMIT?

Verify your capability profile and matching credentials before submitting direct applications.