Data Engineer - Lead

Apply now »

Posted On: 27 Jul 2026

Location: Noida, UP, India

Company: Iris Software

Why Join Iris?
Are you ready to do the best work of your career at one of India’s Top 25 Best Workplaces in IT industry? Do you want to grow in an award-winning culture that truly values your talent and ambitions?
Join Iris Software — one of the fastest-growing IT services companies — where you own and shape your success story.
 
About Us  
At Iris Software, our vision is to be our client’s most trusted technology partner, and the first choice for the industry’s top professionals to realize their full potential.
With over 4,300 associates across India, U.S.A, and Canada, we help our enterprise clients thrive with technology-enabled transformation across financial services, healthcare, transportation & logistics, and professional services.
Our work covers complex, mission-critical applications with the latest technologies, such as high-value complex Application & Product Engineering, Data & Analytics, Cloud, DevOps, Data & MLOps, Quality Engineering, and Business Automation.

Working with Us
At Iris, every role is more than a job — it’s a launchpad for growth.
Our Employee Value Proposition, “Build Your Future. Own Your Journey.” reflects our belief that people thrive when they have ownership of their career and the right opportunities to shape it.
We foster a culture where your potential is valued, your voice matters, and your work creates real impact. With cutting-edge projects, personalized career development, continuous learning and mentorship, we support you to grow and become your best — both personally and professionally.
Curious what it’s like to work at Iris? Head to this video for an inside look at the people, the passion, and the possibilities. Watch it here.

Job Description

Mandatory Skills:

Databricks Workflows, PySpark, Amazon Kinesis, Delta Lake on Databricks

Additional Skills:

CI/CD & Source Control (CI/CD + Jenkins + Git)

Key Responsibilities:

  • Define and drive enterprise data engineering strategy aligned with organizational objectives and data modernization initiatives.
  • Establish data engineering standards, governance frameworks, and best practices across teams.
  • Lead the design of enterprise-scale data processing architectures using PySpark and modern data platform technologies.
  • Define enterprise standards for Snowflake and Delta Lake-based data platforms supporting analytical and operational workloads.
  • Drive real-time and event-driven data architecture initiatives using Apache Kafka or Amazon Kinesis.
  • Establish governance standards for data ingestion, transformation, streaming, and processing frameworks.
  • Define workflow orchestration, scheduling, and operational governance standards using Apache Airflow or Databricks Workflows.
  • Establish data quality, validation, monitoring, and operational excellence frameworks across data engineering ecosystems.
  • Define enterprise standards for data products, data quality ownership, metadata management, discoverability, and trusted business data consumption across the organization.
  • Establish architecture standards for modern Lakehouse platforms, data observability, platform engineering, and scalable cloud-native data ecosystems supporting enterprise analytics and AI initiatives.
  • Define AI-ready data foundation strategies supporting structured and unstructured data processing, vector-enabled architectures, retrieval patterns, and future GenAI and Agentic AI initiatives.
  • Partner with business stakeholders to translate business objectives into scalable data platform capabilities, data products, and enterprise data architecture decisions.
  • Drive adoption of AI-assisted engineering practices across data engineering teams to improve developer productivity, code quality, documentation, testing, and delivery effectiveness while maintaining governance standards.
  • Lead architecture reviews and ensure data solutions meet scalability, reliability, maintainability, and performance objectives.
  • Guide teams on distributed data processing, streaming architectures, modern data platforms, and engineering best practices.
  • Identify platform risks, scalability bottlenecks, operational gaps, and architectural challenges while defining mitigation strategies.
  • Collaborate with various teams and leadership stakeholders to align data initiatives with organizational objectives.
  • Drive continuous improvement initiatives focused on platform maturity, engineering excellence, scalability, reliability, and delivery effectiveness.

 

  •  
  •  
 

 

Mandatory Competencies

Data Science and Machine Learning - Data Science and Machine Learning - Apache Spark
Data & AI - Data Engineering - Data Quality & Validation
Big Data - Big Data - Pyspark
Data Science and Machine Learning - Data Science and Machine Learning - Python
Database - Database Programming - SQL
Data Science and Machine Learning - Data Science and Machine Learning - Databricks
Beh - Communication and collaboration

Perks and Benefits for Irisians
Iris provides world-class benefits for a personalized employee experience. These benefits are designed to support financial, health and well-being needs of Irisians for a holistic professional and personal growth. Click here to view the benefits.

Apply now »