YOUR ROLE

  • Design, develop, and maintain scalable batch and real-time data pipelines using Databricks, PySpark, and Spark SQL.
  • Build and implement Delta Lake-based solutions, including data ingestion, transformation, and Medallion Architecture (Bronze, Silver, Gold) layers.
  • Develop and orchestrate end-to-end data workflows using Databricks Workflows and other scheduling tools.
  • Optimize Spark workloads for performance, scalability, and cost efficiency through effective cluster management and tuning practices.
  • Ensure data quality, reliability, and governance through validation, monitoring, and testing frameworks.Integrate and process data from multiple enterprise sources, including databases, APIs, files, and streaming platforms.
  • Collaborate with cross-functional teams, including data scientists, analysts, architects, and business stakeholders, to deliver data-driven solutions.
  • Contribute to code management, CI/CD implementation, documentation, and continuous improvement initiatives.

YOUR PROFILE

  • 4-6 years of experience in Data Engineering, with strong hands-on expertise in Databricks.
  • Proficiency in Python, PySpark, and SQL for building scalable data processing solutions.
  • Strong understanding of Apache Spark architecture, Delta Lake, and modern data engineering practices.
  • Experience working with cloud platforms such as Azure, AWS, or GCP.Solid knowledge of ETL/ELT processes, data modeling, and data warehousing concepts.
  • Experience with version control systems such as Git and CI/CD processes.
  • Familiarity with technologies such as Unity Catalog, Azure Data Factory, ADLS Gen2, Kafka, Event Hubs, Structured Streaming, or Snowflake is an advantage.
  • Exposure to Infrastructure-as-Code tools such as Terraform is a plus.Databricks certifications will be considered an added advantage.
  • Bachelor's degree in Computer Science, Information Technology, or a related field.