Skip to main content
AgileEngine
Зібрано зDjinniСьогодні
FrontendMiddle

Middle Databricks Data Engineer ID86295

DatabricksPySparkDelta LakePythonSQLApache SparkStructured StreamingAuto LoaderKafkaEvent HubsDatabricks WorkflowsApache AirflowAzure Data FactoryUnity CatalogdbtTerraformAzure DevOpsPostgreSQLDynatraceAWS CloudWatchDatabricks system tablesAnthropicGitHub Copilot
Формат
Remote
Зайнятість
-
Локація
Worldwide
Оплата
Не вказана

Про позицію

We are looking for a Middle Data Engineer to help modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will build batch and streaming pipelines with PySpark and Delta Lake, following a medallion architecture. This role also uses AI tools like Claude and GitHub Copilot to speed up development.

Обовʼязки

  • Design, build, and operate batch and streaming data pipelines on Databricks using PySpark, Delta Lake, and Databricks Workflows.
  • Model and maintain a medallion (bronze/silver/gold) architecture serving analytics, reporting, and machine learning consumers.
  • Migrate legacy ETL and data warehouse workloads onto the Lakehouse with validated data parity and minimal business disruption.
  • Use Claude or Github Copilot as a development accelerator, generating code scaffolding, writing and reviewing tests, creating documentation and prototyping solutions.
  • Write clean, well-tested Python and SQL; maintain high standards through code review and documentation.
  • Optimize Spark jobs and Delta tables for performance and cost, including partitioning, clustering, caching, and cluster sizing.
  • Implement data quality, lineage, and governance controls using Unity Catalog and automated validation checks.
  • Debug, troubleshoot, and resolve pipeline failures, data defects, and production incidents.
  • Collaborate with DevOps, platform, and analytics engineers on observability, security, and compliance best practices.

Вимоги

  • 3+ years of professional experience in data engineering, featuring direct expertise with Apache Spark and cloud-based data architectures.
  • Strong hands-on experience building data pipelines with Databricks, Apache Spark (PySpark), and Delta Lake.
  • Advanced SQL and Python, with strong data modeling skills across dimensional and Lakehouse patterns.
  • Experience with streaming ingestion using Structured Streaming, Auto Loader, Kafka, or Event Hubs.
  • Experience with workflow orchestration (Databricks Workflows, Airflow, or Azure Data Factory).
  • Experience with legacy platform migrations, ETL modernization, or managing data hygiene when porting old systems.
  • Strong problem-solving, collaboration, and communication skills.
  • Familiarity with Unity Catalog, data governance, access control, and PII handling.
  • Experience with dbt or an equivalent transformation framework.
  • Familiarity with secure coding standards and industry security best practices.
  • Experience delivering production data platforms at scale.
  • Upper-intermediate English level.

Переваги

  • Accelerate your professional journey with mentorship, TechTalks, and personalized growth roadmaps
  • We match your ever-growing skills, talent, and contributions with competitive USD-based compensation and budgets for education, fitness, and team activities
  • Join projects with modern solutions development and top-tier clients that include Fortune 500 enterprises and leading product brands
  • Tailor your schedule for an optimal work-life balance, by having the options of working from home and going to the office – whatever makes you the happiest and most productive.
Middle Databricks Data Engineer ID86295
Оригінал