A
AgileEngine
FrontendMiddle
Middle Databricks Data Engineer ID86295
DatabricksPySparkDelta LakePythonSQLApache SparkStructured StreamingAuto LoaderKafkaEvent HubsDatabricks WorkflowsApache AirflowAzure Data FactoryUnity CatalogdbtTerraformAzure DevOpsPostgreSQLDynatraceAWS CloudWatchDatabricks system tablesAnthropicGitHub Copilot
About the Position
We are looking for a Middle Data Engineer to help modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will build batch and streaming pipelines with PySpark and Delta Lake, following a medallion architecture. This role also uses AI tools like Claude and GitHub Copilot to speed up development.
Responsibilities
- Design, build, and operate batch and streaming data pipelines on Databricks using PySpark, Delta Lake, and Databricks Workflows.
- Model and maintain a medallion (bronze/silver/gold) architecture serving analytics, reporting, and machine learning consumers.
- Migrate legacy ETL and data warehouse workloads onto the Lakehouse with validated data parity and minimal business disruption.
- Use Claude or Github Copilot as a development accelerator, generating code scaffolding, writing and reviewing tests, creating documentation and prototyping solutions.
- Write clean, well-tested Python and SQL; maintain high standards through code review and documentation.
- Optimize Spark jobs and Delta tables for performance and cost, including partitioning, clustering, caching, and cluster sizing.
- Implement data quality, lineage, and governance controls using Unity Catalog and automated validation checks.
- Debug, troubleshoot, and resolve pipeline failures, data defects, and production incidents.
- Collaborate with DevOps, platform, and analytics engineers on observability, security, and compliance best practices.
Requirements
- 3+ years of professional experience in data engineering, featuring direct expertise with Apache Spark and cloud-based data architectures.
- Strong hands-on experience building data pipelines with Databricks, Apache Spark (PySpark), and Delta Lake.
- Advanced SQL and Python, with strong data modeling skills across dimensional and Lakehouse patterns.
- Experience with streaming ingestion using Structured Streaming, Auto Loader, Kafka, or Event Hubs.
- Experience with workflow orchestration (Databricks Workflows, Airflow, or Azure Data Factory).
- Experience with legacy platform migrations, ETL modernization, or managing data hygiene when porting old systems.
- Strong problem-solving, collaboration, and communication skills.
- Familiarity with Unity Catalog, data governance, access control, and PII handling.
- Experience with dbt or an equivalent transformation framework.
- Familiarity with secure coding standards and industry security best practices.
- Experience delivering production data platforms at scale.
- Upper-intermediate English level.
Benefits
- Accelerate your professional journey with mentorship, TechTalks, and personalized growth roadmaps
- We match your ever-growing skills, talent, and contributions with competitive USD-based compensation and budgets for education, fitness, and team activities
- Join projects with modern solutions development and top-tier clients that include Fortune 500 enterprises and leading product brands
- Tailor your schedule for an optimal work-life balance, by having the options of working from home and going to the office – whatever makes you the happiest and most productive.
Middle Databricks Data Engineer ID86295
View Original