Skip to main content
Implex
Scraped fromDjinniYesterday
BackendSenior

Senior Data Engineer (GCP, PySpark, Dataproc)

Apache SparkPysparkGcp DataprocBigqueryPostgreSQLSap HanaHadoop S3A ConnectorMinioTls/SslJava TruststoresJdbcOdbcEtlSqlData WarehousingData ModelingScdGcp Secret ManagerKafkaLinuxcURLOpenssl S_Client
Work Type
Remote
Job Type
-
Location
Worldwide
Salary
Not specified

About the Position

We are looking for a Senior Data Engineer with strong hands-on experience in Apache Spark, PySpark and GCP Dataproc to join Implex and work on a long-term data platform project for a public-sector in the Middle East. The role involves developing and supporting data-processing pipelines integrating multiple enterprise data sources.

Responsibilities

  • Develop, maintain, and optimize ETL pipelines using Apache Spark and PySpark.
  • Configure, run, and troubleshoot Spark workloads on GCP Dataproc.
  • Package and submit Spark jobs, manage dependencies, and analyze driver and executor logs.
  • Identify and resolve performance issues related to memory usage, shuffling, partitioning, data skew, and distributed data processing.
  • Integrate Spark workloads with S3-compatible object storage using the Hadoop S3A connector.
  • Configure and troubleshoot Spark connectivity with MinIO, including bucket policies, custom endpoints, path-style access, TLS, signature compatibility, and redirect handling.
  • Configure custom CA certificates and Java truststores for Spark drivers and executors.
  • Implement and optimize scalable Spark JDBC reads and writes for PostgreSQL and SAP HANA, including partitioning, batching, pushdown, and data type mapping.
  • Build reliable incremental data loads, retries, backfills, idempotent processing, and data reconciliation mechanisms.
  • Design and support staging-to-publish data flows, bulk data loads, and data warehouse integration patterns.
  • Contribute to data modeling, schema evolution, Slowly Changing Dimension patterns, data lineage, technical documentation, and data quality practices.
  • Apply secure secrets management and least-privilege access principles across data integrations.

Requirements

  • Apache Spark / PySpark development (Dataproc): driver/executor behavior, job packaging/submission, performance tuning
  • GCP Dataproc operations: cluster configuration, init actions, dependency management, troubleshooting via logs/metrics
  • Hadoop S3A connector: `fs.s3a.*` configuration, endpoint/path-style access, credential providers, S3 semantics
  • MinIO (S3-compatible) integration: bucket policies, TLS endpoints, signature/redirect troubleshooting
  • TLS/SSL & PKI with custom CA: certificate chains, SAN/hostname validation, diagnosing handshake/PKIX errors
  • Java truststores (JKS/PKCS12) & JVM SSL config: `keytool`, distributing truststores, setting driver/executor JVM options
  • PostgreSQL integration: Spark JDBC reads/writes at scale, indexing/performance basics, data type mapping
  • SAP HANA integration: JDBC/ODBC connectivity, driver management, calculation views vs tables, pushdown/performance tuning
  • ETL engineering: incremental loads/CDC concepts, idempotency, retries, backfills, data quality/reconciliation
  • Data Warehousing integration: strong SQL, staging-to-publish patterns, SCD concepts, bulk load strategies
  • Data modeling & governance basics: dimensional modeling, schema evolution, lineage/documentation practices
  • Linux + networking fundamentals: DNS, routing, firewall/LB/proxy basics; tools like `curl`/`openssl s_client` for validation
  • Secure secrets handling: GCP Secret Manager (or equivalent), least-privilege access, avoiding hardcoded credentials
  • KAFKA knowledge if we ever bring KAFKA into the architecture again

Benefits

  • A long-term international project
  • Opportunity to work on a national-scale digital platform used by thousands of users
  • Remote full-time collaboration
  • Professional and supportive team environment
  • Challenging technical tasks and a meaningful product with real-world impact
Senior Data Engineer (GCP, PySpark, Dataproc)
View Original