I
Implex
BackendSenior
Senior Data Engineer (GCP, PySpark, Dataproc)
Apache SparkPysparkGcp DataprocBigqueryPostgreSQLSap HanaHadoop S3A ConnectorMinioTls/SslJava TruststoresJdbcOdbcEtlSqlData WarehousingData ModelingScdGcp Secret ManagerKafkaLinuxcURLOpenssl S_Client
Про позицію
We are looking for a Senior Data Engineer with strong hands-on experience in Apache Spark, PySpark and GCP Dataproc to join Implex and work on a long-term data platform project for a public-sector in the Middle East. The role involves developing and supporting data-processing pipelines integrating multiple enterprise data sources.
Обовʼязки
- Develop, maintain, and optimize ETL pipelines using Apache Spark and PySpark.
- Configure, run, and troubleshoot Spark workloads on GCP Dataproc.
- Package and submit Spark jobs, manage dependencies, and analyze driver and executor logs.
- Identify and resolve performance issues related to memory usage, shuffling, partitioning, data skew, and distributed data processing.
- Integrate Spark workloads with S3-compatible object storage using the Hadoop S3A connector.
- Configure and troubleshoot Spark connectivity with MinIO, including bucket policies, custom endpoints, path-style access, TLS, signature compatibility, and redirect handling.
- Configure custom CA certificates and Java truststores for Spark drivers and executors.
- Implement and optimize scalable Spark JDBC reads and writes for PostgreSQL and SAP HANA, including partitioning, batching, pushdown, and data type mapping.
- Build reliable incremental data loads, retries, backfills, idempotent processing, and data reconciliation mechanisms.
- Design and support staging-to-publish data flows, bulk data loads, and data warehouse integration patterns.
- Contribute to data modeling, schema evolution, Slowly Changing Dimension patterns, data lineage, technical documentation, and data quality practices.
- Apply secure secrets management and least-privilege access principles across data integrations.
Вимоги
- Apache Spark / PySpark development (Dataproc): driver/executor behavior, job packaging/submission, performance tuning
- GCP Dataproc operations: cluster configuration, init actions, dependency management, troubleshooting via logs/metrics
- Hadoop S3A connector: `fs.s3a.*` configuration, endpoint/path-style access, credential providers, S3 semantics
- MinIO (S3-compatible) integration: bucket policies, TLS endpoints, signature/redirect troubleshooting
- TLS/SSL & PKI with custom CA: certificate chains, SAN/hostname validation, diagnosing handshake/PKIX errors
- Java truststores (JKS/PKCS12) & JVM SSL config: `keytool`, distributing truststores, setting driver/executor JVM options
- PostgreSQL integration: Spark JDBC reads/writes at scale, indexing/performance basics, data type mapping
- SAP HANA integration: JDBC/ODBC connectivity, driver management, calculation views vs tables, pushdown/performance tuning
- ETL engineering: incremental loads/CDC concepts, idempotency, retries, backfills, data quality/reconciliation
- Data Warehousing integration: strong SQL, staging-to-publish patterns, SCD concepts, bulk load strategies
- Data modeling & governance basics: dimensional modeling, schema evolution, lineage/documentation practices
- Linux + networking fundamentals: DNS, routing, firewall/LB/proxy basics; tools like `curl`/`openssl s_client` for validation
- Secure secrets handling: GCP Secret Manager (or equivalent), least-privilege access, avoiding hardcoded credentials
- KAFKA knowledge if we ever bring KAFKA into the architecture again
Переваги
- A long-term international project
- Opportunity to work on a national-scale digital platform used by thousands of users
- Remote full-time collaboration
- Professional and supportive team environment
- Challenging technical tasks and a meaningful product with real-world impact
Senior Data Engineer (GCP, PySpark, Dataproc)
Оригінал