Our client is looking for Data Engineers at different experience levels to join their growing company in Malta.
Role Overview:
You build and operate the pipelines and data products that everything else depends on. You make provenance structural: stable record identity, idempotency, defined reprocess behaviour, replayable failures.
You publish against documented, versioned contracts, so downstream teams ship without you writing their code. You move real volumes with Spark on Kubernetes and streaming through Kafka. You carry the SLOs and the pager for what you build.
Responsibilities:
- Build pipelines that carry provenance by construction. Stable record identity. Idempotency keys. Explicit reprocess and overwrite behaviour. Failures that are visible and replayable.
- Model and publish data products against versioned contracts. Agreed record identity and sourcing pattern. Schema registered with owner and project. Documentation good enough that a downstream team ships without you.
- Move data at production volume. Spark and PySpark on Kubernetes. Streaming through Kafka or Redpanda. Durable orchestration. Open table formats with schema evolution, partitioning and compaction.
- Make lineage, quality and audit answerable. A lineage query that returns source, transaction, pipeline run and model version. Quality checks with defined thresholds and defined consequences. Audit coverage of modify, download and share.
- Operate what you build. Published SLOs. Dashboards and alerting. Runbooks that cover the failure modes.
Requirements
- Python for data work with engineering discipline: PySpark, Arrow or pandas
- SQL and PostgreSQL: modeling, indexing, reading query plans, migrations safe against live consumers
- Batch and streaming together: Spark or PySpark, Kafka or Redpanda, and clear reasoning about idempotency, ordering, backpressure, checkpoints and replay
- Open table formats and object storage: Apache Iceberg or Hudi with a catalog, partitioning and clustering, compaction, time travel, and S3-compatible storage semantics
- Durable workflow orchestration: Temporal, Airflow or equivalent
- Data contracts and schema evolution: Avro, Protobuf or JSON Schema, backward and forward compatibility, versioning, and a migration path for consumers
- Data contracts and schema evolution: Avro, Protobuf or JSON Schema, backward and forward compatibility, versioning, and a migration path for consumers
- Must be located in Malta or willing to relocate permanently.
Benefits
- Fully remote
- Home office package and set-up
- Relocation package
Education and experience
- Five to eight years of data engineering working with large datasets
- Ran pipelines at terabyte scale or above, and knows where the naive implementation breaks
- Delivered against a downstream contract with named consumers who depended on it
- Worked in a regulated, multi-tenant or restricted-network environment