Job Description
We are looking for a Lead Streaming Engineer to build the ingestion path for the analytics platform of a national-scale foundational identity programme in West Africa. You will own how data gets out of the identity system and into the lake — change data capture, the streaming buffer, stream processing, and the batch and log paths alongside them. This is the component that determines whether the entire platform can be trusted, and it is the one that touches the production identity environment.
The constraints are the interesting part. The source system cannot absorb query load, so capture is log-based. Delivery is at-least-once, so the pipeline has to deduplicate correctly or every recovery event inflates the national enrolment figures. Events arrive out of order and the source schema changes underneath you. Personal identifiers must be pseudonymised inside the pipeline, before anything is written. And there is a historical migration of over one hundred million records to land alongside the live stream.
If you have spent time explaining to people why their pipeline is silently losing or duplicating data, this role will be familiar.
Access Remote and Full time Vacancies on our Channel as they break!
Requirements
● At least seven years in data engineering, of which three or more building production streaming or change data capture pipelines.
● Production experience of log-based change data capture — Debezium or equivalent — including schema change handling and connector operation.
● Apache Kafka in production: topic and partition design, retention, consumer groups, and the practical consequences of at-least-once delivery.
● Stream processing at scale, in Apache Spark Structured Streaming, Flink or equivalent.
● Has diagnosed and resolved a real duplication or ordering defect in production, and can describe both the diagnosis and the design change.
● Workflow orchestration — Airflow or equivalent.
Desirable
● Writing to an open table format, including compaction and file sizing.
● A large historical migration, in the order of one hundred million records or more.
● MOSIP schema familiarity.
● Ingestion from systems holding personal or sensitive data under a data protection regime.
● Java and Python.
● Kubernetes.
Send CV to: [email protected]
Benefits
___ SPONSORED___
Earn gifts and coins from your post on Payhankey. You don't have to pay a penny to receive gifts from your friends. Just create posts and share with your friends to gift you.
On Payhankey, you don't need to go LIVE or have followers to earn from your posts.
This is the social media we need now. A reward for your DATA.
Bye bye to wasting time on other social media that doesn't pay you
Try out PAYHANKEY TODAY
Facebook is not paying you enough for social media posts?
Join 40K+ others earning monthly on Payhankey
Get Started →Download Professional CV for your Application on our Channel
Get your professional CV tailored for job applications
Get Started →