Role Overview
This role carries end-to-end ownership of Food4Education's data platform — the connectors, orchestrators and warehouse layers that move records out of source systems, and the models, tests and definitions that determine what those records actually mean. Day to day you will profile awkward sources, write and tune Python and SQL, build dbt models and tests, chase down why two systems report different totals for the same metric, and tighten the controls that keep personal data out of reporting. It matters because expansion into new countries, reporting to external bodies and technical support to partner organisations all depend on data that lands intact, is modelled identically wherever it is consumed, and can be run by someone other than the person who built it.
Key Responsibilities
- Source discovery and platform design: Profile and document every feed the organisation relies on — relational databases, spreadsheets, APIs and unstructured data — recording quality issues, data volumes, refresh frequency, key relationships and how each source maps to business entities. Shape and maintain the shared BigQuery model using agreed naming and governance conventions, with partitioning, clustering, retention and historisation choices made deliberately for cost and for portability across countries and partners.
- Pipeline development: Build and maintain extraction, transformation and enrichment pipelines for each source system, favouring incremental loading to keep processing costs down, and deliver migrations, new integrations and first-time ingestion for additional countries.
- Orchestration and reliability: Configure and maintain Apache Airflow or an equivalent scheduler so runs follow business rhythms, with dependency management, retries, backfills and error handling in place. Watch pipeline health continuously and clear failures before they surface downstream.
- Analytics engineering: Design and maintain dbt models from staging through to serving layers, holding grain, naming and conformed dimensions consistent throughout. Translate agreed metric definitions into working logic so dashboards, reports and external submissions all return the same answer, and build dbt tests across the assets that matter most.
- Metric governance: Own the KPI dictionary — definition, logic, source, refresh cadence, as-at date and business owner. Reconcile conflicting figures between systems, publish certified self-service datasets, and escalate definitions that cannot be implemented as written, working with system and business owners to close the gap upstream.
- Data quality, protection and incidents: Place validation checks at extraction and load, maintain quality monitoring dashboards, data dictionaries and lineage, and log, triage, resolve and document incidents against agreed severity levels. Mask or hash personal data inside the ETL layer, enforce row- and column-level access tiers, make sure no model reintroduces personal data downstream, and support access reviews and data protection obligations.
- Documentation, knowledge transfer and enablement: Keep the raw-layer dictionary, KPI dictionary, dependency maps, lineage records, runbooks and transformation documentation current. Train and support BI colleagues on the data models and how to access them, and leave the platform operable and transferable without you present.
Requirements & Qualifications
- Bachelor's degree in Computer Science, Engineering, Statistics or a comparable field.
- At least four years spanning data engineering and analytics engineering, with hands-on production experience in both disciplines.
- A minimum of two years owning dbt and BigQuery in a production setting — models, tests and documentation included — plus a track record of delivering data migrations.
- Strong Python and SQL, and advanced working knowledge of BigQuery or a comparable cloud data warehouse in production.
- Solid dimensional modelling skills: grain, conformed dimensions, slowly changing dimensions and star schema design.
- Production experience building and orchestrating pipelines with tools such as Apache Airflow, Dagster, Cloud Composer or GitHub Actions, including scheduling, dependency handling, retries and alerting.
- Familiarity with change data capture, incremental loading and backfill approaches across a range of source system types.
- Applied use of version control, code review and CI/CD for data work on GitHub or an equivalent platform.
- Practical grounding in data protection controls and data migration procedures.
- The ability to turn a business metric definition into an implementable specification and discuss it directly with non-technical stakeholders.
- Personal qualities that fit a small team: accuracy-focused, documents as a matter of habit, builds systems others can operate, raises problems early, is comfortable disclosing known issues and quality gaps, communicates clearly with non-technical colleagues, system owners and vendors, and stays organised under operational pressure.
- Certifications: a cloud data platform or dbt certification is helpful but not expected.
- Advantageous but not essential: ERP integration experience such as Sage X3 or similar; IoT or telematics data work; supporting month-end financial close so finance feeds are complete, reconciled and ready on time; and prior experience in a small team holding up the full data stack.
What Success Looks Like in the First Year
- Data migrations and integrations delivered inside the agreed timelines, with new ingestions following standard patterns rather than bespoke builds.
- Pipeline availability held above 90%, and alerts answered within four hours.
- Failures caught by monitoring before a stakeholder notices them.
- Data and KPI dictionaries covering more than 80% of core datasets.
- Personal data masked at the point of ingestion, with access tiers applied across all reporting.
- Dependencies and governance grounded in business logic implemented on the core datasets.
- Every incident closed with a documented root cause within five working days.
845 open positions on Semasocial right now
· 11386 open positions in Nairobi County, Kenya
· 24 posted in the last 7 days
Contact Information