Job Details
Job Type:
Full Time
Workplace Type:
On-site
Qualification:
Diploma
Job Experience:
Mandatory
Job Location:
Nairobi County, Kenya
Closing Date:
Undisclosed
Salary:
Estimated: KES 45,000 - KES 300,000 / month
Other Pay:
Benefits
Role Overview
This position sits at the intersection of data plumbing and cloud infrastructure: you will take messy, inconsistently formatted source files and turn them into reliable datasets that other systems and teams can build on. Day to day you will be writing Python, shaping schemas, shipping containerised services to Google Cloud, and gradually extending the platform with AI-driven features such as semantic search and natural-language access to data. The work matters because everything downstream — analytics, product features, and the emerging agent-based tooling — depends on the pipelines you keep accurate and available.
Key Responsibilities
- Design, build and maintain ingestion pipelines that parse raw CSV and TSV files and land them as clean, well-structured, queryable datasets.
- Define custom entity models and translate heterogeneous source data into consistent target schemas.
- Stand up and operate data environments across SQLite, Google Cloud SQL and BigQuery, choosing the right store for each workload.
- Containerise data loaders, embedding generators and data-serving applications, and keep those images current and reproducible.
- Run production services on Google Cloud Run with attention to scaling behaviour, uptime and failure recovery.
- Wire together supporting infrastructure including Redis, VPC networking, Artifact Registry and Cloud Storage.
- Harden cloud deployments using Cloud IAM roles, VPC configuration and Cloud Run ingress rules.
- Automate recurring workflows and releases through GitHub Actions — scheduled ingestion, identifier mapping, embedding jobs and container builds included.
- Add and integrate AI capabilities such as embeddings, semantic retrieval, natural-language querying and LLM-backed workflows, and contribute to early work on agentic integrations and the Model Context Protocol (MCP).
Requirements & Qualifications
- At least two years working in data engineering, cloud engineering or backend development.
- Confident Python developer, with hands-on Pandas experience for manipulating and reshaping data.
- Practical Docker skills and a track record of running containerised applications.
- Solid grounding in Google Cloud Platform, covering compute, managed databases, data warehousing, networking and IAM.
- Strong SQL, with experience across both relational databases and analytical warehouses.
- Experience designing, building and consuming REST APIs.
- A self-directed engineering approach: able to troubleshoot thorny problems alone and own a system from first commit through to production operation.
- Nice to have: AI/ML engineering exposure such as embeddings, Sentence Transformers or LLM applications; familiarity with agentic frameworks, MCP or similar emerging integration patterns; Flask and Jinja; WSL and Google Cloud Shell.
772 open positions on Semasocial right now
· 10705 open positions in Nairobi County, Kenya
· 34 posted in the last 7 days
Contact Information