Role Overview
This is a working leadership role: roughly 60% of your time goes to hands-on engineering and 40% to leading people, so you stay close enough to the platform to shape technical direction and spot trouble before it reaches your inbox. Day to day you set the priorities for the team that runs Remote's core infrastructure and reliability practice, coach and grow the engineers on it, and act as the group's voice in conversations across engineering and with senior leadership. The reliability function here is still taking shape — the SLO framework has landed on its first few teams, and questions about how operational load should be weighed against project delivery are genuinely unresolved — which is exactly why the role matters: the groundwork exists, but the hard problems are still open.
Key Responsibilities
- Guide the full career arc of your direct reports, covering onboarding, regular feedback, performance assessment, promotion readiness and hiring as the team grows.
- Set the team's commitments and sequencing, deciding what reliability work gets taken on, in what order, and how it trades off against day-to-day operational demands and unplanned work.
- Design and own the support and on-call model so that incident coverage stays dependable without burning the team out.
- Stay hands-on enough to review infrastructure changes, push back on designs, and contribute credibly during an incident.
- Carry the SLO and error budget framework out from its early adopters to the remaining engineering teams, and keep the incident review loop honest so that findings turn into durable changes rather than one-off fixes.
- Own the health of the core platform: Kubernetes, AWS, PostgreSQL, DNS and TLS, and the CI infrastructure engineers depend on.
- Partner with the Security team on threat response, patching and infrastructure controls, including the team's audit and compliance obligations.
- Manage the vendor relationships underpinning the platform, including renewals and commercial discussions with backing from your Director, and represent the team's interests to the wider engineering organisation and senior leadership.
Requirements & Qualifications
- Prior experience leading a site reliability, infrastructure or platform engineering team, with genuine ownership of your reports' growth, performance and progression rather than just their ticket queue.
- A track record of coaching both technical craft and interpersonal skills, with concrete examples of people who developed because of your guidance.
- Comfort addressing underperformance early and directly, with clarity and empathy, and skill at reading team dynamics and resolving conflict instead of working around it.
- Hiring experience, including the judgement to distinguish a strong interview performance from a strong engineer.
- Hands-on grounding in site reliability, DevOps or cloud infrastructure engineering, deep enough to review your team's work and be taken seriously in an incident.
- Production Kubernetes experience, including the messy operational realities rather than only the well-trodden path.
- AWS experience at meaningful scale.
- Hands-on involvement in building, enabling and scaling AI infrastructure.
- Solid grounding in observability principles and practice.
- Infrastructure as code with Terraform.
- CI/CD systems such as GitLab CI, GitHub Actions or Jenkins, plus Docker and shell scripting.
- Experience running a reliability practice end to end: incident response, on-call, SLOs, error budgets, and the discipline of converting incidents into lasting change.
- Familiarity with working in regulated environments.
- Excellent prioritisation when operational load and project work compete for the same hours, and the ability to shield your team's focus without letting operational commitments slip.
- Clear, structured writing — Remote is fully distributed and asynchronous, so a great deal of your leadership happens in written form.
- Confidence building relationships across teams; much of an SRE team's value comes from being the group others bring problems to early.
Nice to have
- Working knowledge of a backend language, ideally Elixir, or otherwise Java, Clojure, Node.js, Python or similar.
- Depth in modern observability tooling: OpenTelemetry, distributed tracing, and platforms such as Honeycomb.
- Database operations experience, particularly PostgreSQL or Aurora performance, connection pool health and query tuning.
- Running and configuring Linux systems outside a managed cloud environment.
- Security capability spanning both defensive and offensive perspectives.
- Cloud cost management and FinOps experience.
- Experience growing a team from a small base, including establishing the hiring bar along the way.
Why Join
- The role is deliberately split between engineering and leadership, so you keep your technical edge while building and developing a team.
- The platform foundations are already in place; the open questions — extending the SLO framework, rebalancing operational load against project delivery — are the interesting ones.
- You will have a direct hand in shaping how reliability is practised across the whole engineering organisation, not just within one team.
- Remote is a fully distributed, asynchronous workplace, and the position reflects that way of working.
433 open positions on Semasocial right now
· 11212 open positions in Nairobi County, Kenya
· 7 posted in the last 7 days
Contact Information