Data Platform SRE

Outerlimit
Outerlimit

London, UK

Posted on Sep 25, 2026

Outerlimit is looking for a Senior Site Reliability Engineer to embed within our Data Engineering team. You'll own the reliability, performance, and operability of the data platform — the pipelines, orchestration, storage, and streaming systems that the wider business depends on for analytics, reporting, and product features. This is a hands-on engineering role: you'll write infrastructure-as-code, build observability into data systems from the ground up, lead incident response for data platform outages, and work directly with data engineers to raise the reliability bar on everything they ship.

You'll act as the reliability voice inside the Data Engineering team — not a separate, siloed SRE function — balancing feature velocity against system stability, and bringing SRE practice (SLOs, error budgets, blameless post-mortems, capacity planning) to a team that has historically optimised for delivery speed.

Key Responsibilities

  • Define and maintain SLIs/SLOs for critical data pipelines, warehouses, and streaming systems, and use error budgets to guide the pace of change.
  • Design, build, and maintain the infrastructure that underpins the data platform (compute, storage, orchestration, networking) using infrastructure-as-code primarily focussed in Microsoft Azure.
  • Build and improve observability for data systems — metrics, logging, tracing, and data-quality/freshness monitoring — so issues are caught before they reach downstream consumers.
  • Lead incident response for data platform issues: triage, coordinate, drive root cause analysis, and run blameless post-mortems that result in durable fixes.
  • Own on-call rotation design and participate in on-call for the data platform, working to reduce toil and alert noise over time.
  • Partner with data engineers on pipeline design reviews, capacity planning, and cost/performance trade-offs, embedding reliability practices into their day-to-day workflow rather than gatekeeping after the fact.
  • Automate manual operational work — deployments, scaling, failover, data backfills, recovery procedures — to reduce repetitive load on the team.
  • Drive disaster recovery and business continuity planning for data systems, including backup strategy, failover testing, and documented runbooks.
  • Contribute to the broader platform/infrastructure SRE community at Outerlimit, sharing tooling and practices across teams.
  • Support compliance and audit requirements (e.g. SOC 2, data governance) as they relate to data platform reliability, access control, and change management.

Nice to have

  • Experience with a cloud data warehouse (Databricks, BigQuery, or Redshift) at scale.
  • Experience operating streaming systems (Kafka, Azure Event Hubs, or similar) in production.
  • Good understanding of Python with the ability to read/debug Pyspark jobs and configure tooling.
  • Prior experience formally introducing SRE practices to a team that didn't previously have them.
  • Relevant compliance/security exposure (SOC 2, ISO 27001, or similar frameworks).

What we offer

  • Hybrid working — London office, 2+ days per week, with flexibility around core hours.
  • A genuine seat at the table shaping how the Data Engineering team builds and operates its platform, not a bolt-on ops function.
  • Investment in tooling, training, and conference attendance to keep your SRE practice current.
  • The standard Outerlimit benefits package (details shared during the interview process).