Skip to content
Service

Managed Data Services

Ongoing ownership of a platform after it is built: monitoring, incident response, cost control and continued development against an agreed monthly capacity.

The problem

The platform works. Our Snowflake bill has tripled in a year, nobody owns it, and the engineer who built it has left.

A platform without a named owner degrades in a predictable way. Warehouses stay oversized, because nobody wants to be the person who slowed down a report. Intermittent test failures get muted. Documentation stops being updated at the first schema change that goes unrecorded, and none of it is visible from outside the team until something breaks during a close or an audit.

Our approach

We take named ownership of the platform against an agreed response standard, and we work a visible backlog alongside incident response. Cost is tracked as an engineering metric against an agreed target.

  1. 01

    Takeover

    We document the current state, find the undocumented dependencies, and fix whatever is already failing without an alert being raised. A takeover normally finds at least one job in that condition.

  2. 02

    Operate

    Monitoring and alerting run against a defined response standard, and incidents are handled with written post-incident reviews. The same named engineer works on your platform each month. Requests do not go into a rotating queue.

  3. 03

    Optimize

    We review warehouse sizing and auto-suspend policy, the query patterns that account for most of the spend, storage and retention settings, and whether clustering is reducing enough scanned data to justify its cost. Each change and its measured effect are reported monthly against a spend target.

  4. 04

    Extend

    A prioritized backlog is worked at an agreed monthly capacity, so a new request can be scheduled without an escalation.

What you get

  • Platform documentation brought current, including the dependencies discovered during takeover
  • Monitoring and alert coverage mapped against a defined response standard
  • Monthly operations report covering incidents, spend against target and backlog progress
  • Cost optimization log recording each change and its measured effect
  • Written post-incident reviews
  • Quarterly platform review with your stakeholders

Technology

Platforms

  • Snowflake
  • Databricks
  • Microsoft Fabric
  • BigQuery

Operations

  • Airflow
  • dbt Cloud
  • Monte Carlo
  • Datadog
  • PagerDuty

Cost management

  • Snowflake resource monitors
  • query attribution
  • warehouse right sizing
  • Select.dev

Engagement

Duration
Rolling, on a twelve-month term with an agreed monthly capacity
Team
A named lead engineer with a second engineer for cover, escalating to an architect
Starts with
We start with a platform health check that covers cost, reliability and documentation. The findings become the baseline that the ongoing arrangement is measured against.

The first stage is short, and you can stop after it if the findings do not justify going further.

All services