Question Clearly sourced

Expert knowledge for digital decisions

What Service Levels and Responsibilities Are Needed for Productive AI Operations?

Short answer

A productive AI service requires measurable SLIs and SLOs for availability, response time, error rates, and professional quality, as well as clear responsibilities for platform, model, data, data protection, and professional approval. An SLA describes the external commitment and consequences; it does not replace operational runbooks or quality checks.

Measure First, Then Commit

Google SRE distinguishes between Service Level Indicator, Objective, and Agreement. An SLI is the metric, such as successful requests or p95 time-to-first-token. The SLO is the internal target. An SLA is an agreement with users or customers and can define consequences for non-compliance. Not every internal metric belongs in a contract.

For an AI service, at least availability, error rate, p95 response time, capacity limits, and support times are defined. Professional quality requires its own SLO based on a versioned test set, for example, a minimum rate of correct, substantiated responses and correct rejections. Since generative responses cannot be guaranteed deterministically, no blanket truth promise is formulated. Instead, operational limits, testing procedures, and human approval for critical decisions are established.

The percentage is translated into understandable time. In a hypothetical 30-day month, 99.5% availability corresponds to a maximum of 3 hours and 36 minutes of unplanned downtime, 99.9% to about 43 minutes and 12 seconds, and 99.95% to about 21 minutes and 36 seconds. Whether maintenance windows count, how partial outages are measured, and at which measurement point availability begins must be explicitly defined.

Responsibilities are assigned by name or role-based: Service Owner, platform operations, model responsibility, data responsibility, information security, data protection, and professional approval. A RACI can make responsibility, decision-making, and consultation visible. Runbooks regulate alerts, escalation, communication, rollback, and recovery. SLOs also manage an error budget: if it is consumed, stabilization and error correction take precedence over risky changes. Quarterly reviews are a possible rhythm; frequency and on-call availability depend on business risk and promised service.

Key facts

99.9% in a 30-day month
about 43 minutes and 12 seconds error budget
99.95% in a 30-day month
about 21 minutes and 36 seconds error budget
Areas of Responsibility
Platform, model, data, security, data protection, and professional approval

Sources

All external claims are backed by traceable sources.
  1. 01
    Service Level Objectives Google Site Reliability Engineering
  2. 02
    Embracing Risk – Error Budgets Google Site Reliability Engineering
  3. 03
    Best Practices for Production Services Google Site Reliability Engineering

Ready for your next project?

Free initial consultation - no sales pressure, just clear answers.

Request consultation