Picture an IT provider that meets its 15-minute response SLA. A ticket comes in, an engineer acknowledges it within six minutes, and the target is met. Meanwhile, Microsoft 365 stays down for the whole company for four hours. The response target was met, and the business still lost its morning.
That gap is where most SLA frustration starts. Response, restoration, and resolution are different things, and an agreement that promises only the first can look impressive while protecting very little. The same goes for “99.9% uptime,” which sounds precise until you ask what is measured, where, and what gets excluded.
This guide explains how to read an IT SLA before you sign it, how to negotiate one that reflects real business impact, and how to use it afterward to manage vendor performance. A well-built SLA is not a trap set by the provider. It protects both sides by making expectations, priorities, and responsibilities measurable.
What an IT Service Level Agreement Actually Does
An IT SLA is the measurable performance commitment within a service relationship: what service is covered, how well it must perform, how performance is measured, and what happens when the target is missed. AWS defines it similarly, citing uptime, response time, and resolution time.
An SLA does not always exist as a standalone contract. It may sit inside the main agreement, arrive as a schedule, or be incorporated by reference, so you may need to read several connected documents to learn what the vendor has actually promised.
| Document | What it does | What to check |
| Contract / MSA (master services agreement) | Overall legal and commercial framework: fees, liability, term, termination | Whether its terms limit or override SLA remedies |
| SOW (statement of work) | A specific project: scope, deliverables, milestones, price | Best for migrations and projects, not daily support |
| Service description | The recurring service: systems, users, locations, dependencies | Whether an issue is in scope before any clock starts |
| SLA | How well the service must perform: targets, measurement, reporting, remedies | Whether definitions are precise enough to measure |
Consider how they interact. The service description says Microsoft 365 support is included. The SLA promises a particular P1 (highest priority) response. The main contract contains a broad third-party exclusion that quietly covers Microsoft outages. Read together, the promise may be much weaker than the headline. An order-of-precedence clause, stating which document wins in a conflict, plus clear version control, helps prevent this.
Three related terms also come up. An SLI (service level indicator) is what is actually measured, such as the percentage of successful logins. An SLO (service level objective) is the target for it. The SLA is the customer-facing commitment, with agreed consequences. Atlassian’s overview covers the distinction.
Outsourcing moves day-to-day control to another organization while the business impact stays with you, so the SLA has to work as a shared operating model. Our guide to getting the most from your IT support provider lists scope, response and resolution targets, escalation, channels, metrics, reporting, and review as its core elements.
What a Strong IT SLA Should Contain
An SLA should not be a table of response times with a signature block. The strongest ones connect scope, operating rules, measurement, and governance, so each part supports the others.
Start with scope and exclusions. Which users, sites, devices, applications, and environments are covered, and which are not? If staff cannot tell whether a real problem is covered without calling their account manager, the scope is too vague.
Next comes the operating model: support hours, time zone, holiday calendar, and contact channels. A clock tied to provider business hours in another time zone does little for an office that works late. The SLA should also say how a P1 is reported after hours, because a portal ticket may not reach anyone at 2 a.m.
The core of the document defines performance: incident and priority definitions, response and restoration commitments, clock rules, escalation, and availability. Around them sit the elements that make an SLA verifiable: how measurement works, what is reported and when, what the customer must provide, what remedies apply, and how the agreement is reviewed and amended.
Not every provision applies to every service:
- Managed support: user and device coverage, after-hours rules, ticket reopen rules.
- Cloud and infrastructure: availability measured where users connect, architecture assumptions, maintenance terms.
- Cybersecurity: a clear line between a security event and a confirmed incident, plus triage, containment, and notification duties.
- Backup and disaster recovery: protected systems, retention, restore testing, recovery objectives.
- Onsite support: travel area, dispatch versus arrival commitments, after-hours fees.
Security, backup, and change-management terms often sit outside the SLA, in a security schedule, data-processing addendum, or recovery plan. Check that these documents connect, and that no responsibility falls into the gaps.
Response Time Is Not Resolution Time
Return to the email outage. The provider acknowledged the ticket in six minutes, then spent four hours working through a Microsoft 365 dependency. The response SLA was met. The business still absorbed a major disruption.
A fast acknowledgement has real value because it confirms ownership and starts communication, but it is not restoration. An SLA should treat them as separate stages.
| Stage | What it means | Define it as |
| Acknowledgement | A qualified person has taken ownership or started triage | Valid ticket or alert to first human response. Automated receipts should not count unless agreed |
| Active engagement | The right resolver is working the problem | From acknowledgement to engagement by the correct team |
| Restoration | The business function works again, perhaps via failover or a workaround | From incident report to verified restoration of the agreed function |
| Resolution | The incident record’s success criteria are met | Must state whether a workaround counts and whether customer confirmation is needed |
| Permanent fix | The root cause is corrected | Usually tracked through problem management, not the incident clock |
Resolution needs particular care because tools disagree about it. ServiceNow’s documentation treats an incident as resolved once the user has a temporary workaround or a permanent solution, while Atlassian notes that temporary restoration may leave the underlying cause unfixed. The contract just has to say which meaning applies.
This is why a first-response target alone leaves a gap. For serious incidents, also ask for active engagement expectations, a fixed P1 update frequency, escalation timing before a breach, a restoration target, a clear definition of resolution, and a root-cause review afterward. A workaround may put people back to work while the cause remains, so track the permanent correction separately with an owner and due date.
How SLA Clocks Work
Every target runs on a clock, and the clock rules quietly decide how it performs. The agreement should say when the clock starts (when a valid ticket arrives, or only after the provider accepts it), what pauses it, what resumes it, what stops it, and whether time is elapsed or counted only in business hours. A P1 described as 24/7 but measured in business hours is not 24/7.
Pausing while genuinely waiting for customer information, approval, or access can be reasonable. The problem is pause conditions so broad that difficult tickets effectively disappear from reporting. Good rules require a documented, timestamped request, automatic resumption when the customer responds, and continued work on anything that can progress in parallel. Check reopened tickets too: if reopening starts a new clock, one failed fix becomes two successful-looking records.
Define Incident Priorities Around Business Impact
P1 to P4 labels are not universal standards, and no response time belongs to them by default. Atlassian describes priority as the result of impact (the effect on business processes) and urgency (how soon that impact becomes significant). A severe problem that will not hurt until quarter end is less urgent than a smaller one blocking work now.
A useful model weighs the critical systems affected, number of users, revenue or operational impact, security implications, degree of degradation, and whether a workaround exists.
User count alone should not determine priority. One treasury user locked out of a payment system at a deadline may create more business impact than 30 users with a minor problem in a non-critical application. A security event may also deserve high priority before anyone sees its effects.
Illustrative framework only. The table below is an example for explanation, not FunctionEight’s SLA and not an industry standard. It deliberately carries no time targets.
| Level | Illustrative criteria | Example incidents |
| P1 / Critical | Critical service down or serious active security impact; no workable alternative | Company-wide sign-in failure; ERP unavailable at financial close; ransomware spreading |
| P2 / High | Major degradation, or a department or site blocked; impact substantial but contained | One office offline; failed production integration; email down for a large team |
| P3 / Medium | One or a few users affected; workaround exists | New hire cannot sign in; one user’s application keeps crashing |
| P4 / Low | Minor fault, how-to question, or scheduled request | Software install request; cosmetic issue |
Preventing Priority Disputes
Priority matters because it sets the target. If the provider alone can downgrade a P1 to a P2, the clock changes and so does the apparent compliance. The SLA should state who assigns the initial priority, who can change it, what evidence supports reclassification, whether the customer can challenge a downgrade, and whether the original clock keeps running during a dispute. A jointly agreed impact and urgency matrix, with named critical services and a record of priority changes, prevents most arguments.
What Uptime Percentages Really Mean
Allowed downtime is the measurement period multiplied by one minus the availability target. The table applies that to a 30-day month (43,200 minutes) and a 365-day year (525,600 minutes).
| Availability | Approx. allowable downtime per 30-day month | strong>Approx. allowable downtime per 365-day year |
| 99% | 7 hours 12 minutes | 3 days 15 hours 36 minutes |
| 99.5% | 3 hours 36 minutes | 1 day 19 hours 48 minutes |
| 99.9% | 43 minutes 12 seconds | 8 hours 45 minutes 36 seconds |
| 99.95% | 21 minutes 36 seconds | 4 hours 22 minutes 48 seconds |
| 99.99% | 4 minutes 19 seconds | 52 minutes 34 seconds |
| 99.999% | 25.9 seconds | 5 minutes 15 seconds |
These figures assume continuous 24/7 measurement and no exclusions, and some values are rounded. Calendar months vary.
The percentage is only the headline. Two “99.9%” SLAs may not measure the same thing. Google’s Compute Engine SLA calculates uptime from total minutes minus downtime minutes, while its Maps Platform SLA compares valid requests with failed ones. One counts time, the other counts transactions.
The agreement should therefore define the service being measured (a server, an application, or a journey such as login or payment), the denominator and measurement period, the monitoring point, and what counts as downtime, including whether severe degradation counts. It should also cover planned and emergency maintenance, third-party dependencies, architecture requirements such as multiple zones or circuits, and exclusions.
The monitoring point matters more than most buyers expect. A server can be “up” while employees cannot use it because sign-in, DNS (the internet’s address lookup), connectivity, or another dependency has failed. For important services, agree a source of truth, consider independent endpoint monitoring, and decide in advance how conflicting data will be reconciled.
Maintenance deserves the same scrutiny. Scheduled maintenance can be a legitimate exclusion because patching reduces risk, but whether it is excluded depends on the agreement. Look for limits: permitted windows, advance notice, maximum duration, an emergency process, and whether overruns become unplanned downtime.
Choose SLA Metrics That Match the Service
One generic KPI list does not fit every provider. The right metrics depend on what the provider actually manages.
| Service | Core metrics | Supporting diagnostics |
| Managed IT support | Meaningful first response and restoration by priority; SLA compliance; backlog ageing; reopened or recurring tickets; user satisfaction | First-contact resolution; escalations; volume by service or site |
| Infrastructure and cloud | User-facing availability; time to restore; error or latency limits for critical journeys; major incidents | Capacity; incident rate; change failure rate |
| Cybersecurity | Time to acknowledge and triage; time to contain; escalation and notification; remediation by severity | Detection source; false positives; repeat control failures |
| Backup and DR | Successful backups and exceptions; verified recoverability; restore-test success; recovery objectives (RTO, RPO), not just “backup completed” | Backup age; failed-job ageing; replication lag |
The split matters. Core metrics protect outcomes and may carry consequences. Diagnostics belong in the monthly report, where they explain performance. Not every useful KPI should become a service-credit-bearing commitment.
Be Careful With MTTR
MTTR is overloaded shorthand. It can mean mean time to repair, restore, recover, respond, or resolve, and Datadog’s overview notes that teams draw the start and end points differently. “MTTR under four hours” is not a usable commitment until the SLA states what starts the timer, what stops it, whether detection time counts, and whether a workaround ends it. It should also say whether customer-wait time is removed, whether measurement is 24/7 or against support hours, and which incidents are included.
Averages need care too. An average can hide a few extremely poor outcomes. Google’s SRE guidance makes the same point about latency, recommending a look at the slow tail rather than the average alone. If severe incidents matter, ask for a per-incident maximum, or for worst cases and percentiles alongside the average.
RTO and RPO
RTO (recovery time objective) is how long a system or process can stay unavailable before the business impact becomes unacceptable. RPO (recovery point objective) is how much data loss, expressed as time, you can tolerate. NIST’s contingency planning guide uses the same concepts. They belong to backup, disaster recovery, replication, and business continuity, not to help-desk resolution. If a provider manages backups but not application recovery, an RTO promise may depend on you, your cloud platform, and your software vendors, so test recovery end to end. See also our article on managed services and disaster recovery.
Examine Exclusions and Remedies
Exclusions are not inherently evasive. A provider cannot reasonably guarantee events outside its control. Legitimate examples include documented planned maintenance, customer-caused configuration changes, unsupported systems, denied access, separately contracted telecom failures, certain third-party failures, force majeure, and delays awaiting required customer approval.
The problem is breadth. Warning signs include excluding every third-party failure even when the provider selected or manages that party, unlimited maintenance windows, pauses triggered by any question, partial outages or degraded service that never count, a blanket exclusion for security incidents, and monitoring that cannot see the user-facing failure.
A useful test is whether each exclusion is specific, causal, evidenced, and proportionate. Who controls the dependency, what evidence supports the exclusion, and what must the provider still do (monitor, open a vendor case, communicate) even when the downtime is excluded?
Service Credits and Repeated Failure
Service credits typically work as a percentage of the affected service fee, applied to a future invoice. They may require a formal claim before a deadline, may be capped, and are sometimes the sole remedy for that failure. The Google Maps Platform SLA shows the pattern: a claim within 30 days, a monthly cap of 50% of the affected service amount, credit applied to future use, and sole-remedy wording. Do not assume credits are automatic: check who claims, with what evidence, and by when.
Credits are better understood as a pricing adjustment and accountability mechanism than as compensation. A credit worth a fraction of one month’s fee can be trivial next to lost sales, payroll disruption, recovery labor, or customer harm.
What matters more is what happens when failure repeats. Consider a graduated path: formal escalation, root-cause analysis, a corrective-action plan, executive review, and enhanced monitoring. If expressly negotiated, chronic-failure provisions may add termination rights, triggered by something precise such as a set number of breaches within a rolling period. Contract rights vary by agreement and jurisdiction. This article is practical vendor-management guidance, not legal advice, so have counsel review remedies and liability clauses.
How to Negotiate a Realistic IT SLA
Negotiation should not mean demanding the fastest possible times. Stricter service levels cost more, because 24/7 staffing, specialist availability, rapid onsite response, spare equipment, redundancy, replication, and geographic resilience all take resources. AWS notes that multi-region designs can raise availability but add cost and complexity, and Microsoft advises weighing the cost of disruption against the cost of preventing or recovering from it. Round-the-clock rapid restoration may be right for authentication, payments, or online ordering. It is hard to justify for a single printer or a payroll archive.
A sequence that works:
- Identify critical business services and user journeys, not just devices.
- Decide how much interruption or degradation each can tolerate.
- Map vendors and dependencies.
- Tier services by business impact.
- Define priorities jointly.
- Separate acknowledgement, engagement, restoration, and resolution.
- Write the clock rules.
- Specify how uptime is measured.
- Test each exclusion.
- Choose a focused scorecard.
- Agree reporting, escalation, and remedies.
- Walk through realistic scenarios before signing and ask for historical data calculated with your proposed definitions.
The last step is the most revealing. A scenario such as “the finance director cannot reach the banking portal on the last day of the month” exposes gaps faster than any clause review. The best SLA is aligned to business criticality, not the one with the shortest times.
How to Monitor IT Vendor Performance
Signing is the start. A practical rhythm, scaled to contract risk:
- Continuously: monitoring of major service health and alerts for approaching breaches.
- Monthly: an SLA report covering breaches, backlog, recurring issues, root causes, and improvement actions.
- Quarterly: a trend and business review covering impact, major changes, risks, security, capacity, satisfaction, and an improvement roadmap.
- After a major incident: a review of the timeline, business impact, communications, causes, recovery performance, and corrective actions.
- Annually and before renewal: a review of scope, critical services, support hours, targets, exclusions, pricing, remedies, and future needs.
This is a risk-based recommendation, not a universal rule.
Look Beyond the Overall SLA Percentage
A provider can report excellent compliance while delivering poor outcomes. Hundreds of P4 tickets can dilute several serious P1 failures. Automatic acknowledgements can make response times look flawless. A workaround can close a ticket that returns every week. The provider’s infrastructure can look available while users cannot log in. Tickets can be closed quickly and reopened repeatedly. An average can hide a handful of severe outliers.
Ask for results by priority and by service, with raw counts and denominators. “98% SLA compliance” says far more when the report shows that 490 of 500 tickets met target, identifies the priorities of the ten misses, and explains their business impact. Also look at breach magnitude, reopened tickets, recurring incidents, ageing backlog, root-cause actions, and whether corrective actions actually get completed.
When Should an SLA Be Revised?
Review performance monthly, but revise the agreement itself deliberately. Reconsider it formally at least once a year and before renewal, and earlier after significant change:
- Growth, new offices, or new operating hours
- Cloud migrations or major infrastructure changes
- Mergers and acquisitions
- New security or privacy requirements
- Significant scope changes
- Persistent SLA failures, or persistent over-performance suggesting targets no longer reflect reality
An early review a few months into a new service also helps. By then you will know whether priorities, volumes, reporting, and clock rules work in practice.
Cloud Dependencies and Local Considerations
A managed service provider’s (MSP’s) SLA and the SLA of an underlying cloud platform are different agreements. AWS, Azure, and Google Cloud each promise platform availability under their own definitions, conditions, and credit processes. The MSP may still be responsible for monitoring, troubleshooting, communication, vendor escalation, workarounds, failover, and recovery coordination. A cloud outage does not automatically answer what the MSP must do during it.
Ask whether the MSP’s response and restoration clock keeps running, who escalates to the cloud provider, who submits any SLA credit claim, and who receives the credit. Ask too whether the architecture meets the platform’s own SLA conditions. Those credits adjust the platform bill. They do not compensate your business loss.
Singapore and Hong Kong
Local rules matter here only because they shape incident-response responsibilities. This is not legal advice.
In Singapore, under the PDPA, an organization that determines a data breach is notifiable must notify the Personal Data Protection Commission (PDPC) as soon as practicable and no later than three calendar days after that determination, according to the PDPC. A data intermediary that discovers a breach must tell the organization it acts for without undue delay, and the organization stays responsible for notification. An MSP may therefore need to supply information, logs, escalation, containment help, and evidence quickly enough for you to meet your own obligations. That does not transfer statutory responsibility to the MSP, and not every security incident is a reportable breach. Singapore’s Cybersecurity Act applies to specifically designated entities such as critical information infrastructure owners, not to ordinary businesses by default.
In Hong Kong, data users remain responsible for protecting personal data when processing is outsourced, and the privacy regulator (PCPD) says they must use contractual or other means to ensure appropriate protection. At the time of writing, breach notification to the PCPD is recommended practice rather than a statutory requirement, so there is no Singapore-style statutory deadline. In both places, document how quickly the provider must notify you and what cooperation, evidence, and containment support to expect.
Questions to Ask Before Signing an SLA
- Which users, sites, systems, and activities are in scope, and which are excluded?
- Which services are 24/7, whose time zone and holiday calendar apply, and how are P1s reported after hours?
- How are priorities defined, who assigns them, and how can we challenge a downgrade?
- Does “response” mean a human or an automated acknowledgement?
- What are the separate targets for acknowledgement, engagement, updates, restoration, and resolution?
- What starts, pauses, resumes, and stops each clock, and does a workaround count as resolution?
- How is availability calculated, and does degraded service count as downtime?
- Which exclusions apply, and what evidence supports each one?
- Can we see the data behind the reports, and past performance under these definitions?
- What happens after repeated failures, and are credits the sole remedy?
- How often is the SLA reviewed, and what triggers an interim revision?
Frequently Asked Questions
Does a 15-minute response SLA mean the problem will be fixed in 15 minutes?
No. It usually means acknowledgement or initial engagement. Restoration and resolution need their own targets.
What does 99.9% uptime mean per month?
About 43 minutes and 12 seconds in a 30-day month, before exclusions, subject to the contract’s definition of downtime.
Who decides whether an incident is P1 or P2?
The SLA should say. Ideally criteria are agreed jointly, changes are recorded, and the customer can challenge a downgrade.
Do SLA clocks stop while waiting for the customer?
Often, but only for narrow reasons such as missing information, approval, or access, with a timestamped request and automatic resumption.
Are service credits applied automatically?
Not necessarily. Many require a claim within a deadline and apply to future invoices.
Treat the SLA as a Working Tool
The best SLA is not necessarily the one promising the fastest response or the most impressive uptime percentage. It is the one that reflects your critical services, defines performance precisely enough to measure, allocates responsibilities clearly, and creates a process for improving service when things go wrong.
Filed away after signing, an SLA is just paperwork. Reviewed monthly, tested against real incidents, and revised as the business changes, it becomes a vendor-management tool that works for both sides. Precise terms protect providers too, because they show what was promised, measured, and outside anyone’s control.
Looking for More Reliable IT Support?
Clear expectations, measurable service, and regular reviews are what make an IT support arrangement dependable. FunctionEight has more than 20 years of experience providing managed IT support in Singapore and Hong Kong, including help desk and onsite support. If you are comparing providers, reviewing your current agreement, or looking to improve how your IT support is delivered and measured, we can help you work out what your business needs. Contact FunctionEight to discuss your IT support requirements.








