Problems & Escalation — Unit 17, Upper-intermediate English
Collocations for reporting issues, describing shortfalls, and escalating professionally.
Reading: Incident Report — Login Failures
**Incident summary · 14 June**
Ops **flagged a concern** at 02:10 UTC after **widespread** login failures across EMEA. Errors were **intermittent** for three hours before stabilising. Uptime **fell short of** the 99.9% SLA for the window, and users reported **service degradation** on mobile clients. Roughly twelve percent of sessions failed during the peak.
At 02:14 we declared a **critical incident** and applied a **containment measure** on the API gateway. Support issued a **workaround** within forty minutes. Marketing emails were **put on hold** until **business impact** was assessed. It was **all hands on deck** until traffic normalised. Leadership joined the bridge within fifteen minutes.
RCA points to an **underlying issue** in DNS failover — not the app layer. The misconfigured records acted as a **single point of failure**. Engineering is **getting to the bottom of** the fault and will **remediate** configs; health checks will **mitigate** repeat impact. The **root cause** should be documented before Friday.
Given **severity**-1 impact, we **escalated to** Tier-3 at 03:00 via the standard **escalation path**. Custom reporting fixes are **out of scope** for this bridge. This is a **time-sensitive** fix for enterprise clients — **escalate promptly** if data loss is suspected. I'd like to **raise an issue** with vendor response times in tomorrow's review. No customer data was lost, but trust recovery will take longer than the technical fix.
**Post-incident review · 21 June**
The technical remediation is complete. The review focused instead on the eighty-four minutes between the first customer report and the declaration of a **critical incident**, because that interval — not the DNS fault — accounts for most of the customer impact.
**What happened in those minutes.** The first report arrived at 00:46 from a single enterprise tenant. Tier-1 followed policy correctly: one report, no reproduction, no alert firing, ticket logged as normal severity. Three further reports arrived by 01:20, all handled by different agents on different shifts, none aware of the others. The pattern became visible at 02:10 only because an engineer happened to notice regional latency on an unrelated dashboard.
**Why nobody escalated.** Every individual decision in that chain was correct under the policy as written. The policy requires either an alert or a reproducible failure to **escalate promptly**, and the failure was **intermittent** enough that no agent could reproduce it. Reports were spread across three agents, and nothing in the tooling aggregates unreproduced reports by symptom.
**The uncomfortable finding.** Our monitoring did not detect this outage at all. It was found by a person looking at the wrong dashboard. Health checks queried a DNS resolver inside the same failure domain as the records that failed, so from the checks' perspective the system was healthy throughout. **This is a monitoring failure of a kind that would recur under any policy change**, and it is the more serious of the two findings.
**Actions.** External synthetic checks from outside the failure domain, by 5 July. Symptom-based aggregation of unreproduced reports across agents and shifts. Explicit permission for any agent to declare a **critical incident** on pattern alone, without reproduction, with no requirement to be right.
**Not an action.** We are not adding an approval step to declaration. Two people proposed it; the delay it introduces is precisely the failure mode under review.
Vocabulary from this unit
| Phrase | Use |
|---|---|
| raise an issue / flag a concern | open a problem formally |
| fall short of | miss target or standard |
| widespread / intermittent | scope and pattern |
| severity / root cause | classify and diagnose |
| Phrase | Use |
|---|---|
| underlying issue | deeper cause |
| escalate to | hand to higher tier |
| remediate / mitigate | fix vs reduce impact |
| workaround | interim relief |
| Phrase | Use | Example |
|---|---|---|
| critical incident | severe outage or failure | We declared a critical incident at 02:14 UTC. |
| service degradation | partial loss of performance | Users reported service degradation in EMEA. |
| get to the bottom of | find the true cause | Engineering is getting to the bottom of the DNS fault. |
| all hands on deck | everyone must help urgently | It is all hands on deck until stability returns. |
| business impact | effect on operations or revenue | Assess business impact before communicating externally. |
| escalation path | defined route to raise severity | Follow the escalation path if SLA breach continues. |
| time-sensitive | urgent; deadlines matter | This is a time-sensitive fix for enterprise clients. |
| containment measure | action to limit damage | We applied a containment measure on the API gateway. |
| out of scope | not included in current responsibility | Custom reporting is out of scope for this sprint. |
| single point of failure | one component whose failure breaks the system | The load balancer was a single point of failure. |
| put on hold | pause until issue is resolved | We put marketing emails on hold during the outage. |
| escalate promptly | raise severity without delay | Escalate promptly if data loss is suspected. |
This is the free sample from this unit. The full unit adds 1 more reading passage, comprehension questions with an answer key, the listening exercises, flashcards for the vocabulary above, and a speaking task — with your progress tracked so the next unit unlocks when you are ready for it.
Common questions
What level is this unit?
Upper-intermediate — roughly CEFR B2. If you are not sure of your level, the course places you with a short diagnostic before you start rather than making you guess.
Do I need to pay to use this?
The reading passage and vocabulary on this page are free to read. The rest of the unit — the remaining passages, the exercises, the listening, the answer key and the progress tracking — is part of the paid General English course.
Is this British or American English?
British English spelling and vocabulary, which is what most learners in Azerbaijan are taught and what IELTS expects, though both are accepted in the exam.