When a service fails, the first status update often arrives before anyone knows the cause. That is a reasonable time to write. People need to know whether the problem they see is shared, which work it affects and when they should look again.
You can give them those details without guessing at the repair. Our usual first draft has three parts: the affected service, the observed impact and the time of the next update. Someone else reads it before publication when the incident leaves enough time for that check.
Start with the effect
“API requests are timing out. We are investigating and will update this page by 14:30 UTC.” This gives a reader something they can use. It does not claim that every endpoint is affected, and it does not predict when service will return.
If you know which requests fail, name them. If you have only reports from a few locations, say where those reports came from. Avoid putting internal component names in the first sentence unless readers already know what they mean. A customer trying to submit a form should not need your architecture diagram to understand the update.
Keep a distinction between a workaround and a fix. If retrying a request might create a duplicate record, do not recommend it until you have checked that risk. A short update with one verified fact is better than a longer instruction built on a guess.
Make the next update a commitment
An update time is a promise to communicate. Set one you can meet even if the investigation produces little new information. “We are still investigating” can be useful when it also confirms that the team is present and names the next checkpoint.
Tern keeps written updates alongside the incident timeline and publishes them to your status page. That lets the person taking over read both the technical events and the messages people have already seen. It also reduces the chance of two people posting conflicting recovery estimates.
When a check recovers, take a moment to describe what you have verified. A successful request may be the first sign of recovery, but queued jobs or delayed data can still affect users. Say what is working and what remains under observation.
After the incident, review the updates as part of the follow-up. Look for statements that were hard to interpret, missed update times and workarounds that caused confusion. Keep the corrections in your next draft template. The page should help people make decisions while you repair the service, and it should leave a readable record afterward.



