Stop Confusing Incidents and Problems — Why the Distinction Is the Foundation of Service Stability
The Simple Rule
Incidents are about restoring service; Problems are about preventing recurrence.
This simple distinction is the foundation of effective service management. Yet many organizations fail to maintain the separation, leading to recurring issues and frustrated users.
ITIL Definitions
Incident: An unplanned interruption to an IT service or reduction in the quality of an IT service.
Problem: The underlying cause of one or more incidents .
Key Differences
|
Dimension |
Incident |
Problem |
|
Purpose |
Restore service quickly |
Prevent recurrence |
|
Reported by |
Customers/Users |
Internal IT members |
|
Focus |
Symptom |
Root cause |
|
Resolution |
Workaround or fix |
Permanent elimination |
|
Timeline |
Immediate |
Can take months |
|
Priority |
Service impact |
Business impact and recurrence |
Why the Distinction Matters
1. Avoiding Recurrence
If you only fix incidents and never investigate problems, the same issues will happen again and again. Incident management restores service; problem management prevents future disruptions.
2. Resource Allocation
Incident management requires immediate attention; problem management can be planned. Without distinguishing between the two, urgent incidents can prevent strategic problem investigation.
3. Knowledge Management
Known errors (problems with documented workarounds) enable faster incident resolution. Without problem management, this knowledge isn't captured.
4. Continuous Improvement
Problem management drives improvement by identifying and eliminating systemic issues.
When Does an Incident Become a Problem?
An incident should become a problem when :
- A Major Incident has occurred
- A pattern of recurring Incidents suggests an underlying cause should be addressed
- An Event has occurred where an underlying cause should be addressed
Real-World Example
Incident: A server crashes. Teams restart the server, restoring service.
Problem: Investigation reveals the server crashed because of a memory leak in the application.
Problem resolution: The application is patched to fix the memory leak.
Result: The incident doesn't recur.
Common Mistakes
|
Mistake |
Consequence |
|
Treating every incident as a problem |
Overwhelmed problem management team |
|
Never escalating incidents to problems |
Recurring issues never fixed |
|
Using the same process for both |
Problems treated as urgent fixes |
|
Not documenting known errors |
Same incident repeats |
Conclusion
The distinction between incidents and problems is not academic—it's operational. Organizations that maintain this separation will have fewer recurring incidents, better knowledge management, and more stable services.
Action Items for Your Organization
- Ensure your ITSM platform has separate Incident and Problem record types
- Train your team on the incident vs. problem distinction
- Create clear criteria for when an incident should become a problem
- Document known errors when problems are resolved
- Measure recurring incident rates
Topic 11: Known Errors and Workarounds — The Knowledge Foundation of Problem Management
Headline: When You Can't Fix It Permanently — How Known Errors and Workarounds Keep Services Running
What Is a Known Error?
When Problem Management identifies the underlying cause and develops a workaround, the problem becomes a "known error" .
A known error is a problem that has been diagnosed and has a documented workaround.
Key characteristics:
- Root cause is identified
- Workaround is documented
- May not be permanently fixed yet
- Knowledge is available for future incidents
What Is a Workaround?
A workaround is a temporary method for achieving the given task when the planned method is not working due to the Problem. The workaround is abandoned when the Problem is fixed .
Workarounds help reduce service interruptions until the Problem is fully resolved .
Characteristics of a good workaround:
- Restores service
- Is documented clearly
- Is easy to implement
- Is safe to use
The Known Error Database (KEDB)
Known Error articles are documented both in the Problem Record and as articles in the IT Service Management tool's Knowledge Base .
The KEDB is the repository for known errors and their associated workarounds. It serves as a critical knowledge resource for incident management teams.
What the KEDB contains:
- Problem description
- Root cause
- Symptoms
- Workarounds
- Resolution (if available)
- Related incidents
Why Known Errors Matter
|
Benefit |
Impact |
|
Faster resolution |
Incidents resolved using known workarounds |
|
Reduced downtime |
Services restored quickly |
|
Increased confidence |
Teams can resolve issues reliably |
|
Knowledge sharing |
Tribal knowledge is captured |
|
Better reporting |
Known errors inform trend analysis |
The Known Error Workflow
- Problem diagnosed: Root cause is identified
- Workaround identified: A temporary fix is developed
- Known error created: Documented in the KEDB
- Incidents linked: Related incidents associated with the known error
- Incident resolution: Agents use the workaround to restore service
- Problem resolution: Permanent fix is implemented
- Known error updated: Resolution is documented, workaround retired
Creating Effective Known Error Articles
Essential information:
- Problem description
- Symptoms
- Root cause
- Workaround steps
- Implementation considerations
- Related incidents
- Resolution (when implemented)
Best practices:
- Use clear, plain language
- Include steps in chronological order
- Note any limitations of the workaround
- Keep articles current
- Link to related knowledge
When Known Errors Are Most Valuable
Known errors are particularly valuable for:
|
Scenario |
Why |
|
High-volume incidents |
Same issue repeatedly resolved faster |
|
Complex systems |
Expert knowledge is captured and shared |
|
New team members |
Access to institutional knowledge |
|
Service desk triage |
First-line support can resolve more issues |
Conclusion
Known errors and workarounds are the foundation of effective incident and problem management. By documenting known errors, organizations capture institutional knowledge, enable faster incident resolution, and build a learning culture.
Action Items for Your Organization
- Establish a Known Error Database (KEDB)
- Create a template for known error articles
- Train teams on how to create and use known errors
- Link known errors to incident records
- Measure how often known errors are used in incident resolution