Re-establishing Problem Solving & Incident Management in IT

Blog

History

I have seen bad, worst & horrific problem solvers, managers, directors, problem solving team(s) & team members during my 12 years of IT career. Their negative classification was drawn from their ability to confuse people during the troubleshooting, discussing the matters which are not related to the problem &/ or pressured the teams or team members. These bad examples certainly provided me with good knowledge of what I should not be doing while resolving any problem, issue or incident.

Introduction

An incident means something is broken somewhere and need to be fixed as soon as possible. Problem solving is an art. Some managers have mastered it while some are still learning. Biggest issue is with people who don’t understand what problem solving is or they have no intention to solve any problem at all.It is important to remove the grass from the root to ensure your plants flourish and grow. Relatively, every manager, director, VP, engineer, customer and vendor must be aware of what exactly is problem solving is and how to solve it promptly & efficiently. It will simplify problem solution in the future, If we understand the whole process bottom to top one more time.

ITIL provides us with the framework for incident management. However, it does not provide insight or guidelines on how to manage people. No IT framework has answer to problem creators or consequence experts & other toxic people. Who has unique talent of amplifying the problem consequences or questioning every proposed solution.

How to solve an incident or problem efficiently & effectively

1. Identify the problem

Problem is an incident and in simple words, incident is an unplanned activity which is disrupting the business, business functions or IT function(s). Identification of the problem is the first step in avoiding it or resolving it. It entirely depends on the incident / issue. Some of these problems can be very huge due to its impact while some can be minor.

2. Understand then define the incident or issue

Based on the identification & description of the issue. Define the problem as short as possible. Such as. X application is down or  Unable to do booking in app or database down, etc.

3. Engage all relevant teams ONLY

Any incident manager would be aware of this task. However, it is a growing trend in IT which is called “all hands on the deck”. It basically means that every team must have their representative on the troubleshooting call, regardless of their requirement. Such as, network team must be on the call even if it is a server issue.  This must be stopped. If you want faster and effective solution then you must engage only the relevant teams. Including all would mean waste of time for every person who is unrelated. Setting up wrong examples & practices in the organization. These people can continue to focus on other tasks if they are not disturbed with unrelated stuff.

4. Less management is better

It is highly recommended to let Incident manager run the show. If he or she is ineffective then Service Delivery Manager or Service Manager can step in to support. All unrelated managers, management, client and unrelated vendors must be excluded from the troubleshooting. All these management folks and unrelated people will simply build pressure on the engineer who is trying resolve the issue. High level management notification will be sent out by incident manager periodically.

5. Give time and freedom to the engineers

Rocket is neither made in a day nor pilot learns to fly plane in a day. similarly, engineers are also human and need time to understand system / application behavior before they take any action. Engineers can do excellent job if they are given freedom to figure out the best & fast solution in any given situation.

6. Explain to all the stakeholders again and again that we are resolving an incident and NOT identifying the root cause.

Some managers are beyond stupid. They will try to dig hole in the middle of the ocean and ask nonsensical questions all the time. Root cause is the best weapon of all stupid people in IT. They don’t understand nor they care that incident resolution is the most important task at hand instead of identifying a root cause. Root cause analysis  must be done after all services are restored and verified. IM / SDM / SM must remind the priorities to these folks every time they ask for root cause or why it happened.

Conclusion

It is like football you must keep repeating the same shot until you get it right. Similar to that, educating everyone again and again is the only solution. We the most experience people in IT has to ensure problem is resolve in the most efficient & effective way. Even if we need to omit some managers  or inform client to wait for an update. Manager is like a general and engineers are the army. No general can win a battle or step into a war with unhappy or under pressure soldiers. Engineers must be provided with time, freedom and adequate resources to continue their job. It will certainly work wonders if followed correctly.

Share It..

You may also like

keyboard_arrow_up