Operations and Support

    Monitoring and availability

    By Redaktion techport.ai, IT-Beratung · Last updated on

    The most common way of noticing an incident is still a call from a user. That means time passes between the problem occurring and its discovery in which nobody acts, and that IT has to reconstruct the situation from accounts rather than from data.

    Monitoring solves that if it watches the right things. Monitoring that reports everything gets ignored within two weeks and is then worse than none, because it creates a false sense of security.

    How you notice it

    • Incidents are reported by users, not by systems.
    • There are alerts but nobody responds because too many of them are unimportant.
    • After an incident the sequence of events cannot be reconstructed.
    • There are availability commitments but no measurement of whether they are met.

    Why this happens

    Monitoring tools get put into service with their default settings, and those report everything technically noticeable regardless of business relevance. Within a few weeks a volume of messages builds up that nobody reviews any more. At the same time the measuring points that actually matter are missing: not whether a service is running but whether a user can enter an order. That business view has to be defined by someone, and that requires the departments.

    How we go about it

    1. Define from the business. We establish with the departments which workflows are critical and what an outage means specifically, for example order entry, shipping, production reporting or payment runs. Those workflows get monitored, not just individual machines.
    2. Set measuring points and thresholds. We define a few measuring points per critical workflow with thresholds based on experience, and separate information, warning and alert strictly. An alert means someone acts now.
    3. Tie alerts to people. We define who is reachable when, through which channel alerting happens and what occurs if nobody responds. An alert without a named recipient is a log entry.
    4. Measure and report availability. We measure the availability of the critical workflows, report it regularly and use it as the basis for conversations with departments and providers. Without measurement, availability commitments in contracts are worthless.

    What you gain

    • Incidents that surface before the first request arrives.
    • Fewer false alarms and therefore alerts that get taken seriously.
    • Solid figures for performance conversations with providers.

    From our projects

    When setting up monitoring, the hardest part is not the technology but deciding what will not be monitored. We therefore deliberately start with a few measuring points per critical workflow and extend only when an incident occurs that an additional point would have caught earlier. The second recurring finding concerns alerting outside working hours: in many companies everything is technically in place but it is not settled who is reachable at night and at weekends, or whether that person is even authorised to act. That question belongs answered before the first alert fires at three in the morning.

    Häufige Fragen

    Which availability metrics make sense?

    Four are useful: availability per critical workflow as a percentage of the agreed service time, time to detect an incident, time to restore, and the number of incidents per month by cause. The last is the most important, because it shows whether you are working on symptoms or on causes.

    Should we run monitoring ourselves or buy it?

    If you have no round the clock on-call cover, external monitoring with an agreed response is often the better choice, because it works in exactly the hours when nobody is looking. Defining what is critical remains your task in any case. No provider can take that part off your hands.

    Let us talk about Monitoring and availability

    In a thirty minute first call we work out where your biggest lever sits and whether we are the right people for it.

    Further reading

    Back to the field Operations and Support