# IT monitoring and availability: spot incidents before users call

> Monitoring that helps instead of adding noise: choose the right measuring points, set sensible thresholds, tie alerts to people, measure availability and be able to hold commitments.

URL: https://techport.ai/en/it-beratung/betrieb-und-support/monitoring-und-verfuegbarkeit

---

1.  [IT Consulting](/en/it-beratung)/
2.  [Operations and Support](/en/it-beratung/betrieb-und-support)/
3.  Monitoring and availability

[Operations and Support](/en/it-beratung/betrieb-und-support)

# Monitoring and availability

By Redaktion techport.ai, IT-Beratung · Last updated on 21 August 2026

The most common way of noticing an incident is still a call from a user. That means time passes between the problem occurring and its discovery in which nobody acts, and that IT has to reconstruct the situation from accounts rather than from data.

Monitoring solves that if it watches the right things. Monitoring that reports everything gets ignored within two weeks and is then worse than none, because it creates a false sense of security.

## How you notice it

*   Incidents are reported by users, not by systems.
*   There are alerts but nobody responds because too many of them are unimportant.
*   After an incident the sequence of events cannot be reconstructed.
*   There are availability commitments but no measurement of whether they are met.

## Why this happens

Monitoring tools get put into service with their default settings, and those report everything technically noticeable regardless of business relevance. Within a few weeks a volume of messages builds up that nobody reviews any more. At the same time the measuring points that actually matter are missing: not whether a service is running but whether a user can enter an order. That business view has to be defined by someone, and that requires the departments.

## How we go about it

1.  **Define from the business.** We establish with the departments which workflows are critical and what an outage means specifically, for example order entry, shipping, production reporting or payment runs. Those workflows get monitored, not just individual machines.
2.  **Set measuring points and thresholds.** We define a few measuring points per critical workflow with thresholds based on experience, and separate information, warning and alert strictly. An alert means someone acts now.
3.  **Tie alerts to people.** We define who is reachable when, through which channel alerting happens and what occurs if nobody responds. An alert without a named recipient is a log entry.
4.  **Measure and report availability.** We measure the availability of the critical workflows, report it regularly and use it as the basis for conversations with departments and providers. Without measurement, availability commitments in contracts are worthless.

## What you gain

*   Incidents that surface before the first request arrives.
*   Fewer false alarms and therefore alerts that get taken seriously.
*   Solid figures for performance conversations with providers.

## From our projects

When setting up monitoring, the hardest part is not the technology but deciding what will not be monitored. We therefore deliberately start with a few measuring points per critical workflow and extend only when an incident occurs that an additional point would have caught earlier. The second recurring finding concerns alerting outside working hours: in many companies everything is technically in place but it is not settled who is reachable at night and at weekends, or whether that person is even authorised to act. That question belongs answered before the first alert fires at three in the morning.

## Häufige Fragen

Which availability metrics make sense?

Four are useful: availability per critical workflow as a percentage of the agreed service time, time to detect an incident, time to restore, and the number of incidents per month by cause. The last is the most important, because it shows whether you are working on symptoms or on causes.

Should we run monitoring ourselves or buy it?

If you have no round the clock on-call cover, external monitoring with an agreed response is often the better choice, because it works in exactly the hours when nobody is looking. Defining what is critical remains your task in any case. No provider can take that part off your hands.

## Let us talk about Monitoring and availability

In a thirty minute first call we work out where your biggest lever sits and whether we are the right people for it.

[Arrange an initial call](/en/kontakt)[Our software](/en/loesungen)

## Further reading

[Operations and SupportPlanning IT infrastructureInfrastructure that carries: plan capacity and replacement ahead, size network and sites, make dependencies visible, set maintenance windows and keep outages manageable.](/en/it-beratung/betrieb-und-support/it-infrastruktur)[Security and ResilienceEmergency management and recoveryWhat happens when IT stops: determine critical processes and time targets, write the emergency plan, rehearse recovery, settle crisis communication and learn from exercises.](/en/it-beratung/sicherheit-und-resilienz/notfallmanagement)[IT Strategy and SteeringSteering IT providersWhat you do yourselves and what you buy: make the sourcing decision, steer managed services, measure performance, limit dependency and negotiate contracts that hold at the exit.](/en/it-beratung/it-strategie-und-steuerung/sourcing-und-dienstleister)[KnowledgeIT metricsDefinitions and formulas read the same way across the company.](/en/it-beratung/kennzahlen)[KnowledgeIT glossaryTerms from IT, software and security, briefly explained.](/en/it-beratung/glossar)

Back to the field [Operations and Support](/en/it-beratung/betrieb-und-support)

Rt

Written by

[Redaktion techport.ai](/ueber-uns), IT-Beratung

Mehr als 15 Jahre Erfahrung in IT-Projekten des Mittelstands, Auswahl und Einführung von Unternehmenssoftware, Aufbau von IT-Betrieb und Informationssicherheit in wachsenden Organisationen.

[More about us](/en/ueber-uns)

More from techport.ai

[

Software

Custom process software for mid-sized companies.

](/en/loesungen)[

HR consulting

People processes and the systems behind them.

](/en/hr-beratung)[

IT maturity check

Ten minutes to a clear position.

](/en/it-beratung/reifegrad-check)[

HR maturity check

24 statements, a result per field.

](/en/hr-beratung/reifegrad-check)[

Funding

BAFA grant plus more than 50 programmes for delivery.

](/en/foerderung)[

Process in practice

How workflows become reliable software.

](/en/sop-praxis)[

Data and AI

Analysis, forecasts and assistance systems.

](/en/daten-ki)[

About us

The people behind techport.ai.

](/en/ueber-uns)
