Monitoring systems for availability and speed, and flagging deviations before users experience problems, is a task that can largely be taken over. Not because AI has suddenly gained expertise in infrastructure, but because the task itself has for years already been built around metrics, thresholds and repetition. That is a different conclusion than for tasks where judgment or negotiation form the core.
Three axes are decisive here: volume, structuredness and error costs.
The volume is high. Monitoring runs continuously, day and night, across hundreds or thousands of metrics. A human keeping track of that manually looks at dashboards at intervals and misses what happens in between. A system that measures every second and compares against a threshold value does not miss that.
The structuredness is high. The task consists of: measuring, comparing against a norm, and issuing a signal in case of deviation. That is a fixed procedure, not an open-ended question. Compare that to answering a user question about software, where the question is phrased differently each time and requires context.
The error costs are low to moderate. A missed or late notification is annoying, but usually recoverable: the system sends a new warning as soon as the deviation persists, and most threshold values are set with margin. That is different from a task where one missed step directly results in a user without a working application, such as resolving a first-line IT incident.
Two axes hold the picture back: judgment latitude and creativity stand at 2 and 1 respectively.
Flagging a deviation is something different from understanding what that deviation means for the organization. A spike in memory usage could be harmless, or it could be the start of a problem that takes down the webshop in two hours. Determining that meaning, and the decision to escalate to an engineer who intervenes, remains human work. AI flags the deviation; a person with knowledge of the environment assesses what that deviation is worth.
That is also precisely why this task does not stand on its own. The signal that a monitoring tool issues must land somewhere: as a registered incident with the right urgency. How that proceeds is described at registering and prioritizing IT incidents, a task that is slightly less tightly defined than the monitoring itself.
The assessment is: an agent. Not a separate script that checks one threshold value, but a system that measures continuously, combines multiple signals, and itself determines whether a pattern is worth a notification before a human sees it. That is a step further than mere alerting, and a step back from full autonomy: the agent flags and categorizes, an administrator decides what happens with the signal.
Two boundary conditions determine whether that works. There must be configured threshold values, tuned to what is normal for those specific systems: a threshold that is strict enough for one application can constantly trigger false alarms for another. And there must be automated alerting in place that actually delivers the signal to someone. Without those two, there is nothing to take over: no norm to measure against, no channel to pass it on through.
At a company with a few servers and fixed office hours, monitoring is often still a matter of checking in every now and then. The volume gain from automation is then limited, simply because the volume is low. At a company with many systems, customers who expect access at any time of day, and a history of incidents that arose at night, the picture is different: there, the FTE capacity freed up by continuous automated monitoring quickly adds up, because the alternative is a human who must be permanently on standby.
The error-cost axis also shifts per company. At an internal test environment, a missed notification has no consequences. At a system that directly touches payment transactions or medical data, the compliance bar is higher, and that pushes the compliance axis, which already stands at 4 here, even further toward mandatory logging and demonstrable follow-up.
This is not a statement about personnel. Whether and how an organization redeploys the freed-up capacity of a systems administrator elsewhere is a choice for the employer, with its own legal requirements where that touches on decisions about roles. This page describes only the work, not the people who currently do it.
Monitoring is rarely an isolated task. It is connected to performing and checking backups, to incident handling, and to the planning of who needs to be available for follow-up and when. Anyone wanting a broader picture of what AI can take over within the IT function as a whole will find a starting point on the page about work and planning.
This page provides an assessment based on the task in its general form. How much that means for a specific company depends on the number of systems, the configured threshold values and the consequences of a missed notification. An indication for your own situation can be obtained with the free quickscan: twelve questions, no account required, with an indication of what portion of the hours in this profile can be taken over by AI today. The full work scan, which breaks down the work of an entire company into tasks, is still under construction.
Vraag maar. Ik ken de kennisbank van deze site; wat ik niet weet, zeg ik erbij.
Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.