ftetoai Join the waiting list

Kennisbank

Human oversight of AI in practice

Why category 2 is the hardest part

At [ftetoai.com](https://ftetoai.com) we divide tasks into three categories: AI can take over the task, AI can partly do it with human oversight, or it remains human work. In practice, that middle category is the most important one, because it does not work on its own. "Oversight" is not a fixed concept with a single interpretation; it only gains meaning once you define who checks what, how often, and what happens in the event of a rejection. This page describes that mechanism. It does not replace legal or employment law advice and does not provide grounds for decisions about personnel — separate statutory requirements and a separate advisor apply for that.

Who approves: which role, which level

Human oversight presupposes a person who is authorized and competent to assess the outcome of a task. That is not automatically the employee who used to perform the task. For a task such as answering personnel questions, the assessment often lies with someone with subject-matter expertise, while for a task such as planning, the assessment is more likely to lie with a manager who can oversee the operational consequences. The question of "who approves" therefore differs per task and depends on what could go wrong if the AI output is used without any check.

What is checked

Oversight without a clear assessment framework is difficult to carry out in practice. In practice it comes down to a limited number of questions: is the outcome factually correct, is it complete, is it in line with applicable rules or agreements, and is it appropriate for the specific situation of the person or case. For recording hours, for example, this means checking for deviating patterns; for tracking leave balances, it more often involves checking exceptions such as special leave or transitional arrangements that are not standardly processed in a system. The assessment framework should be documented in writing, so that the check is reproducible and does not depend on the memory of a single person.

How often: sampling or full review

The frequency of oversight is a separate choice, apart from the question of who approves and what is checked. Two extremes: every outcome is checked before it is used, or a sample is assessed afterwards while the rest proceeds. Intermediate forms are common, for example full review for a new application that is phased down to a sample once the outcomes prove stable. Which frequency is appropriate depends on the consequences of an error: the greater the impact on an individual or the organization, the less room there is for a low check frequency. For tasks that fall under high-risk AI in the workplace, a lower check frequency is generally not advisable, and separate obligations regarding oversight also apply that must be verified separately.

Approving and rejecting with a reason

A core element of workable oversight is that a rejection is given a reason. Without a reason, a rejection is just an isolated observation; with a reason, it becomes a signal that can be used to adjust the AI application, the instruction, or the process. In practice, this usually means a fixed, short format: what was the outcome, what was incorrect, and what adjustment follows from it. This does not need to be a heavy system — for smaller applications a log is often sufficient — but without any record it cannot be demonstrated that oversight was actually exercised, and that demonstrability can become relevant as soon as a task touches on the processing of personal data or on decisions that have consequences for individuals. For the data protection side, see the GDPR when automating tasks.

When you may phase down oversight

There is no fixed percentage or fixed term after which oversight may be reduced; that depends on the nature of the task, the error-sensitivity of the outcomes in practice, and any statutory requirements that apply to that specific task. What does hold as a general mechanism: phasing down is defensible once there is sufficient recorded experience showing that the margin of error remains within an accepted limit, and that limit was set in advance and not chosen afterwards to justify the phase-down. For tasks with consequences for individual employees — think of leave, hours, or pay — a gradual phase-down while retaining sample checks is more common than an abrupt transition to category 1. Whether and when full phase-down is warranted is an assessment that differs per organization and per task, and one that cannot be separated from any sectoral or statutory standards.

Oversight is not a substitute for personnel policy

This page describes how oversight of AI outcomes works as a control mechanism. It says nothing about whether, how much, or which personnel will have less work as a result of AI deployment. That question falls outside the mechanism of oversight and outside this page; decisions about it are subject to their own statutory requirements and belong to the employer and their advisor, not to a tool that categorizes tasks.

What you can do now

To determine whether a task in your organization falls into category 1, 2, or 3, and thus whether and how much oversight is needed, it helps to first have a concrete picture of the tasks that qualify. The free quickscan from ftetoai consists of twelve questions, requires no account, and gives an indication of what share of the hours in your profile could be taken over by AI today. The full work scan, which goes deeper into individual tasks and forms of oversight, is still under development — we therefore do not yet offer that here, but we do offer the quickscan as a first, no-obligation step.

KIPPde assistent van de werkscan

Vraag maar. Ik ken de kennisbank van deze site; wat ik niet weet, zeg ik erbij.

Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.