A task that fell into category three last year — human work, no doubt possible — can end up in category two this year: partly transferable, with human oversight that approves or rejects with reason. That is not an error in the first measurement. It is the consequence of how we measure, and why an outcome is always a snapshot.
Your company may not change. But the eight axes on which we assess a task do, independently of your organization. We look at structuredness of the input, judgment scope, customer-contact sensitivity, physical component, creativity, cost of error, compliance sensitivity, and volume/repetition. Of these eight, there are two that are most subject to change outside your company: structuredness of the input and judgment scope.
Structuredness of the input improves as underlying models get better at processing messy, unstructured text — a handwritten note, an email with three subjects at once, a phone call with background noise. What was too unstructured to process reliably last year may be within reach this year. That is exactly why can AI take over answering the phone and transferring calls gives a different answer than two years ago: not because your telephony changed, but because speech recognition and call classification have moved forward a step.
Judgment scope shifts more slowly, but that axis is not static either. As systems become better at recognizing borderline cases — and at explicitly indicating when they are uncertain — part of what was previously entirely human work shifts to the area where AI makes a proposal and a human assesses it. That is the core of category two: not AI that decides, but AI that submits with reason, and a human who approves or rejects.
If two of the eight axes can shift, and the other six are relatively stable because they depend on the nature of the work itself — cost of error and compliance sensitivity do not change because a language model gets better — then the outcome of a measurement is always a combination of something stable and something movable. That translates into a bandwidth, not a percentage.
A wide bandwidth is not a weakness in the measurement. It is an honest representation of the uncertainty that genuinely exists. For a task such as can AI take over triaging a general email inbox, the outcome depends strongly on how much of the incoming email is genuinely unstructured and how much already follows a recognizable pattern. An inbox with mostly standard questions scores differently than an inbox where half the messages describe a unique situation. A single label for "email triage" therefore does not exist without that nuance, and a bandwidth that acknowledges this is more accurate than a seemingly precise number.
There are situations in which a score at a given moment has little predictive value for the following year. This applies especially to tasks that lie close to the border between category two and three, and whose judgment scope depends heavily on exceptions that occur rarely but weigh heavily. In document management, for example, the question can AI take over archiving documents and checking them against retention periods touches strongly on compliance sensitivity: an incorrect check of a retention period can have legal consequences disproportionate to how often the error occurs. For such tasks, a score may turn out favorable this year based on technical progress, while the cost-of-error axis keeps tipping the balance and the task remains stuck in category two or three regardless.
A score also says little when the underlying task description is too broad. "Customer contact" or "administration" are not tasks but categories of tasks, and a re-measurement of these mainly tells you that you need to break the task down further before the outcome means anything.
Because two of the eight axes are movable, a re-measurement over time does make sense — not as an annual formality, but as a response to concrete signals: a new version of an underlying model, a task that has since been organized differently, or doubt that could not yet be resolved last year. A re-measurement that does nothing more than repeat the previous outcome was not worth the effort; a re-measurement that reveals a shift in structuredness or judgment scope, is.
This method does not address what an employer does with that outcome. Decisions about personnel and deployment are subject to their own legal requirements, which are not addressed or substantiated here.
The free quickscan from ftetoai consists of twelve questions, requires no account, and gives an indication of what portion of the hours in your profile can be taken over by AI today. The full work scan, which assesses tasks on the eight axes individually, is still under construction — we deliberately do not offer that here yet, but the quickscan already provides an initial direction.
Vraag maar. Ik ken de kennisbank van deze site; wat ik niet weet, zeg ik erbij.
Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.