Two organizations describe the same task. At one, it falls into the category 'AI can take over the task'; at the other, 'it remains human work'. That is not an inconsistency in the method, but the result of the method. The outcome does not depend only on what a task involves, but on the circumstances in which that task is carried out. This page explains which factors make that difference and what that means for the reliability of an estimate.
We never assess a task on a single characteristic. Every task scores on eight axes: structuredness of the input, judgment latitude, customer-contact sensitivity, physical component, creativity, cost of error, compliance sensitivity, and volume/repetition. Two organizations with 'the same' set of tasks can differ sharply on these axes. A planning task with fixed input and high volume turns out differently than that same planning task when the input varies and exceptions are the rule. So it is not about the title of the task, but about the combination of scores behind it.
At low volumes, incidents weigh more heavily than at high volumes. A task that occurs five times a year offers little basis for recognizing a pattern; a task that occurs hundreds of times a day shows more quickly where AI performs stably and where it does not. High volume often makes an estimate more specific, because there is more repetition to test against. Low volume makes an estimate broader, not because the task is unclear, but because there is simply less to generalize from. This plays a role, for example, in the question of whether AI can take over monitoring the progress of goals and projects: with a handful of projects a year, the range is wider than with a portfolio of hundreds of ongoing initiatives.
A task that takes place entirely within structured systems is assessed differently than a task that largely lives in people's heads, email, and loose documents. Not because the task itself changes, but because the extent to which AI has access to reliable, up-to-date information differs greatly. This is clearly visible with budget monitoring: whether AI can take over monitoring budgets and flagging deviations depends heavily on whether the underlying figures are already available in structured, up-to-date form, or whether they first have to be manually gathered and interpreted. Same task, different system setup, different outcome.
The cost of an error is not universal. A misinterpretation in an internal overview has different consequences than a misinterpretation in customer communication or a compliance report. With a high cost of error, the outcome shifts more readily toward category two — AI produces a proposal, a human approves or rejects it — even if the task's scores on the other axes would point toward AI taking over. This is one of the reasons why we never make a determination based solely on 'what the task involves': the context in which an error lands counts just as heavily.
A range is not vagueness, but an honest reflection of what we do not know with certainty. For tasks with a lot of judgment latitude — where two experienced employees would also differ from each other — every estimate is by definition broader. For tasks with a lot of customer contact, it matters that tolerance for AI interaction differs by sector and by customer group, something we cannot measure without your own figures. And for tasks that have just been set up or have just switched systems, there simply is not yet a stable pattern to go on. In all these cases, we opt for a wider range rather than a precise number that would suggest more certainty than we can substantiate.
There are situations in which a single measurement has little value: a process that has just changed, a task that rarely occurs, or an organization still in transition to new systems. A score at such a moment is a snapshot, not a prediction. That is also why a re-measurement is relevant: the same task, measured later, can produce a different outcome because circumstances have changed, not because the first measurement was wrong.
This method says something about tasks, not about people or roles. An outcome is not grounds for a decision about personnel; decisions affecting employment are subject to their own legal requirements, which you assess separately. We provide facts about tasks; you draw the conclusions that fit your organization. As an illustration: even a task like chairing meetings cannot be assessed uniformly, as shown by the question of whether AI can take over leading the weekly team meeting — the answer differs by team, by meeting culture, and by the amount of decision-making latitude in those meetings.
If you would like an initial impression of how this plays out for your own situation, you can fill in the free quickscan from ftetoai: twelve questions, no account required, with an indication of what portion of the hours in your profile could be taken over by AI today. The full work scan, which goes deeper into the eight axes per task, is still under construction — it does not yet exist at this time, and we would rather say so honestly than offer it anyway.
Vraag maar. Ik ken de kennisbank van deze site; wat ik niet weet, zeg ik erbij.
Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.