AI oversight

The error that looks right

Emanuel Flury·23 September 2026·6 min read
Two lines of text on a light ground — in the word «клиentseitig», three letters are Cyrillic

A language model slipped three Cyrillic letters into the middle of a German sentence. The sentence read perfectly — and that is exactly the problem.

The sentence on the cover comes from a working session at Skopa. It is grammatically correct, it is factually right, and it contains an error nobody noticed at the time: in the word «клиentseitig», the first three letters are Cyrillic. It surfaced hours later on a phone screen, because the type is larger there. Errors of this kind are why machine-made work needs a second pair of eyes — and why that second pair of eyes has to be organised differently from how most firms picture it.

Not garbled characters, a whole word

An encoding fault damages text at random: accents collapse, characters turn into question marks, the result looks broken. Something else happened here. «кли» is the start of «клиент», the Russian word for client. Mid-sentence, in German, the model reached for the same idea in another script and finished the word in German. It is a substantive slip wearing a typographic disguise.

The mechanism behind it is not a malfunction but a property. A model that handles twenty languages holds the same idea in twenty forms, and those forms sit close together. A word that means the same thing across several languages and sounds alike — Klient, client, клиент — sits especially close. While producing the rest of the word, a small deviation is enough for the form from the wrong language to appear. The model did not reason wrongly; it rendered the right idea in the wrong script.

It also picked the wrong term along the way: in German, «Klient» is a lawyer's client, whereas the software is a «Client». Two errors in one word, and both would have tripped a spell checker. A human reading at speed tripped neither.

Why the eye fails here

That such words get through is neither chance nor carelessness. Several Cyrillic letters are barely distinguishable from Latin ones, and in some typefaces not at all. The Unicode Consortium catalogues these pairs in a standard of their own and calls them «confusables»: strings that look alike but come from different writing systems (Unicode Consortium 2026). Its best-known example is an address spelling «paypal» with two Cyrillic letters — identical to the eye, a different address to the machine.

What is deliberate in a fraud attempt was pure accident here. The effect is the same. And it lands precisely where a business can least afford it:

  • Master data: a company name carrying one foreign letter no longer turns up in the next search. The record is there and stays unfindable.
  • Matching: two spellings that look identical count as different to any matching routine. The customer appears twice, the payment goes unallocated.
  • Reporting: a filter rule fails to bite, a grouping falls apart, and a total is short by an amount nobody misses.
  • Code and configuration: an identifier that looks identical without being so produces an error message that sends you looking in the wrong place.

The human who stops looking

The obvious answer is that someone will read it over. Research on how people work with automation holds an uncomfortable finding here. As early as 1997, Parasuraman and Riley described how people misuse automated systems in predictable ways — they lean on them too heavily once things have gone well for a while (Parasuraman and Riley 1997). Two years later, Skitka and colleagues showed experimentally that people working with a reliable aid miss errors they would have caught without it (Skitka et al. 1999).

That is the real trap. A machine that delivers ninety-nine clean texts trains its reader to stop looking at the hundredth. The better the tool, the more distracted the check — and the more expensive the one case that slips through. An approval that amounts to a single click stops being an approval after two weeks and becomes a habit.

«A check that grows bored is no longer a check.»

Oversight is not distrust, it is ownership

Regulation now says something concrete about this. The EU's AI regulation requires high-risk systems to be built so that people can effectively oversee them — the text states explicitly that the person doing the overseeing has to understand the system's limits and be able to override its output (Regulation (EU) 2024/1689, Art. 14). The framework published by NIST points the same way and asks for risk to be measured across the whole life cycle rather than once at rollout (NIST 2023). For a Swiss SME running neither a high-risk system nor a standards programme, one plain question remains: who answers for what the machine produced, and how does that person tell when something is off?

The answer implies a division of labour that holds up in practice: what a machine can check reliably, let a machine check. What a person has to judge should not reach them unchecked in the first place.

  • Let the machine check the checkable: the error on the cover can be caught in a single line — a test that rejects any character outside the Latin range, with a short list of permitted exceptions. Tests of that kind belong in the build, not on a checklist.
  • Put only matters of judgement in front of a person: tone, figures in context, commitments made to a customer. Anyone also asked to hunt for typos ends up catching neither.
  • Tie approval to an act: whoever approves changes something — a line, a date, a form of address. An approval that demands nothing turns into a formality.
  • Record what the machine proposed and what the person changed. Only that trail shows, three months on, where the system is weak.
  • Spot checks in peace: one text a week, read in full, on paper or in large type. That is exactly how the error in this article came to light.

What follows for everyday work

A tool that produces errors is not a bad tool. A spreadsheet produces them too, and so does a calculator given the wrong input. The difference is that a language model's errors look good: they are grammatically clean, plausible on the facts and pitched in exactly the right tone. That is why they travel further than an obvious mistake, and why the question of oversight is not one of trust in the technology but of how work is organised.

At Skopa that oversight sits at every point where something leaves the house: the draft comes from the machine, the release from a person, and whatever can be checked mechanically is checked by a rule in the build — as of this article, writing systems too. This text was written with machine help, read back and rewritten in places. The error it describes comes from the same family of tools. Both belong together, and both are worth saying.

How we work, and why every run gives the same result, is on its own page. Method →

AI oversightReviewAutomationSMEs

Sources

  1. Europäische Union (2024) Verordnung (EU) 2024/1689 über künstliche Intelligenz, Artikel 14: Menschliche Aufsicht, Amtsblatt der Europäischen Union
  2. NIST (2023) Artificial Intelligence Risk Management Framework (AI RMF 1.0), National Institute of Standards and Technology
  3. Parasuraman und Riley (1997) Humans and Automation: Use, Misuse, Disuse, Abuse, Human Factors 39(2), 230–253
  4. Skitka et al. (1999) Does automation bias decision-making?, International Journal of Human-Computer Studies 51(5), 991–1006
  5. Unicode Consortium (2026) Unicode Security Mechanisms (UTS #39), Version 18.0.0

written by

Emanuel Flury
Emanuel Flury

Founder of Skopa. Nearly ten years of process automation in Fortune-500 environments, today for Swiss SMEs.

intro call

Have a process we should talk about?

An intro call is non-binding and concrete: we look at a real workflow and tell you honestly whether and where automation pays off.