A few weeks ago I saw a video of an American CEO speaking Japanese. His lips formed the Japanese words and his voice sounded like his own. The gestures matched too. The catch: he doesn’t speak Japanese.

The tool behind the clip is called lip-sync translation, and the praise it got was deserved on the merits. Software took his original English, translated the meaning, remapped his mouth movements to the new words and reproduced his voice in Japanese. That is a hard problem and it worked.

Strip away the enthusiasm, though, and a plainer description sits underneath. An AI altered the face of a real person so that he appears to say something in a language he does not speak, in words he never uttered. That description has a name: deepfake. Nobody in the coverage I saw used it. Companies selling the tool reach for softer terms instead: localisation, personalisation, scaling content across markets.

I’ll take the eye-rolling when I say, perhaps pedantically, that localisation and forgery use the same technology; the difference is the intent and whether it is disclosed. Whether the video is a tool or a deception turns on whether the viewer knows a translation happened. How many viewers actually know is an open question, since the labelling is not always there.

This is an ethics question, and ethics rarely comes up in the technical excitement around the tool. I’ve written several essays about ethics. For me this is one too.

A society should be able to agree on one thing: a person’s words belong to that person. A quote in an interview can be checked. Someone can dispute it or take it out of context, but the record shows it was said.

Now run the same setup with synthetic speech instead of a quote. A politician appears in a lip-synced, voice-matched video saying something he never said. The clip drives an actual decision. Nothing about that scenario needs invention; the technology in the CEO video already does the work.

Intent is where the line actually runs. A company that puts its CEO’s voice into twenty languages wants reach. Forgery is not the goal. That use is legitimate, and I don’t doubt the people building it mean it that way.

The technology still does more than swap words. It interprets, the way a human translator interprets, except it is easier to check and correct afterwards than a live interpreter working in real time under pressure.

A good human translator takes a thought and finds an expression for it in another language. The result sounds different from the original. That is correct, because the two languages are different, and the gap between them is an honest one. It shows that a translation happened and that the words came from somewhere else first.

Lip-synced translation, left unmarked, removes exactly that gap. With it goes the only sign that a translation happened at all. The video looks like an original, while the translation behind it stays invisible. Companies selling it call the result localisation. Strictly, it is a deepfake, and disclosure is what decides which name applies.