Will Machine Translation Replace Interpreters? Where the Line Really Is

A conference organizer gets a proposal from a vendor: swap the interpreting booth for an automatic speech translation system, and the budget instantly shrinks. An interpreter who hears about these services reacts the opposite way - as a threat to the profession. Both reactions have a point, and both are incomplete. Here is what automatic speech translation actually does well, where it predictably breaks down, and why the line between "good enough" and "not good enough" does not run where it first appears.

What Machine Speech Translation Already Does Well

Where speech is predictable, automatic translation services perform reasonably well. One speaker reads a prepared text, the room is quiet, the vocabulary is neutral - under these conditions recognition rarely confuses words, and the translation comes out coherent. This is exactly the scenario the technology was built for: a narrow topic, a single voice, minimal surprises.

The same logic applies to written translation: a rough machine translation of a document is a useful working tool if a human then reads and edits it. Our written translation page describes exactly this approach - the machine speeds up the first pass, and the translator stays responsible for the result. The problem starts not where the technology is used as an assistant, but where it is put in place of a person for a situation it was never built to handle.

How the Recognition, Translation, and Synthesis Chain Works

Automatic speech translation is not one process but three steps in sequence: the system recognizes the sound and turns it into text, the text is translated into another language, and the translation is voiced by a synthesized speaker or shown as subtitles. Each step works only with what it received from the one before it and cannot correct that step's mistake - it simply does not see it.

If the first step mishears a word and substitutes something that sounds similar, the second step translates the already distorted text - and translates it smoothly and confidently: the output never signals "I am not sure here." A listener cannot tell an accurate translation from a distorted one.

A human interpreter has a mechanism this chain lacks - feedback. Hearing an odd phrase, a simultaneous interpreter asks again or clarifies from context, and in consecutive interpreting simply asks the speaker to repeat. The recognition-translation-synthesis chain has no feedback loop: a mistake made at the first step is guaranteed to reach the listener.

Where It Breaks Down at a Real Event

The difference between lab conditions and a live room is the difference between the one scenario the technology works for and everything else. At a real event, the conditions described in the first section are almost never met.

  • Accents and dialects. A speaker for whom the working language is not native pronounces words with deviations from the standard - and this is exactly where recognition makes the most errors.
  • Room noise. Air conditioning, rustling, sound from the back rows, and echo in a large room overlap the speaker's voice and reduce recognition accuracy.
  • Interruptions and several people speaking at once. A panel discussion, a question from the floor while someone is answering - an ordinary situation for a person, but a source of garbled, overlapping speech for a recognition system.
  • Unfinished sentences. A speaker starts a thought, cuts it off, and restarts it in different words. A person intuitively discards the first half; the system translates that half too.
  • Industry jargon and abbreviations. To the system these are rare words, and it substitutes the nearest familiar word from its general vocabulary instead.
  • Names and job titles. Surnames, company names, and job titles are not translated - they need to be recognized precisely, and precision is exactly what gets lost most often at the first step of the chain.
Situation at the event How automatic speech translation behaves
One speaker reads a prepared text in a quiet room Recognition works almost flawlessly, the translation comes out coherent
Several people interrupt each other Speech overlaps, and the translation turns into fragments with no clear logic
A strong accent or regional dialect Recognition regularly gets words wrong, and the error moves straight down the chain
Industry jargon, abbreviations, names, and job titles The system substitutes the closest-sounding word from its general vocabulary instead of the industry term
Room noise, echo in a large space Part of the audio is recognized with errors or not recognized at all

The Cost of a Mistake: A Typo in a Post vs. an Error in Supply Negotiations

A machine translation mistake costs something different depending on what is being translated. If a service garbles a paragraph on social media, the reader will at worst smile and get the point from context - the cost is low, and the mistake is easy to spot and fix. In supply negotiations, the cost is different: a quantity, deadline, or payment term rendered imprecisely becomes part of the agreement the parties will later have to honor. The mistake is not visible in the moment here - it surfaces once it is already too late.

This is also why applying text-translation logic to spoken language works worse than it seems. Russian text is on average longer than English, and a simultaneous interpreter constantly applies compression - deliberately tightening a phrase, deciding on the fly what can be dropped and what cannot be touched. A machine does not distinguish the essential from the secondary: it either translates everything and falls behind the pace of speech, or trims mechanically by length - and risks cutting an essential term of the deal.

Bottom line. A typo in a translated post and an error in supply negotiations are events of a different scale, even if in both cases it was "just one word translated wrong." It is the risk level of the event, not the complexity of the speech itself, that determines whether a human interpreter is needed.

Accountability: You Cannot Ask a Machine

A human interpreter at the negotiating table carries professional and personal responsibility for accurately conveying meaning. If a wording is disputed, you can turn directly to the consecutive interpreter and get clarification right there, before the decision is recorded in the minutes. An automatic service offers no such option: what you have in front of you is an algorithm that bears no responsibility for the result. If a disputed wording has already made it into the minutes and no transcript was kept, there is often nothing left to check what was actually said.

A separate question is where the content of the conversation technically goes. Most automatic services process audio on external servers: the audio stream leaves the room before it comes back as a translation. By comparison, infrared simultaneous interpreting systems transmit the signal by infrared beam within the room - it does not pass through walls and cannot be intercepted from the next office, which even radio systems cannot guarantee. For deal negotiations, the difference is fundamental: a human interpreter with local equipment keeps the conversation inside the room, while a cloud service takes it outside by definition.

Where Technology Already Helps the Interpreter

None of this means automatic tools have no place in interpreting - they do, just not as a replacement for the interpreter, but as an assistant. Before an event, automatically processing submitted materials speeds up building the glossary - collecting terms, participant names, and company names. Afterward, automatic transcription of the recording saves time on producing a text version of what was said. And when the outcome of a meeting calls not for interpreting but for written translation of a document, a machine draft is a reasonable starting point for a human to edit, not the final result.

Technology works in a similar way in hybrid formats, where part of the audience is in the room and part joins remotely. The RSI X remote interpreting platform is an example of technology extending what a human interpreter can do rather than replacing them: the interpreter works exactly as in a regular booth, only the sound reaches remote listeners through the platform. Automatic subtitles have a place alongside this as a supporting layer for those who only need the general sense. But as soon as the official part with minutes begins, the audience gets a human interpreter - more on these formats on the online interpreting page.

How to Decide: Three Questions

Before deciding whether an event needs a human interpreter or can get by with an automatic service, it is worth answering three questions.

  1. What happens if the meaning of a phrase gets distorted? A minor misunderstanding that gets rephrased on the spot is low risk. A phrase that becomes part of an agreement or a legal decision is high risk, and automatic translation is not suitable here at all.
  2. Is it acceptable for the content of the conversation to leave for an external server? For an open presentation, this is not an issue. For negotiations over commercial terms, it is a key one.
  3. How many speakers are there, are there interruptions and jargon, or is it one voice reading a script? The closer it is to a live, unpredictable conversation among several people, the smaller the chance that the automatic chain will handle it without losses.

If even one answer points to high risk, there is only one option - a human interpreter. This is more predictable than it looks: rates are calculated in shifts of up to 4 or up to 8 hours rather than by the hour, and interpreters in the booth always work in pairs, switching every 20-30 minutes, because accuracy drops beyond that and numbers and proper names are the first casualties. There is no such thing as "one simultaneous interpreter for eight hours."

Language Simultaneous interpreting, shift up to 4 h Consecutive interpreting, shift up to 4 h
English from RUB 32,000 from RUB 20,000
French from RUB 34,000 from RUB 18,000
Arabic - from RUB 43,000

Full pricing for all languages is on the simultaneous and consecutive interpreting pages.

Bottom line. A human interpreter is needed not because technology is distrusted in general, but because it is risky to hand a machine something you will later have to answer for - to a partner, or to your own meeting minutes.

What This Means for Those Studying to Become Interpreters

For those choosing the interpreting profession, the honest answer is this: the profession is not disappearing, but it is shifting. Translating a single voice reading neutral text from a script is becoming increasingly automated - competing with the machine there will only get harder. Handling interruptions, noise, and jargon, and taking responsibility for the meaning that ends up in the minutes, remains a human's territory - it requires understanding context, not just knowing the language. Automatic tools are worth learning not as an alternative to the profession but as part of the working toolkit: preparing glossaries, transcribing recordings, and working with rough drafts all get faster with their help today. The basic principles of this work are covered on the interpreter training courses page.

Frequently Asked Questions

Can automatic translation be used if there was no time to find a human interpreter?

For a low-risk situation - for example, to get a general sense of a meeting's content after the fact - yes, it is better than nothing. For negotiations where the decision is recorded in the minutes, it is better to move the date than to rely on a translation that no one can double-check.

Is it worth turning on automatic subtitles alongside a human interpreter at a conference?

As an extra layer for stream viewers who have not been given a dedicated translation - yes. As a replacement for the human interpreter for the audience in the room - no: the difference in quality will be obvious right away, especially next to the work of a simultaneous interpreter.

Will automatic translation replace written translators?

For a first rough pass on neutral text, automation is already widely used. Where legal, technical, or medical wording needs to be conveyed precisely, the final word stays with the human who checks the result and puts their name to it.

Should translators and interpreters learn these tools themselves?

Yes, and the sooner the better: it is not a competitor but a tool that frees up time on preparation and transcription for the part of the work a machine cannot replace.

What to Do Next

If the event involves negotiations, minutes, or a narrow industry topic, choosing a human interpreter is justified almost every time, and it is easier to work out which format fits your case in conversation with a vendor. Contact us and describe the event format - we will tell you where a human simultaneous interpreter is essential and where a simpler solution is enough. More articles on interpreting and translation are in our blog.

Зайцев Дмитрий

Выпускник переводческого факультета МГПИИЯ им. Мориса Тореза (1984). Там же прошёл курс синхронного перевода и получил диплом профессионального синхронного переводчика (1990).

Comments

No comments yet. Be the first!

Leave a comment

Comment will be published after moderation

Contact Us

Choose your preferred contact method

Or leave a request