
Interpreter vs translator is the question buyers ask, and it’s the wrong first question. The right one is: text or speech? A translator works with written text. An interpreter works with live speech. And since 2024 there’s a third answer that the top-ranked pages on this topic still barely mention: AI real-time translation, which handles low-stakes speech at roughly one-hundredth the cost of a human. This guide gives you the 30-second answer, a five-question decision tree, the cost-per-minute math, and four worked buyer examples, so you can make the call without a vendor in the room.
Key takeaways
• Translators write, interpreters speak. If the source is a file, you want a translator. If it’s a voice in a room or on a call, you want an interpreter. That single split settles most of the decision.
• AI real-time translation is now the third option. For internal webinars, training, and support chat in 2026 it’s good enough. For sworn testimony or medical-of-record work, it isn’t, and won’t be soon.
• The cost gap is four orders of magnitude. Human simultaneous interpretation runs $5–$13 per minute; AI real-time runs $0.001–$0.20. The right choice is set by stakes and law, not by the price tag.
• Latency decides whether AI feels live. Under 800 ms feels natural; past 2 seconds listeners start talking over the translation. Every 2026 AI system lands somewhere in the 300 ms–4 s band.
• Certification is the buying signal for human work. Ask for ATA (translators), AIIC or a court credential (interpreters). For AI, ask for word error rate on your own audio, not a marketing number.
Interpreter vs translator vs AI: the 30-second answer
In everyday English, “translator” gets used for both jobs. In the language industry, the two words name two different professions, and the mix-up costs money when you brief a vendor.
Translators work with text. Contracts, websites, manuals, marketing copy, subtitles. The work is written, reviewed, and often proofread by a second linguist. A mistake can be caught and corrected before the document ships.
Interpreters work with speech. Meetings, depositions, conferences, doctor visits, hearings, broadcast. The work happens live. There’s no draft. A mistake is in the room the moment it’s made.
Both professions translate meaning, not words. A literal one-to-one mapping rarely survives a language boundary intact, so a good translator and a good interpreter both reach for the equivalent idea, idiom, register, and emotional weight. The skill is the same. The medium is different, and so is the clock: translation is asynchronous (hours to days), while simultaneous interpretation compresses the whole loop into a 1–3 second lag.
AI real-time translation is the newcomer that reshuffles the table. It listens, translates, and speaks or captions, live, at a cost so low it changes which meetings get language support at all. It doesn’t replace certified humans where the record matters. It does replace a lot of the work nobody used to bother translating. The figure below is the whole comparison on one screen.

Figure 1. The three options across the five factors that actually decide the call.
Why trust this comparison
We’re Fora Soft. We build the systems this article compares. Since 2005 we’ve shipped 250+ video and real-time communication projects, and a chunk of them sit exactly on the interpreter-vs-AI seam: platforms where a human interpreter and a machine model do the same job in the same call.
We built TransLinguist, a video platform with a marketplace of 30,000+ certified interpreters across 75+ languages that also runs AI speech-to-speech in 16+ languages; it won the NHS national framework for language services in the UK. We built VOLO, AI real-time translation deployed at Black Hat 2025 (22,000+ attendees), HIMSS, and GDC. And we stabilized Rafiky, a cloud simultaneous-interpretation platform running 200+ languages with a machine-translation fallback for when a live interpreter drops. So the trade-offs here aren’t theoretical. We’ve shipped both sides, watched them fail in production, and priced them for real clients.
One honest caveat up front: we sell custom software, so when the answer is “buy an off-the-shelf interpreter from an agency,” this guide says so plainly. Most of the time it is the answer. The interesting question is the slice where it isn’t.
Not sure which side of the line you’re on?
30 minutes with a senior engineer who has shipped real-time translation on WebRTC, LiveKit, and custom SFUs. Bring your room, your stakes, and your budget.
The four modes of interpretation
Once you know you need spoken-language work, the next question is which mode. There are four, and the setting picks it for you.
| Mode | Setting | Latency | Equipment | Typical cost |
|---|---|---|---|---|
| Simultaneous | Conferences, UN sessions, broadcast | 1–3 seconds | Booth, headsets, two interpreters | $150–$400/hr each |
| Consecutive | Depositions, medical visits, meetings | Pause-and-translate (2× meeting length) | None required | $35–$120/hr |
| Whispered (chuchotage) | 1–2 listeners inside a larger meeting | 1–3 seconds | None | Near simultaneous, slightly lower |
| Sight translation | Courts: statements read aloud | Real-time read-out | None | Rolled into consecutive rate |
Simultaneous has the interpreter speaking while the speaker is still talking, one to three seconds behind. Two interpreters per language isn’t a luxury, it’s a working requirement: the cognitive load is brutal, and AIIC, the International Association of Conference Interpreters, treats a two-interpreter rotation as standard for sessions over 30 minutes.
Consecutive has the speaker talk for a minute, then pause while the interpreter delivers. No booth, usually one interpreter, and the meeting takes twice as long because every point is said twice. It’s the default for depositions and medical visits, where accuracy beats speed.
Whispered interpretation (chuchotage) has the interpreter sit beside one or two listeners and speak the translation into their ear in real time. It taxes the interpreter harder than a booth, with no soundproofing and no rotation partner, so sessions cap at 30–60 minutes.
Sight translation has the interpreter read a written document and deliver it aloud in the target language on the spot. Courts use it for exhibits and sworn statements that were never translated in advance. Most certified interpreters train for it.
Reach for a human interpreter when: the words are spoken, a wrong one has legal or clinical consequences, and the setting is a court, a clinic, a negotiation, or a public stage. Certified humans are the standard because the record has to survive scrutiny.
AI real-time translation in 2026: what it actually is
AI real-time translation is a chain of three things running live: speech-to-text on the source, machine translation in the middle, and text-to-speech or captions on the output. Some systems fuse all three into one model, like OpenAI’s Realtime API or Google’s Gemini Live. For the listener, the result is captions or a synthesized voice in another language, 300 milliseconds to 4 seconds behind the speaker.
The reason buyers are searching this comparison in 2026 is that this third option got good enough to matter, and the association pages that rank for the query still write as if it’s 2019. Four systems lead enterprise procurement now: DeepL Voice, KUDO AI Speech Translator, Interprefy Aivia, and Meta’s open-source speech models. For the vendor-by-vendor numbers, our real-time speech translation vendor benchmarks break down what each one actually publishes.
Where AI real-time translation works
The common thread across the wins is a forgiving room. Listeners are participating, not litigating, and the cost of an occasional wrong word is a little friction, not liability. That describes a lot of business communication: corporate webinars and town halls, internal training and onboarding, customer-support chat, live class translation in e-learning, and pre-recorded video dubbed for a global audience. VOLO runs exactly this pattern for events with tens of thousands of attendees, and nobody in the room is signing an affidavit.
Where it breaks
The failures cluster where a transcript becomes evidence or a mistranslation harms someone. Legal proceedings of record (sworn testimony, depositions, immigration and custody hearings) need a certified human because the transcript has to stand up in court. So do medical interactions where the record matters: informed consent, surgical pre-op, psychiatric assessment. High-stakes public events, where the audience paid to hear a specific speaker rendered a specific way, still want human simultaneous. And diplomacy runs on subtext and the interpreter’s read of the room, which no model has.
Reach for AI real-time when: the audience is internal or best-effort, no transcript goes on a legal or medical record, and you want language coverage for meetings that were never going to get a human interpreter at $200 an hour. That’s most webinars, support queues, and training sessions.
Accuracy and latency: the 2026 numbers
Two numbers govern whether AI translation is usable: how often it’s wrong, and how far behind it runs. Both are measurable, and both are worse than the marketing implies.
On accuracy, the input stage is the weak link. Word error rate for leading speech-to-text on clean conference audio in a major language pair sits around 8–15% in 2026; on accented English, technical jargon, or noisy rooms it climbs to 20–30%. The Hugging Face Open ASR Leaderboard tracks these figures for models like Whisper-large-v3. On translation quality, DeepL Voice scored 96.4 out of 100 in an independent Slator benchmark, though the vendors don’t publish comparable word error rates on standardized voice test sets, so treat cross-vendor accuracy claims as unproven until you run your own audio.
Latency is the one users feel. Under 800 milliseconds, translated speech feels live. Past two seconds, listeners start talking over the interpretation and the conversation collapses into cross-talk. Every 2026 system lands in the 300 ms–4 s band, and where it lands inside that band matters more than a fractional accuracy edge.

Figure 2. Usability holds up to about 800 ms, then falls off a cliff by 2 seconds. Pick a vendor by where it sits on this curve for your region.
The decision tree: five questions
Walk these five questions in order and the answer falls out. Three come from the room, two from legal and finance, so pull both groups into the conversation early.

Figure 3. The five-question walk. Follow the trunk down until a branch sends you to an answer.
1. Is the source text or speech? Text goes to a translator, and you’re done. Speech continues down the tree.
2. Is a certified, legally defensible record required? Court, immigration, sworn statements, contract-binding medical, regulated financial advice: you need a certified human interpreter, full stop. The legal or medical team can use AI privately for prep, but it doesn’t go on the record.
3. Is real-time delivery required? If not, a recorded transcript translated afterward is cheaper and more accurate than live interpretation. If yes, continue.
4. Is the audience over 500, or a high-stakes public event? If yes, human simultaneous interpretation (remote or booth) is the default, with AI as a backup or an accessibility caption track. If no, continue.
5. Is the budget under $2 per minute, or the audience internal-only? If yes, AI real-time is the fit. If no, run a hybrid: AI captions plus a human interpreter on the dominant language pair.
Interpreter vs translator vs AI, side by side
Figure 1 is the glance. This is the detail, including the column most vendor pages skip: where each option breaks.
| Dimension | Human interpreter | Human translator | AI real-time |
|---|---|---|---|
| Medium | Live speech | Written text | Live speech or text |
| Turnaround | 1–3 s live | Hours to days | 0.3–4 s live |
| Accuracy ceiling | Highest, human judgment | Highest, reviewed twice | 8–30% WER on input |
| Legal standing | Court-admissible if certified | Certifiable for filings | Not for the record |
| Cost per minute | $5–$13 | $0.10–$0.30/word | $0.001–$0.20 |
| Scales to many languages | Costly (team per pair) | Linear with word count | Near-flat, add channels |
| Where it breaks | Price and scheduling at scale | Useless for live speech | Accents, jargon, the record |
Reach for a human translator when: the deliverable is written and durable (a contract, a product page, a certified filing), brand voice or legal precision matters, and you have hours or days rather than seconds. Machine translation can draft it; a human should sign off on anything that ships.
Cost per minute: the reference table
Here’s the number people bookmark. These are 2026 list bands from public agency pricing; negotiated, multi-day, and committed-volume rates run lower.
| Option | Cost per minute | What drives it |
|---|---|---|
| Human simultaneous | $5–$13 | $150–$400/hr each, two-interpreter team |
| Human consecutive | $0.60–$2 | Single interpreter; meeting runs 2× longer |
| Hybrid (AI + human on lead pair) | $1.50–$5 | Splits cost and risk across languages |
| AI real-time (cloud API) | $0.04–$0.20 | Per-minute API pricing at moderate volume |
| AI real-time (self-hosted) | $0.001–$0.05 | Amortized on saturated GPUs |
For written work the unit is different: human translation runs $0.10–$0.30 per word, and machine-translation post-editing (a model draft, a human cleanup) brings it to $0.05–$0.15 at some loss of nuance. A 5,000-word site into five languages is $2,500–$7,500 in total at full human rates, roughly $500–$1,500 per language. The figure below plots the per-minute spread on a log scale, because a linear axis can’t show a range this wide.

Figure 4. Four orders of magnitude separate self-hosted AI from a human simultaneous team. Stakes, not price, should pick the row.
Four worked buyer examples
Find the one closest to your situation. The recommendation and the arithmetic come with it.
Example 1: quarterly all-hands, 4 languages, 1,200 employees
Internal session. The CEO presents in English, with Spanish, Portuguese, and Mandarin for distributed offices, plus Q&A. No external attendees. What we’d do: AI real-time captions in all three target languages, with a bilingual employee per region spot-checking the dominant pair on a side channel. Record it, then run machine-translation post-editing on the transcript before posting. Cost: $400–$800 per session in AI fees versus $4,000–$8,000 for full human remote simultaneous. The internal reviewers add no fee. At four all-hands a year, that’s roughly $2,000 against $24,000.
Example 2: deposition with a Spanish-speaking witness
Civil litigation. The witness testifies under oath, the transcript becomes a legal record, and opposing counsel will scrutinize every word. What we’d do: a court-certified Spanish consecutive interpreter, in person or by video link, with credentials from the relevant state or federal body. No AI on the record; the legal team may use it privately to search prior testimony. Cost: $250–$500 per hour, four-hour minimum, non-negotiable. A struck transcript costs far more than the interpreter.
Example 3: marketing website into 5 languages
8,000 source words, growing 15% a quarter. Spanish, French, German, Japanese, Brazilian Portuguese. Brand voice and SEO both matter. What we’d do: human translators with localization specialists for the homepage, pricing, top 20 landing pages, and any legal copy; machine translation plus human post-editing for long-tail pages; a translation-memory tool so an edit never triggers a full retranslation. Cost: $0.05–$0.15 per word on post-edited pages, $0.18–$0.30 on native human pages. Initial rollout $8,000–$20,000; updates $1,500–$4,000 a quarter.
Example 4: multilingual support chat, 50,000 monthly users
Mid-market SaaS, English product, customers writing in nine languages, median ticket 80 words. What we’d do: AI real-time translation on both sides of the chat, a custom glossary for product terms, and a human flag on anything tagged billing, security, or churn risk, auto-escalated to a native-speaking agent. Cost: under $0.001 per message at scale; the flag-and-escalate workflow is the expensive part, and it pays for itself in retention. Human translation of every ticket would run $0.30–$0.80 per message at bulk post-edited rates, or $200,000–$500,000 a year at this volume.
Building translation into a product, not just an event?
We’ve shipped interpreter marketplaces and AI speech-to-speech into live video. Tell us your room and we’ll map the stack and the cost.
What we learned building TransLinguist
The situation. TransLinguist came to us to build a video platform that had to do the thing this whole article is about: put a human interpreter and an AI model on the same call and let the buyer pick per session. Certified interpreters for the courtroom and the clinic; AI speech-to-speech for the webinar and the town hall. The catch was that public-sector clients (the NHS, councils, police, fire services) demanded both certified-human accountability and a cost line that survived a government tender.
The plan. We built a marketplace of 30,000+ certified interpreters across 75+ languages for the work that needs a human, and layered AI speech-to-speech in 16+ languages with closed captioning in 22 for the work that doesn’t. The media path runs on MediaSoup and WebRTC; the AI path chains Google Cloud speech-to-text, Deepgram, and Speechmatics into translation and text-to-speech. It plugs into Zoom, Google Meet, and Microsoft Teams so clients don’t change how they meet.
The result. TransLinguist won the NHS national framework for language services, covering NHS organizations, councils, schools, and emergency services across the UK. Clients report 50% cost savings, an 80% reduction in interpreting costs on the sessions that moved to the platform, 53% higher attendance, and a 2× return in the first two years. The lesson we carry into every build since: the interpreter-vs-AI decision isn’t a one-time procurement, it’s a per-session routing rule, and the platform’s job is to make switching cost nothing. Want a similar assessment? Book a 30-minute call.
Hybrid workflows: AI plus humans
The most common answer in practice isn’t interpreter or AI. It’s both, split by risk. Hybrid setups put AI on the cheap, high-volume, low-stakes channels and a human on the one pair or moment where a mistake would hurt.
Three patterns cover most of it. AI captions plus a human on the lead pair: the language most of your audience speaks gets a certified interpreter, the long tail gets AI captions. AI first pass plus human review: the model produces a draft transcript or subtitle track and a human cleans it before anything is published, which is standard for dubbed video and legal-adjacent transcripts. Human primary plus AI fallback: a live interpreter carries the session, and machine translation with voice-over keeps it running if the interpreter drops, which is exactly how Rafiky holds a session together when a booth goes quiet.
Hybrid costs land at $1.50–$5 per minute, between pure AI and full human, and buy you most of the accuracy for a fraction of the all-human price. The design work is in the routing: deciding, per language and per moment, which lane a given utterance belongs in.
Reach for a hybrid setup when: one language pair carries most of the audience and the rest are long-tail, or when the stakes vary within a single event, so a human covers the keynote and AI covers the breakout rooms.
What to ask before you hire
Whether you’re briefing an agency or a vendor, asking the right question signals you know what you’re buying and gets you a better answer.
Ask about certification. For translators, ATA (American Translators Association) certification tests a specific language pair. For interpreters, look for AIIC for conference work, a state or federal court credential for legal work, and CCHI or NBCMI for medical interpreting in the US; NAATI covers Australia and NRPSI the UK. Certification is competency tested by exam, not a self-declared title.
Ask about mode and direction. For interpretation, confirm the mode (simultaneous, consecutive, whispered, sight) fits the venue. Most interpreters work into their A, or native, language, not out of it, so two-way coverage usually means two interpreters.
For AI, ask for numbers on your own inputs. Word error rate on your domain audio, latency measured from your region, data residency, and whether there’s a custom-glossary endpoint. Insist on a pilot with your real audio before you sign an annual contract; a vendor’s benchmark on clean studio speech tells you nothing about your accented, jargon-heavy conference call.
A decision framework in five questions
If the tree is the diagnostic walk, this is the quick lookup: match your top priority to the pick.
| Your top priority | Pick | Why |
|---|---|---|
| Legal defensibility | Certified human interpreter | Only a credential holds up in court or on a filing |
| Lowest cost per minute | AI real-time (self-hosted) | Under $0.05/min once GPUs are saturated |
| Brand voice in writing | Human translator + memory | Nuance and consistency need a person who signs off |
| Global reach on a budget | AI captions + human QA | Add languages for the price of channels, review the top one |
| Nuance in a small meeting | Consecutive interpreter | Two passes per utterance, no booth needed |
Five pitfalls that blow up multilingual events
1. Booking one interpreter for a long simultaneous session. Solo simultaneous degrades fast past 30 minutes. Budget for two per language pair, always, or the second hour is worse than no interpretation.
2. Putting AI on the legal record. A best-effort transcript with a 15% word error rate isn’t evidence. If a lawyer will ever quote it, you needed a certified human from the start.
3. Trusting a vendor’s accuracy number. The 95%-plus figures come from clean studio audio. Your conference room, with cross-talk and a regional accent, is a different test. Pilot on your own audio or you’re buying a benchmark, not a result.
4. Ignoring latency until go-live. A system that tests fine at 700 ms from the vendor’s datacenter can run 2.5 seconds from your region, which is the difference between usable and unusable. Measure from where your speakers and listeners actually are.
5. Skipping the glossary. Product names, drug names, and legal terms are exactly what a general model gets wrong. A custom glossary is an afternoon of setup that prevents your brand name from being “translated” into something embarrassing.
KPIs: how to know your choice worked
Quality. Track word error rate on a sampled set of real sessions, not vendor demos, and a human-rated adequacy score on a weekly sample. For interpreters, track complaint and clarification rates. Anything above 15% WER on your core language pair is a signal to add human review.
Business. Measure attendance and engagement by language cohort (TransLinguist saw 53% higher attendance once sessions were covered), support CSAT by customer language, and cost per translated minute against the all-human baseline you replaced.
Reliability. Watch end-to-end latency at the 50th and 95th percentile from each region, caption uptime during live events, and fallback rate: how often the human or the backup path had to take over. A rising fallback rate means your primary lane is under-provisioned.
When not to use AI translation
Honesty sells better than a pitch, so here’s the counter-position. Don’t use AI real-time translation when a transcript will become evidence, when a mistranslation could harm a patient, or when the whole point of the event is that a specific person is heard in a specific voice. In those rooms, a certified human isn’t a nice-to-have, it’s the standard of care and, often, the law.
Two subtler cases. Don’t use AI for a language pair where the leading models are still weak (low-resource languages, some dialect pairs), because the word error rate that’s fine in Spanish is unusable there. And don’t use it when the audience is tiny and the relationship matters more than the cost, a one-on-one negotiation, a board conversation, a family medical decision, where a human’s read of the room is the product you’re actually buying.
How Fora Soft builds real-time translation
If your answer is “buy an interpreter for one event,” hire an agency and you’re done. If your answer is “translation needs to live inside our product,” that’s the work we do. We embed real-time translation into video calls, events, and support flows: the speech-to-text, translation, and text-to-speech chain, the routing between AI and human lanes, and the WebRTC or LiveKit transport underneath it.
The architecture questions we work through with clients are covered in our real-time speech translation guide, our multilingual translation for video calls reference, and the engineering detail in our /learn article on multilingual speech translation in calls. When it’s time to scope a build, our AI integration service and the language-interpretation team are where to start.
Want a second opinion on interpreter vs AI for your product?
30 minutes with a senior engineer. No deck, no pitch, just where you sit on the decision tree and what it costs to build.
FAQ
What is the difference between an interpreter and a translator?
A translator works with written text and has time to revise and review. An interpreter works with live spoken language in real time, with no second draft. Both translate meaning rather than words; the difference is the medium and the clock.
Can AI real-time translation replace a human interpreter?
For low-stakes, internal, real-time speech in 2026, often yes. For legal proceedings, medical-of-record interactions, sworn testimony, and high-stakes public events, no. Choose on the consequences of a mistake, not the cost per minute: a wrong word that means annoyance is an AI job, a wrong word that means liability is a certified human job.
Who gets paid more, interpreters or translators?
It depends on mode and specialization, not the job title. Conference simultaneous interpreters command the highest rates, $150–$400 per hour per language pair and up for legal or medical work, because the skill is rare and the cognitive load is extreme. Translators are usually paid per word ($0.10–$0.30), so a high earner is one with volume and a specialized, well-paid domain.
How much does a simultaneous interpreter cost in 2026?
$150–$400 per hour per interpreter for conference-grade work, booked as a two-interpreter team per language pair, so roughly $5–$13 per minute of session, with $200–$500-plus per hour for legal, medical, or financial specialization. Remote simultaneous interpretation (RSI) runs 20–40% cheaper than in-person because there’s no travel or on-site equipment.
What are the four types of interpretation?
Simultaneous (real-time, booth, conference), consecutive (pause-and-translate, depositions and meetings), whispered or chuchotage (real-time into one or two listeners’ ear), and sight translation (reading a written document aloud in another language). Setting picks the mode.
Is a medical interpreter the same as a medical translator?
No. A medical interpreter handles spoken interactions (a patient visit, informed consent) and should hold a CCHI or NBCMI credential in the US. A medical translator handles written documents (records, consent forms, drug labels). Both are high-stakes, and neither should be replaced by AI where the interaction goes on the medical record.
How accurate is AI real-time translation?
The speech-to-text stage runs about 8–15% word error rate on clean conference audio in a major language pair in 2026, rising to 20–30% on accents, jargon, or noise. Translation quality is high for major pairs (DeepL Voice scored 96.4/100 on an independent Slator benchmark), but vendors don’t publish comparable word error rates, so test on your own audio before committing.
How do I choose between hiring an interpreter and using AI translation?
Walk the five-question tree: if the content is speech, real-time is required, certified accuracy is not legally mandated, the audience is internal or moderate-stakes, and the budget is tight, AI is usually right. Move toward a human interpreter as any of those conditions tightens, and consider a hybrid when one language pair carries most of the audience.
What to read next
Vendor benchmarks
Real-Time Speech Translation Vendors in 2026
DeepL Voice, KUDO, Interprefy, and Meta’s models on public data.
Pillar guide
Real-Time Speech Translation for Live Video
Architecture overview and the engineering constraints that bite.
Video calls
Multilingual Translation for Video Calls
Design patterns for embedding translation into WebRTC.
Learn: AI for video
Streaming ASR: Deepgram, Whisper, AssemblyAI
The speech-to-text layer under every real-time translator.
Interpreter, translator, or AI: your next move
Three options, five questions, and cost ranges four orders of magnitude apart. Text goes to a translator. Speech that carries legal or clinical weight goes to a certified interpreter. Everything else, the webinars, the training, the support queue, is now AI territory, often with a human on the one pair that matters. The choice is set by stakes and law, not by the feature list.
If you’re weighing where AI fits your specific room, our vendor benchmarks cover what the leading systems actually publish. If translation needs to live inside your product, that’s a build, and it’s the one we do.
Ready to put translation inside your product?
Bring your room, your audience, and your budget. We’ll map the interpreter-vs-AI split and what it costs to build.

