Enterprise language interpretation software connecting global teams with real-time AI translation

A five-minute video remote interpreting call bills at $45.00 on Washington State’s published contract rate, not the $7.50 the per-minute figure suggests, because the contract charges a 30-minute minimum. That one clause moves more money than any vendor negotiation, and none of the pages currently ranking for this term mentions it. Fora Soft is a software development company that builds interpretation into video products, so this is the version we wish existed: the four federal requirements a VRI feature has to satisfy, prices you can open the source for, the latency budget, and the point where building beats buying.

Key takeaways

Video remote interpreting is the only remote modality federal law writes a performance spec for. 28 CFR 35.160(d) names four requirements, in force since 15 March 2011 — and they read like acceptance criteria, not aspirations.

The billing minimum sets the bill. At $1.50 a minute with a 30-minute floor, a 5-minute call costs $9.00 per effective minute. Fix the minimum before you shop the rate.

The FCC has already priced automation. For Fund Year 2026-27 it pays $0.95 a minute for captioning done by speech recognition alone and $1.45 with a human in the loop — a 34% discount, set by a regulator.

There is no US federal performance standard for machine interpreting. Section 1557’s machine-translation clause, 45 CFR 92.201(c)(3), governs text. Interpreting is unregulated — which is a liability question, not a licence.

A VRI product is two products. Real-time media plus a roster, routing, rate-card and reconciliation layer. Teams that budget for one ship late.

Why Fora Soft wrote this VRI playbook

A VRI platform is two products — a real-time media path and an interpreter roster with routing, rate cards and billing — and we have shipped both. TransLinguist is a remote simultaneous interpreting platform we engineered on MediaSoup and WebRTC, with AI speech-to-speech in 16+ languages and closed captioning in 22. TransLinguist reports 75+ languages and a marketplace of 30,000+ registered interpreters, and in February 2024 it was appointed a preferred supplier on the NHS NOE CPC National Framework for Language Services — on the non-spoken-language lot, available to NHS bodies, councils, schools, police and fire and rescue.

The other half is duller and matters more. For a hospital client we built an over-the-phone interpreting system on SIP and FreeSWITCH: a doctor picks up any ward landline, selects a language from an IVR menu and is connected to a live interpreter — no app, no training. Behind it sits the part nobody demos: queue and priority routing, interpreter schedules, call history, payments. That is where interpreting projects actually fail.

Fora Soft has shipped 250+ projects since 2005 with 50 in-house engineers, most of them real-time video and AI, and this guide was written by the engineers on those two interpreting builds rather than by a content team. Every number below is linked to the document it came from — a state contract, a Federal Register rule, an FCC order, a vendor price page or a peer-reviewed paper — and dated. Where a figure is a vendor claim rather than a measurement, we say so.

Adding interpreting to a product you already ship?

We’ll walk your stack, name the modality that fits, and tell you what the compliance surface costs before you commit engineers.

Book a 30-min scoping call →WhatsApp →Email us →

What is video remote interpreting

Video remote interpreting (VRI) is interpreting delivered over a live video link by an off-site interpreter, so that a spoken-language or signed-language conversation can happen without anyone travelling. US regulation treats it as a specific auxiliary aid: 28 CFR 36.303(b)(1) lists “qualified interpreters on-site or through video remote interpreting (VRI) services” side by side, and 28 CFR 35.104 defines a qualified interpreter as one who can interpret effectively, accurately and impartially “via a video remote interpreting (VRI) service or an on-site appearance.”

Two things follow that trip up most product teams. First, the parties do not have to be in the same room. Several vendor pages assert that VRI means both parties share a device while the interpreter is remote, and coin “virtual interpreting” for the everyone-remote case. That distinction is marketing, not law — the regulation is silent on where the participants sit. Second, VRI is not VRS. Video Relay Service is a telephone service for deaf callers, funded by the FCC-administered Telecommunications Relay Service Fund, free to the eligible user and available only when the parties are in different places. If you are building a product, VRS is a compensation scheme you probably cannot join; VRI is a feature you can ship.

Interpreting modalities compared: published US rates, billing minimums and the real cost of a 5-minute call

Figure 1. Human lanes priced from Washington contract 18222 (2023 award); the automated lane from cloud list prices, July 2026. The 5-minute column is where budgets actually go.

Reach for VRI when: the conversation needs visual information — a signed language, a gesture, a wound, a document held up to camera — and an on-site interpreter cannot be there inside the clinical or operational window.

How does video remote interpreting work

A VRI session has five moving parts: a request, a match, a media session, an interpreted turn cycle, and a billing record. The request carries language pair, modality, urgency and often a skill tag (medical, legal, certified deaf interpreter). The match is a queue: the platform finds an available interpreter with that pair and skill, and the clock on your service level starts here, not when the video connects.

The media session is ordinary WebRTC with two unusual demands. The interpreter needs a stable view of both parties at a frame rate high enough for hands, and the room needs audio clear enough for two people speaking two metres from a tablet. Then the turn cycle: for spoken languages the interpreter usually works consecutively (speak, pause, interpret) in short encounters and simultaneously in long ones; for signed languages it is always simultaneous, which is why the video path cannot drop frames.

The billing record is the part engineers forget and finance never does. Start time, end time, language pair, modality, interpreter ID, minimum applied, rounding applied. Get that wrong and reconciliation becomes a monthly spreadsheet argument — we have inherited three of those.

What good looks like on the clock

Public contracts are the most honest benchmark available, because a vendor has to sign them. Washington State’s statewide interpreter contract requires that the contractor “answer all VRI calls within thirty (30) seconds, based on a monthly average of all VRI calls”, and separately requires 95% quarterly fill on listed languages and 80% across all requested languages. Vendor marketing quotes anything from 13 seconds to under 60 with no stated methodology; a contractual 30-second monthly average is a number you can hold someone to.

VRI vs OPI vs RSI vs on-site

There are five ways to deliver interpreting remotely or in person, and they differ less in price per minute than in what they can carry and what they cost when the encounter is short: on-site interpreting, video remote interpreting (VRI), over the phone interpretation (OPI), remote simultaneous interpretation (RSI) and an automated lane. VRI is not RSI: VRI serves a two-party conversation, RSI serves one speaker and an audience in several languages at once, and they are separate products with separate standards.

ModalityCarriesPublished unit rateBilling floorWhere it breaks
On-site interpretingEverything, including signed languages and physical context$50.27–$71.82 per hour2 hoursLead time and travel. A 15-minute appointment costs two hours
Video remote (VRI)Speech plus visual information, if framing and frame rate hold$1.26–$1.80 per minute30 minutesDevice placement. A patient who cannot see the screen gets nothing
Over the phone (OPI)Speech only$0.58–$0.83 per minute30 minutesBlind to gesture, to documents and to signed languages
Remote simultaneous (RSI)Conference-grade speech, multiple language channels at oncePer event or per hour, rarely publishedPer eventOverkill for two-party encounters; needs interpreter pairs and relay
Automated laneSpeech, one direction at a time, no signed languages$0.03–$0.15 per minutenoneAccuracy on names, numbers, dosages; no legal performance standard

Rates come from Washington State DES statewide contract 18222 (awards to Lionbridge and Prisma, 2023) and, for the automated lane, cloud list prices retrieved in July 2026. Treat them as a floor for public-sector buying and a sanity check on any quote you receive: a vendor asking $4.00 a minute for spoken-language VRI is asking two to three times what a US state pays.

Reach for OPI when: the encounter is speech-only, unscheduled and short — a pharmacy callback, an eligibility question, a delivery dispute. It is roughly half the price of VRI and connects faster because the interpreter pool is larger.

Reach for RSI when: one speaker addresses an audience in several languages at once — a town hall, an investor call, a training session. This is a different product with its own standard, ISO 24019:2022, and we cover it separately in our simultaneous interpretation software guide.

If you ship VRI into any US public entity or place of public accommodation, four requirements are already written for you. They sit in 28 CFR 35.160(d) for public entities and 28 CFR 36.303(f) for public accommodations, added by Attorney General Orders 3180-2010 and 3181-2010 (75 FR 56183 and 56253, 15 September 2010) and effective 15 March 2011. Quoting the regulation, a public entity that provides qualified interpreters via VRI must ensure it provides:

  • Real-time, full-motion video and audio over a dedicated high-speed, wide-bandwidth video connection or wireless connection that delivers high-quality video images that do not produce lags, choppy, blurry, or grainy images, or irregular pauses in communication”
  • A sharply delineated image that is large enough to display the interpreter’s face, arms, hands, and fingers, and the participating individual’s face, arms, hands, and fingers, regardless of his or her body position”
  • A clear, audible transmission of voices
  • Adequate training to users of the technology and other involved individuals so that they may quickly and efficiently set up and operate the VRI”

Read those as engineering acceptance criteria and they get sharp fast. Requirement 1 makes frame rate, freeze events and jitter pass/fail conditions, which means you need per-session telemetry you can produce later. Requirement 2 is a camera, framing and screen-size rule: a tablet clamped to a cart and pointed at a face fails the moment the patient signs below the frame. Requirement 3 is room audio, not headset audio. Requirement 4 makes onboarding a product feature — the metric is how long an untrained nurse takes to reach an interpreter, and we hold ours under 60 seconds.

The four ADA video remote interpreting requirements turned into testable engineering acceptance criteria

Figure 2. 28 CFR 35.160(d) on the left, what you have to be able to measure on the right.

Video remote interpreting in healthcare: the second rulebook

Video remote interpreting in healthcare answers to two rulebooks, not one. Healthcare adds a second set of rules that is almost the same and not quite. The HHS Section 1557 final rule (89 FR 37522, published 6 May 2024, effective 5 July 2024, language-access provisions fully implemented by 5 July 2025) restates the VRI standard at 45 CFR 92.201(f). But 92.201(f)(2) asks only for “the interpreter’s face and the participating person’s face regardless of the person’s body position.” The arms, hands and fingers are gone.

That is not a licence to frame tighter. It means a covered health programme is simultaneously subject to the ADA’s hands-and-fingers rule and 1557’s face rule, so you build to the stricter one and document that you did. Section 1557 also adds something the ADA has no analogue for: 45 CFR 92.201(g) sets a three-part standard for audio-only remote interpreting, including real-time audio over a dedicated high-speed connection “without lags or irregular pauses.” If your fallback path drops video and keeps interpreting, that fallback has its own compliance surface.

Reach for a documented framing spec when: any customer is a hospital, court, school district or public entity. “Our camera is HD” is not an answer to 35.160(d)(2); a written framing and screen-size guideline with a test procedure is.

The accessibility deadline that moved in April 2026

One more date, because most content on the web still has it wrong. DOJ’s ADA Title II web and mobile rule (28 CFR 35.200–35.205, 89 FR 31337, 24 April 2024) incorporates WCAG 2.1 Level A and AA by reference. The original deadlines were April 2026 and April 2027. An interim final rule published 20 April 2026 (91 FR 20902, amending 28 CFR 35.200(b) via AG Order 6742-2026) moved them to 26 April 2027 for public entities serving 50,000 or more people and 26 April 2028 for smaller entities and special districts. HHS made a parallel move for Section 504 at 45 CFR 84.84, now 11 May 2027 and 10 May 2028. If your roadmap was built on the April 2026 date, you have a year you did not know about — and if a vendor is still selling urgency on the old date, that tells you something.

What the law lets you automate

The short answer: more than most compliance teams assume, and with less cover than most vendors imply. There is no US federal performance standard for machine interpreting. 45 CFR 92.201(f) and (g) do not mention AI, machine interpreting or automated speech translation at all.

The clause everyone cites is a different clause. 45 CFR 92.201(c)(3) requires that where a covered entity uses machine translation and “the underlying text is critical to the rights, benefits, or meaningful access of an individual with limited English proficiency, when accuracy is essential, or when the source documents or materials contain complex, non-literal or technical language, the translation must be reviewed by a qualified human translator.” And 45 CFR 92.4 defines machine translation as automated translation that is “text-based and provides instant translations between various languages, sometimes with an option for audio input or output.” So the clause is written around text and HHS guidance applies it to documents — but read the definition again: audio input or output is inside it, and the HHS Office for Civil Rights letter of 5 December 2024 applies the human-review duty to a live exchange, giving the example of an EMT using machine translation while a qualified interpreter is being found. Do not treat your audio path as out of scope.

So automated interpreting sits in a gap, and the gap is wider than the compliance decks admit. Read the nearest constraint precisely: 92.201(e)(4) bars a covered entity from relying on “staff other than qualified interpreters, qualified translators, or qualified bilingual/multilingual staff”. A model is not staff, so that clause does not cleanly reach it — which is exactly why we would not lean on it either way. What does bite is 92.4’s definition of a qualified interpreter for an LEP individual: one who interprets “without changes, omissions, or additions and while preserving the tone, sentiment, and emotional level of the original oral statement.” Nobody has yet argued in court that a model meets that. We would not want to be the first defendant to try.

The one number a regulator has published

The FCC has done what no vendor will: priced the human. In its Fund Year 2026-27 compensation order (1 July 2026 to 30 June 2027) the FCC compensates IP Captioned Telephone Service at $0.95 per minute for service using only automatic speech technology and $1.45 per minute for service using a communications assistant, plus a $0.23 supplement where the assistant is paid above a threshold wage. That is a US regulator valuing the human contribution at roughly 34% of the per-minute rate — the closest thing to an official exchange rate between automated and human language work, and a better anchor for a business case than any vendor deck.

If you sell into the EU, Article 50 applies from 2 August 2026

EU AI Act Article 50 transparency obligations apply from 2 August 2026. Where an AI system interacts directly with a natural person in a genuine two-way exchange, the person must be told they are dealing with AI “from the start of the first interaction.” The only relief is a narrow grace period to 2 December 2026 for the Article 50(2) marking obligation on systems already on the market before 2 August. Penalties reach €15 million or 3% of worldwide annual turnover. The 2026 Digital Omnibus on AI delayed parts of the high-risk regime; it did not amend Article 50, whatever you may have read. In product terms this is a one-sprint change: a disclosure at session start, a marking on synthetic audio output, and a record that both happened.

Not sure whether your interpreting feature is compliant or just plausible?

We audit the video path, the framing, the fallback and the consent trail against 35.160(d), 92.201 and Article 50, and hand you the gaps in writing.

Book a 30-min compliance review →WhatsApp →Email us →

What video remote interpreting cost looks like in 2026

Spoken-language VRI runs $1.26 to $1.80 per minute on Washington State’s statewide interpreter contract, against $0.58 to $0.83 for phone interpreting and $50.27 to $71.82 an hour on site (contract 18222, as awarded in 2023). Those are award prices: the contract escalates Exhibit B every 1 February against Bureau of Labor Statistics indices, so the 2026 numbers are higher. We quote contracts rather than ranges because government contracts are public, and because almost no page on this topic publishes a price you can open the source for.

SourceOPIVRIOn-siteMinimums
WA DES contract 18222, Lionbridge award (2023)$0.59/min Spanish, $0.79/min others$1.50/min all languagesNot in this award30 min OPI/VRI
WA DES contract 18222, Prisma award (2023)$0.58 and $0.83/min$1.26 and $1.80/min$50.27 and $71.82/hourSame
Pennsylvania courts fee schedule (2022)Hourly, same as in personHourly, same as in person$45–$80/hour by certification2 hours, applied to video and phone too
FCC TRS Fund Year 2026-27VRS $4.35–$8.61/min by tierRegulated compensation, not a purchase price
Cloud automated lane, July 2026$0.03–$0.08/minSame stackNone

Three notes on reading that table. The Pennsylvania schedule is the cruellest: its 2-hour minimum applies equally to in-person, video and telephonic assignments, so remote delivery saves travel and saves nothing on the invoice. The FCC row is compensation to providers, not a price you pay — included because it is the only per-minute number for signed-language video that a US agency publishes. And the platform vendors most often shortlisted for enterprise interpreting — Wordly, KUDO, Interprefy, Boostlingo — publish no dollar prices at all as of 26 July 2026. The “$75 an hour” figure for Wordly that circulates in comparison posts traces back to a 2022 blog post on its own site. If you see it quoted as current, the author did not check.

What the automated lane costs, line by line

Three components, list prices retrieved 26 July 2026. Streaming speech recognition: Google Cloud Speech-to-Text v2 at $0.016 per minute falling to $0.004 at volume; AWS Transcribe streaming at $0.01 per minute; Deepgram Nova-3 at $0.0048 per minute, which its own pricing page labels a limited-time promotional rate on streaming. Translation: AWS Translate at $15.00 per million characters, Google NMT at $20.00. Speech synthesis or an end-to-end path: OpenAI gpt-realtime-translate audio at $0.034 per minute, Azure AI Speech on the Standard tier at $2.50 per audio hour for real-time speech translation, ElevenLabs Agents at $0.080 per additional call minute.

Worked example. One hour of two-way automated interpreting on the end-to-end path: 60 minutes × $0.034 = $2.04. On an assembled path with Azure Live Interpreter, $1.00 per input audio hour plus $1.50 per standard-voice output audio hour plus $10.00 per million output characters — call it $3.00 for a talkative hour. The same hour of human VRI at the Washington contract rate: 60 × $1.50 = $90.00. The automated lane is 30 to 44 times cheaper, which is exactly why the interesting question is not cost.

The billing-minimum trap

Here is the arithmetic that should shape your procurement. Washington’s contract, clause 6.15(c): “For OPI and VRI Interpretation assignments, Contractor/Interpreter shall be entitled to a minimum of 30 minutes of compensation at their quoted minute rate regardless of whether the assignment lasts the entire 30 minutes.” Clause 6.15(b) adds that beyond the minimum, VRI and OPI may be rounded up to the nearest quarter hour.

Substitute real encounter lengths into $1.50 a minute with a 30-minute floor and the effective rate looks nothing like the rate card. A 5-minute call: billed $45.00, effective $9.00 a minute. Ten minutes: $45.00, effective $4.50. Fifteen: $45.00, effective $3.00. Only at 30 minutes do you finally pay list. Then quarter-hour rounding takes over — a 32-minute call bills as 45 minutes, $67.50.

Effective cost per minute of a VRI call under a 30-minute billing minimum, from 5 to 60 minutes

Figure 3. The same rate card, six encounter lengths. Short calls cost six times list.

The operational consequence is counter-intuitive: your cheapest lever is not a lower rate, it is fewer, longer sessions. Batching a ward round’s four short conversations into one 30-minute interpreter booking turns $180.00 into $45.00 at the same rate. The second lever is contractual — per-second or per-minute billing with no floor exists, and it is worth more than a 20% discount on a rate you will rarely pay. Ask for the minimum in writing before you ask for the rate.

Reach for a no-minimum vendor when: your median encounter is under 15 minutes. Below that line the minimum, not the rate, is your entire interpreting budget.

Where automation actually pays

Not where you would guess. Because the automated lane has no minimum, its advantage is largest exactly where the human lane is worst: brief, unscheduled, high-frequency contact. A 90-second wayfinding question at reception costs $45.00 through a contracted human interpreter and about five cents through an automated one. Thirty of those a day is $1,350 versus $1.53.

And it is smallest where you might expect it to shine. A 45-minute consultation is $67.50 human versus $1.53 automated — still 44×, but now $66 of difference sits against a conversation where an omission has clinical consequences. That is not a cost decision any more.

The accuracy gap, measured

The best head-to-head evidence is the IWSLT 2023 evaluation campaign, which had humans score English-to-German simultaneous output on a 1–5 scale. On the Non-Native test set — noisy audio, non-native speakers, which the organisers call substantially harder — the human interpretation scored 2.79 (95% CI 2.71–2.87) against 2.38 (2.30–2.46) for the best machine system. Non-overlapping intervals, so on that material the gap is real. Two caveats worth stating, because they cut the other way: the interpreter was not scored on the clean TED-talk set, where the best systems reached 3.10 and 3.08, and the human output was rated as a transcript rather than as audio. Read it as evidence that hard audio still separates humans from models, not as a league table.

For text, the healthcare literature is more pointed. Khoong et al. (JAMA Internal Medicine, 2019) ran 100 emergency-department discharge instructions through Google Translate: 92% of Spanish and 81% of Chinese sentences were rendered correctly, but 2% and 8% respectively carried potential for clinically significant harm. Taira et al. (2021) extended the method to seven languages and 400 statements — overall meaning retained 82.5% of the time, from 94% in Spanish down to 55% in Armenian. Both studies tested the machine translation of their day, 2018 to 2020, so treat the absolute numbers as dated. The per-language spread is the part that has not changed and the part to design around: a lane that is acceptable in Spanish may be indefensible in Armenian, and one global confidence threshold hides exactly that.

Which gives you the routing rule. Automate the short, low-stakes, high-volume contact in your best-supported languages. Escalate to a human on low confidence, on clinical or legal content, on request by either party, and unconditionally for signed languages. Log which lane handled each turn, because that log is your defence.

Reference architecture for a VRI platform

A video remote interpreting platform has one media path and two supply lanes behind a single interface, so a session can move between automated and human interpreting without renegotiating anything. Here is the shape we build to.

VRI reference architecture: SFU media path, human and automated interpreting lanes, routing and evidence layer

Figure 4. One media path, two supply lanes, and the evidence layer a compliance review asks for.

Media

Use an SFU rather than an MCU, with per-track simulcast so the interpreter can be sent a different quality than the participants. Signed languages get their own treatment: prioritise frame rate over resolution, keep denoising and beauty filters off the signing track, and avoid aggressive temporal compression, because fingerspelling is high-frequency motion in a small region and that is precisely what motion-compensated codecs throw away first. On the audio side, the reference figures for interpreting equipment are stricter than telephony: the European Commission’s 2023 specification for portable interpreting equipment asks for 125 Hz to 15,000 Hz ±3 dB, latency at or under 10 ms and SNR at or above 95 dBA, while ITU-T G.722 — standard wideband telephony — has a nominal 3 dB bandwidth of 50 to 7,000 Hz. Negotiate a fullband codec for the interpreter leg and stop treating interpreters as phone users.

Supply

The human lane needs an interpreter console, not a participant view: both parties visible at once, a hand-over button for interpreter pairs on long sessions, relay support for language combinations with no direct interpreter, and mute-on-turn so the interpreter’s own voice does not feed back. The automated lane needs partial-hypothesis streaming so output starts before the sentence ends, and one audio mix per listener language rather than a single shared translated track.

Evidence

Four records, because these are what a Section 1557 or ADA enquiry asks for and none of them can be reconstructed after the fact: session telemetry (P95 latency, freeze events, frame rate, packet loss), a consent ledger (who consented, to what, when, in which language), a modality record (human or automated, and who escalated), and a retention policy per tenant covering recordings, transcripts and BAA scope. On that last point — a VPN is a transmission-security measure, not a substitute for a business associate agreement, and a VPN on its own is not HIPAA compliance. What HIPAA wants is a BAA with each vendor that touches PHI, encryption in transit and at rest, and access controls. We have seen “we use a VPN” offered as the whole compliance answer more than once.

Reach for a blended architecture when: you have enough volume that per-minute vendor fees hurt, but not enough interpreter supply to guarantee fill. Own routing and the roster; rent media and the automated lane.

The latency budget

Interpreting has two clocks and people conflate them. The transport clock is ITU-T G.114 (05/2003), which says that below 150 ms one-way mouth-to-ear delay “most applications, both speech and non-speech, will experience essentially transparent interactivity,” and that delays above 400 ms are “unacceptable for general network planning purposes.” That governs turn-taking. The interpreting clock is ear-voice span: how long after a source word the interpretation of it arrives.

Ear-voice span has been measured. A 2025 PLOS One study of simultaneous interpreting with text (15 professionals, 30 trainees) found professionals’ mean EVS fell from 5,307.87 ms (SD 2,101.44) at a slow source rate to 3,500.05 ms (SD 1,095.08) at a fast one; trainees started at 8,333 ms. So a professional interpreter delivers three to five seconds behind the speaker by design — they need the clause before they can render it.

Latency budget of an automated interpreting lane against ITU-T G.114 and measured human ear-voice span

Figure 5. A realistic automated budget lands near 1,150 ms, roughly three times faster than a human interpreter.

Add up a realistic automated pipeline — capture 20 ms, endpointing 200 ms, streaming recognition to a stable partial 300 ms, translation 150 ms, first synthesised audio byte 250 ms, network and jitter buffer 150 ms, playout 80 ms — and you land near 1,150 ms. Three times faster to first word than a professional human on that measurement. For reference, the IWSLT track caps systems at 2 s Average Lagging for speech-to-text and a 2.5 s starting offset for speech-to-speech, so 1,150 ms is competitive rather than heroic.

The design conclusion is uncomfortable for anyone selling AI interpreting on speed. Latency was never the human bottleneck. Keep the transport under 150 ms because that is what makes turn-taking feel natural, then stop optimising the pipeline and start optimising escalation.

ASL over video: frame rate before resolution

For video remote interpreting ASL work, frame rate matters more than resolution, and the legal reason is 28 CFR 35.160(d)(2): the image must show the face, arms, hands and fingers of both parties regardless of body position. That is a framing and field-of-view constraint, and a tablet on a rolling stand at bedside satisfies it only while nobody moves. Signed languages are where a VRI feature most often fails while every dashboard stays green.

Encode for motion, not for sharpness. Fingerspelling and rapid handshape changes are high-temporal-frequency detail in a small area of frame, and non-rigid handshape change defeats block-based motion prediction, so the residual gets coarsely quantised by rate control. Halving the frame rate degrades legibility materially at the same bitrate. Practical guidance we apply: hold the signing track at 30 fps and let resolution give first under congestion, disable any temporal denoise or frame-rate-adaptive smoothing on that track, avoid virtual backgrounds entirely because segmentation eats fingers, and light from the front so hands are not silhouetted.

The evidence that this matters is blunt. Kushalnagar et al. surveyed 555 deaf ASL users in 2016–2018 and found only 41% rated the quality of VRI in health care as satisfactory. The strongest association in that data is telling: respondents who felt VRI interfered with disclosing health information were roughly three times more likely to be dissatisfied. That is not a bandwidth problem; it is a stack of product decisions about privacy, framing and who else is in the room. And on the automated side, the World Federation of the Deaf and the World Association of Sign Language Interpreters caution jointly against using signing avatars as a replacement for human signers, so we do not offer an automated lane here. If signed languages are in scope, budget for human interpreters and build the video path to their standard.

Reach for a certified deaf interpreter when: the deaf party uses non-standard signing, is a child, has limited language access, or the encounter is legal or high-stakes clinical. A CDI works alongside a hearing interpreter and materially changes comprehension.

Interpreter scheduling and management software

Interpreter scheduling software answers when an interpreter is available; interpreter management software answers who, at what rate, and did we get paid. This is the half of the product nobody demos, and the half that decides whether an interpreting business makes money.

Interpreter scheduling software vs interpreter management software

Interpreter scheduling software answers when: calendars, availability, assignment offers, confirmations, cancellations, no-shows. Interpreter management software answers who, how much, and did we get paid: credentials and expiry, language pairs and skills, rate cards per pair and per modality, minimums and rounding rules, invoicing, and reconciliation against what the client was billed.

The two get sold as one category and are not. If you buy scheduling and your rate card lives in a spreadsheet, you have bought half a system — and the half you skipped is the one that leaks margin. On the hospital OPI system we built, queue and priority routing, interpreter schedules, call history and payment processing all sat in one admin panel for exactly this reason: the routing decision and the billing consequence are the same event.

CapabilityScheduling toolsManagement systemsWhy it matters
Availability and assignmentYesYesFill rate depends on it
Credential and expiry trackingRarelyYesAn expired certification invalidates a court or clinical assignment
Rate card per language pair and modalityNoYesSpanish OPI and Armenian VRI are not the same price
Minimums and rounding rulesNoYesThis is where the invoice diverges from the session log
Interpreter payout and client invoice in one ledgerNoYesOtherwise reconciliation is a monthly manual argument
Fill rate, ASA and abandonment reportingPartialYesContracts like WA DES 18222 make these numbers contractual

Two metrics belong on the wall from day one. Fill rate: share of requests met by a qualified interpreter, tracked per language, because your aggregate 97% can hide 40% in Burmese. Average speed of answer: measured from request to interpreter connected, not from video session start, and reported as a monthly average if you want it to mean what the Washington contract means by it.

Building the roster-and-routing half of an interpreting platform?

That is the system we have shipped twice — a 30,000-interpreter marketplace and a hospital IVR routing layer. We’ll map yours in half an hour.

Book a 30-min architecture call →WhatsApp →Email us →

Video remote interpreting companies and services compared

Video remote interpreting services come in four shapes, and the shape decides what you are actually buying: a language service provider with its own interpreter network, a meeting platform that expects you to bring interpreters, an FCC-funded relay provider, or your own build. Compare them on six things: the billing minimum, contractual speed of answer, fill rate per language, how they satisfy the ADA framing rule on their own hardware, certified deaf interpreter capacity, and what data they hand back. Most “top VRI companies” lists rank vendors without publishing a single criterion. Here are those six in the order that has saved our clients the most money and grief.

1. What is the billing minimum, and does rounding apply after it? Ask for clause-level language, not a sales answer. This is worth more than the rate.

2. What is the average speed of answer, measured how, and will they sign it? Published marketing claims run from 13 seconds to under 60 with no methodology. A contractual monthly average, as in the Washington contract’s 30-second requirement, is the standard to hold.

3. What is fill rate per language, for the ten languages you actually use? Aggregate fill rate is a vanity metric. Ask for the tail.

4. How does the vendor satisfy 28 CFR 35.160(d)(2) framing on its own hardware? If the answer is a resolution number rather than a framing and screen-size guideline, they have not read the rule.

5. Signed languages: are certified deaf interpreters on the roster, and how fast? Spoken-language fill rates tell you nothing about ASL capacity.

6. What data does the vendor hand back? Session telemetry, consent records, modality records, per-language reporting, and an API. If the vendor keeps the data, you cannot answer an enquiry or negotiate a renewal.

What you are buyingWho supplies interpretersAre prices publishedBest fit
Language service providerThe provider, from its own networkOnly where a public contract exists (Washington 18222 names Lionbridge and Prisma)You want one invoice and no roster to manage
Meeting or interpreting platformYou do, or a partner doesNo — Wordly, KUDO, Interprefy and Boostlingo all publish contact-sales pricing as of 26 July 2026You have interpreter supply and need the delivery layer
FCC-funded relay (VRS)The relay providerCompensation is set by the FCC, not sold to you ($4.35–$8.61/min for Fund Year 2026-27)Never — it is a service for deaf callers, not a product input
Your own platformYou doYou set them; the rate card becomes a featureInterpreting is the product and volume justifies the build

One editorial note on the vendor field as of July 2026. The large enterprise providers — LanguageLine, Sorenson, AMN Language Services, Boostlingo, Certified Languages International, Propio — differ far more in supply depth and contract terms than in technology. The platform vendors — Wordly, KUDO, Interprefy — sell the meeting layer and expect you to bring or buy interpreters. Neither group publishes prices. Treat any comparison article that quotes their pricing as out of date until you have seen the quote yourself.

Mini case: interpretation at NHS framework scale

Short version: we rebuilt a remote interpreting business from a bolt-on video tool into a platform that could win public-sector tenders, and the deciding feature was not the media stack but the records it produced.

The situation. A remote interpreting business needed to move from a bolt-on video tool to a platform it could take into public-sector tenders. Interpreters were coordinated by email and spreadsheet, sessions ran on a third-party conferencing tool with no session records, and the language catalogue was limited by whoever happened to be online. Tender questionnaires asked for fill rate per language, session telemetry and an audit trail — three things the setup could not produce.

The plan. We built TransLinguist as a purpose-built interpreting platform: MediaSoup and WebRTC for media, a microservice back end on Node.js with RabbitMQ, and an interpreter marketplace with scheduling, session records and usage-based billing. On top of the human lane we added AI speech-to-speech in 16+ languages, closed captioning in 22, session transcription, a sign-language interface and a speaker-slowdown indicator — plus integrations into Zoom, Google Meet and Microsoft Teams, because enterprise buyers will not move their meetings.

The outcome. TransLinguist reports 75+ languages and a marketplace of 30,000+ registered interpreters, and in February 2024 it was appointed a preferred supplier on the NHS NOE CPC National Framework for Language Services on the non-spoken-language lot — reachable by NHS bodies, councils, schools, police and fire and rescue. On its own published figures, customers see around 50% cost savings and 2× ROI inside two years; those are vendor-reported, not measurements we ran. What we can vouch for is the engineering: the platform now produces the fill-rate, telemetry and audit records that the tenders asked for, which is what made it biddable. Want the same kind of assessment on your stack? Book a 30-minute call and we will tell you which half of the build you are underestimating.

Build, buy or blend in five questions

Buy per minute below roughly 2,000 billable minutes a month, blend when you own interpreter supply but not the media, and build only when interpreting is the product itself. Five questions get you there. Work them in order — skipping to architecture is how interpreting projects stall for a quarter.

Decision tree for interpreting: buy per minute, blend, or build the platform, by four screening questions

Figure 6. The first four questions as a tree; the fifth, below, is the one that overrides the other four.

Q1. Are signed languages in scope? If yes, you are buying human interpreting whatever else you do, and your engineering job is the video path, not the supply. WFD and WASLI oppose signing avatars as a substitute, so there is no automated shortcut here.

Q2. Is interpreting the product, or a feature of the product? If customers pay you for interpreting, the routing, rate card and audit trail are your moat and you build them. If it is a feature of a telehealth or contact-centre product, you are integrating, not building — and the integration is a few weeks of product engineering, not a platform programme.

Q3. Are you under roughly 2,000 billable minutes a month? At the Washington rate that is about $3,000 a month of interpreting. One engineer-week costs more. Buy per minute and negotiate the minimum away.

Q4. Do you own the interpreter supply? If you have a roster, blend: own routing and the roster, rent the media and the automated lane. That is the fastest route to margin for a language service provider and it is where most of our interpreting work sits.

Q5. What is your compliance surface in 24 months? The question that changes the answer to the other four. If your roadmap includes US public entities, NHS-style frameworks or EU deployment, the evidence layer is not optional, and vendors who will not expose session telemetry become disqualifying. Which is the honest reason platform teams end up building: not the media, the records. If you want a second opinion on where your line sits, that is what our language interpretation engineering team does in a first call.

Five failure modes that stall VRI rollouts

1. The device cannot be positioned. Every failed deployment we have reviewed had a hardware story: a cart that will not fit beside a bed, a screen too small for two-party framing, a stand that puts the camera above the patient’s eyeline. Requirement 2 is a physical requirement. Pilot in the actual room, with the actual bed, before you buy 200 units.

2. Nobody defined the wait-time policy. When no interpreter answers in 45 seconds, what happens? Most products have no answer, so the clinician gives up and uses a family member — which 45 CFR 92.201(e) is written specifically to prevent. Define the escalation ladder in the product: retry, switch to OPI, switch to captions, page a coordinator.

3. Consent is collected in English. Consent to record, to transcribe, and to use an automated lane has to be obtainable in the language the person actually speaks, before the session, and stored per session. Retrofitting this into a shipped product is a schema migration and a legal review.

4. The invoice does not match the session log. Minimums and quarter-hour rounding mean the billed duration is a derived value, not the session duration. If your platform stores only start and end timestamps, finance will reconcile by hand forever. Store the applied minimum and applied rounding on the record.

5. The automated lane has no confidence signal. Shipping machine interpreting without a per-utterance confidence score and a threshold means you cannot escalate, cannot audit and cannot tell a regulator why a turn was handled by a model. Given the 94%-to-55% per-language spread in the Taira data, one global threshold is not a policy either.

KPIs to track from day one

Quality KPIs. Per-session P95 glass-to-glass video latency (we hold 400 ms as an internal ceiling; the audio path is budgeted separately against G.114’s 150 ms mouth-to-ear figure), freeze events per session (target zero — requirement 1 forbids irregular pauses), sustained frame rate on the signing track (30 fps floor), and escalation rate from the automated lane to a human, broken out per language. That last one is your accuracy proxy and the number that will move most over a year.

Business KPIs. Fill rate per language (not aggregate), average speed of answer measured from request rather than from session start, billed minutes versus session minutes — the ratio that exposes how much minimum you are paying — and cost per completed encounter by modality. If billed minutes run more than 1.4× session minutes, your minimums are the problem, not your rate.

Reliability KPIs. Session setup success rate, abandonment before interpreter connect, fallback activations by type (video to audio, audio to captions, automated to human), and time-to-interpreter for an untrained user — the operational form of requirement 4, and the one we hold under 60 seconds.

When NOT to use video remote interpreting

Do not use video remote interpreting when the person cannot use the screen, when the encounter is short and speech-only, when privacy is the deciding factor, or when the real gap is interpreter training rather than technology. VRI is the wrong choice more often than vendors admit, and saying so is how you keep a hospital client for a decade.

When the person cannot use the screen. A patient who is prone, sedated, restrained, visually impaired or in acute distress cannot participate in a video-mediated conversation. The Health Resources and Services Administration says so in its own VRI technical assistance sheet: “the use of in-person interpreting is an industry best practice”, followed by a section on when VRI should not be used. Design the on-site path; do not treat it as the failure case.

When the encounter is short, speech-only and unscheduled. Use OPI. It costs roughly half as much, connects faster because the pool is larger, and loses nothing that a spoken-language exchange needs.

When the encounter turns on privacy or on trust. In Kushalnagar et al. (2019), the respondents who reported that VRI interfered with disclosing health information were about three times more likely to be dissatisfied with it. A screen in a shared room changes what a patient will say. Where that is the risk, an on-site interpreter — and for complex language access, a certified deaf interpreter — is the answer, not a better codec.

When the real problem is interpreter quality. This is the finding worth internalising. Flores et al. (Annals of Emergency Medicine, 2012) coded 1,884 interpreter errors across 57 pediatric emergency encounters: 18% had potential clinical consequences overall — 12% with professional interpreters, 22% with ad hoc ones, 20% with none. And the significant predictor was training, not experience: interpreters with 100 or more hours of training made a median of 12 errors versus 33, and 2% versus 12% with potential consequences. Lindholm et al. (2012) found that LEP inpatients who did not receive professional interpretation at both admission and discharge stayed 0.75 to 1.47 days longer. Meanwhile Locatis et al. (2010) found providers and interpreters preferred in-person, but patients rated in-person, video and phone the same. Modality is a smaller lever than professionalism. If you are choosing between a better video platform and better-trained interpreters, the literature says pick the interpreters.

FAQ

What is video remote interpreting?

Video remote interpreting is interpreting delivered over a live video connection by an off-site interpreter, used for both spoken and signed languages. US regulation names it explicitly as an auxiliary aid at 28 CFR 36.303(b)(1) and sets four performance requirements for it at 28 CFR 35.160(d) and 36.303(f).

How much does video remote interpreting cost in 2026?

Published public-sector rates ran $1.26 to $1.80 per minute for spoken languages on Washington State DES contract 18222 as awarded in 2023, and that contract escalates prices every 1 February, so current figures are higher. The rate is also only half the story: the same contract bills a 30-minute minimum, so a five-minute call costs $45.00, an effective $9.00 per minute. Ask for the minimum before the rate.

What is the difference between VRI and VRS?

VRI is a service an organisation buys to communicate with the people it serves. VRS is a telephone relay service for deaf callers, funded through the FCC-administered TRS Fund at $4.35 to $8.61 per minute for Fund Year 2026-27, free to the eligible user and available only when the parties are in different locations.

Is video remote interpreting HIPAA compliant?

It can be, and the platform alone does not make it so. You need a business associate agreement with every vendor that touches protected health information, encryption in transit and at rest, access controls and a defined retention policy for recordings and transcripts. A VPN is a transmission-security measure, not a substitute for a BAA, and not compliance on its own.

Can AI replace human interpreters?

Not yet for anything consequential, and not at all for signed languages. On the IWSLT 2023 Non-Native test set, human interpretation scored 2.79 out of 5 against 2.38 for the best machine system, with non-overlapping confidence intervals; machine translation of clinical text carried potential for significant harm in 2% of Spanish and 8% of Chinese sentences (Khoong et al., 2019). The World Federation of the Deaf and WASLI caution against signing avatars as a replacement for human signers.

Does US law regulate AI interpreting?

No federal performance standard exists for machine interpreting. 45 CFR 92.201(f) and (g) do not mention it, and the Section 1557 machine-translation clause at 92.201(c)(3) is written around text. Note that 92.201(e)(4) bars relying on staff other than qualified interpreters, translators or qualified bilingual staff, so it does not cleanly reach a model; the binding text is 92.4's definition of a qualified interpreter, which requires interpreting without changes, omissions or additions.

What internet speed does video remote interpreting need?

The regulation sets a behavioural standard rather than a bandwidth figure: no lags, choppy, blurry or grainy images, and no irregular pauses. In practice hold 30 fps on any signing track and let resolution degrade first, keep one-way latency under 150 ms where you can (ITU-T G.114), and instrument freeze events per session so you can prove it.

What is the difference between interpreter scheduling and interpreter management software?

Scheduling answers when: availability, assignment offers, confirmations, cancellations. Management answers who and how much: credentials and expiry, language pairs and skills, rate cards per pair and modality, minimums and rounding, invoicing and reconciliation. Buying only scheduling leaves the margin leak in place.

How long does it take to add interpreting to an existing video product?

For an integration against an existing vendor API with one language lane, a small team ships a production path in weeks rather than months. Building the roster, routing, rate-card and evidence layer is a separate product and should be scoped as one. The compliance surface, not the media, is what stretches timelines.

Do we need a video remote interpreting app, or can it live in our product?

It should live in your product. A separate app adds a device, a login and a training burden, and requirement 4 of the ADA standard makes fast setup by untrained staff a legal expectation. Embedding interpreting into the workflow that already exists is the reason our hospital OPI system used the ward landline and an IVR menu instead of an app.

Platform build

AI Interpretation Platform Development in 2026

The architecture-first companion: how to build the platform this article tells you when to build.

Conferences

AI Simultaneous Interpretation: The 2026 Playbook

RSI for conferences and multi-language events, including ISO 24019 delivery-platform requirements.

Vendors

Real-Time Meeting Translation: 3 Best Platforms

Vendor-by-vendor comparison for meeting translation, with 2026 pricing and accuracy notes.

Captions

Real-Time Transcription in Video Calls: 2026 Build Guide

The captioning lane, including the 2027 FCC caption mandate for video conferencing.

Learn

Real-Time Speech Translation for Live Video

Our knowledge-base pillar on translation architecture, latency and accuracy trade-offs.

Ready to ship interpreting that survives an audit?

Three things decide whether a video remote interpreting feature works: the framing and frame rate that 28 CFR 35.160(d) actually demands, the billing minimum that turns a $1.50 rate into $9.00 a minute, and the evidence layer that lets you answer a Section 1557 enquiry a year later. Everything else — codec choice, vendor logos, model names — is downstream of those three.

The automated lane is real and 30 to 44 times cheaper per minute, and it belongs on the short, low-stakes, high-frequency contact in your best-supported languages. It does not belong on signed languages, on clinical detail, or anywhere you cannot escalate to a human in one click and prove afterwards that you could. Build the escalation path first and the cost savings follow.

If the automated lane is the part you are least sure about, that is the work our AI integration team does most often: wiring speech recognition, translation and synthesis into a product that already has users, with the confidence thresholds and escalation seams designed in rather than bolted on.

We have built both halves — a 30,000-interpreter marketplace on an NHS national framework, and a hospital interpreting system that answers on a ward landline. If you are deciding between integrating and building, that is a 30-minute conversation, not a discovery project.

Want a straight answer on buy, blend or build?

Bring your volumes, languages and compliance surface. We’ll tell you which of the three paths fits — buy, blend or build — and what the other two would cost you.

Book a 30-min call →WhatsApp →Email us →

  • Technologies