Turning Speech into Business Insight

Verbit’s Yair Amsterdam joins Yoel Israel to discuss the company’s move beyond transcription, the demands of enterprise voice technology, and why accuracy depends on more than a generic AI model.

Billions of conversations take place every day, but much of what is said remains unstructured and difficult for organizations to use.

Verbit began as a transcription company. Today, Yair Amsterdam describes it as a voice intelligence platform that captures spoken content, understands its context, and applies that information to specific business needs.

The company serves more than 2,500 customers across media, education, legal services, and government. Its technology can support transcription, captions, translation, dubbing, summaries, accessibility, and real-time analysis.

In his conversation with Yoel, Yair explained why enterprise voice technology requires specialized models, human preparation, and deep knowledge of each industry.

What Is Voice Intelligence?

Voice intelligence transforms spoken content into structured information that organizations can search, analyze, and use.

Transcription is the first step, but Yair sees many possibilities beyond producing a written record. A platform can summarize a live class, translate a broadcast, generate a quiz, identify inconsistencies in legal

Title tag: How Verbit Turns Speech Into Business Insight

Meta description: Yair Amsterdam explains how Verbit is moving beyond transcription to provide enterprise voice intelligence across media, education, and legal work.

Yair Amsterdam on Verbit’s Voice Intelligence Shift

Verbit’s Yair Amsterdam joins Yoel Israel to discuss the company’s move beyond transcription, the role of people in its AI operations, and how spoken words can become useful business information.

Billions of conversations take place every day, but much of what is said has historically remained unstructured and difficult for software to use.

Verbit began by addressing one part of that problem: turning speech into accurate text. Today, the company is expanding into voice intelligence, using spoken content to generate translations, summaries, accessibility tools, and industry-specific insights.

In his conversation with Yoel, Yair explained how Verbit serves more than 2,500 customers across media, education, legal services, and government. He also discussed why enterprise voice technology requires more than a general transcription engine.

What Is Voice Intelligence?

Voice intelligence captures spoken language, converts it into structured information, and applies that information to a practical outcome.

Transcription remains an important part of the process, particularly for accessibility. Captions allow people who are deaf or hard of hearing to follow a conversation, while written transcripts help audiences consume content in environments where they cannot play audio.

Once speech has been captured accurately, more can be done with it.

Verbit can translate and dub content, provide real-time summaries, generate quizzes, and identify potential inconsistencies in legal conversations. It also supports audio description, which explains visual activity on screen for blind and visually impaired audiences.

Yair describes Verbit’s direction through three actions: capture, understand, and apply. The company started with capture, but its growth increasingly depends on understanding what was said and helping professionals use that information.

Why General Transcription Is Not Enough

Many transcription tools are designed for individual consumers. They can record a meeting, generate text, and provide a summary that is accurate enough for personal use.

Enterprises have different requirements.

A law firm, court, broadcaster, or university may need stronger security, privacy protections, technical integrations, and consistent performance. Its transcription tools may also need to connect directly with existing systems through APIs.

Accuracy becomes more difficult when the subject matter is specialized. A general engine may struggle with legal terminology, medical products, sports commentary, unfamiliar names, or several people speaking over one another.

Verbit prepares its technology for particular industries and individual events. The engine used for a basketball game, for example, can be trained for sports rather than general conversation.

Before a game, the team can provide information about the players, coaches, referees, locations, and relevant terminology. When the engine hears a name that resembles one on the roster, it can prioritize the correct spelling.

The same principle applies to a legal deposition or political debate. Preparing the context before the event helps reduce errors involving important names, places, and technical terms.

How Verbit Combines People and AI

Verbit does not operate as a fully automated AI company. Yair said one of its differentiators is the continued involvement of people.

In the past, a human captioner might transcribe an entire broadcast or conversation in real time. Today, Verbit moves much of that human involvement to the preparation stage.

The people involved are becoming analysts rather than traditional captioners. They collect the terminology needed for a session, select the appropriate engine, and prepare it for the expected speakers and subject matter.

Verbit’s “lingo team” monitors current events for new names, locations, and phrases that may appear in broadcasts. The team also prepares for scheduled events such as basketball games, award ceremonies, and political debates.

This model allows AI to perform more of the live transcription while people concentrate on the context required for a professional result.

Why Accuracy Depends on Context

Yair said a generic transcription engine may achieve approximately 90% to 95% accuracy. For informal use, that may be sufficient because the reader can understand the overall meaning despite occasional mistakes.

Professional settings require more.

According to Yair, a properly prepared Verbit session using the right engine, terms, and context can reach approximately 99% accuracy. That level matters when captions are part of a national broadcast or a transcript will be used in court.

A small error involving an ordinary word may not change the meaning of a conversation. An incorrect name, product, or legal term could be much more consequential.

Voice Signatures and Speaker Identification

Verbit also uses voice signatures to distinguish between speakers.

Yair explained that a system can begin identifying different voices after only a few sentences. Once a voice is connected with a person, the platform can recognize that speaker during later conversations.

This can be particularly useful in legal proceedings, where a transcript must distinguish among a witness, attorney, and other participants.

Voice signatures can incorporate characteristics such as tone, frequency, accent, and other speech patterns. Yair noted that Verbit’s goal is to recognize normal speech rather than test people deliberately attempting to disguise or imitate a voice.

He also emphasized that using a person’s voice for services such as AI dubbing must begin with consent.

Expanding Access to Global Content

Translation and dubbing can help broadcasters, educators, and creators reach audiences in additional markets.

For sports media, AI can turn a broadcast into captions and translated content for audiences watching in different languages. For universities, the technology can support international students attending courses remotely.

Yair said the shift to remote education during COVID created significant demand for Verbit. As universities and organizations moved online, more lectures and meetings needed to be captioned and transcribed.

The same opportunity now extends to new media. YouTubers, podcasters, and other creators want to distribute their work across languages and cultures.

AI dubbing is still developing. Matching the words is only one part of the task. The system must also consider emotion, timing, and the relationship between the audio and the speaker’s lip movements. Verbit works with partners specializing in this area.

Applying Voice Intelligence to Legal Work

Verbit’s expansion into industry-specific applications is also changing its competitive landscape.

The company developed LegalVisor, a tool that captures depositions and provides information during the conversation. Documents and exhibits can be uploaded and indexed before the proceeding begins.

As people speak, the tool can compare their statements with the available evidence and identify possible inconsistencies.

Yair gave the example of a witness claiming to have gone to sleep at 9:30 p.m. while an invoice shows that the person purchased gas at 10 p.m. Connecting the spoken statement with the document in real time could help an attorney identify an important contradiction during the deposition.

Verbit has also launched a simulation tool for training litigators. As the company enters these areas, it increasingly competes with legal technology and professional training providers rather than only transcription companies.

How AI May Change Verbit’s Team

Verbit currently has just under 200 employees, according to Yair.

As automation improves, he expects operations to become more efficient. Developers and researchers can also produce more with AI tools, reducing the need for rapid headcount growth in those areas.

Yair expects more growth in go-to-market roles. Enterprise customers still need people who can explain the product’s value, manage complex sales processes, and understand the customer’s market.

This reflects a broader challenge within Israeli technology. Yair considers Israel’s technical talent a lasting advantage, but he sees a gap in go-to-market execution. His approach is to build sales and marketing teams in the countries where Verbit’s customers operate.

Verbit’s Move Beyond Transcription

Transcription and captioning are becoming increasingly common. For Verbit, the next stage is applying spoken information to the way professionals work.

That may mean helping a broadcaster serve a global audience, giving a student immediate access to a lecture summary, or showing a litigator an inconsistency during a deposition.

The value is no longer limited to recording what someone said. It comes from understanding the words, connecting them with the relevant context, and delivering information that can support the next decision.

Recent Posts

Yair Amsterdam explains how Verbit combines specialized AI and human expertise to turn speech into enterprise insights across multiple industries.
Deep33 partner Yael Barsheshet explains where AI infrastructure is reaching its limits and why Israeli deep tech can help solve them.
Itai Green explains why corporations need startups, how to run effective pilots, and why slow decisions threaten long-term competitiveness.