Business

Boson AI voice model targets cheaper enterprise speech systems

Alex Smola’s startup plans to release Higgs RealTime, saying its speech-to-speech AI costs one-tenth as much as rival systems.

Hana Yoshida

By Hana Yoshida · Markets Reporter

3 min read

Boson AI voice model targets cheaper enterprise speech systems
Photo: Fortune

Boson AI voice model plans are putting Alex Smola’s three-year-old startup into a crowded race with OpenAI, Meta and Microsoft. The Santa Clara, California, company is preparing to release Higgs RealTime, its first speech-to-speech model, and says it can offer live audio AI at far lower cost than competing systems.

Smola, a former distinguished scientist at Amazon and a machine learning researcher, told Fortune that voice has reached a point where it can change how people interact with machines. He said Boson’s systems are “about an order of magnitude more affordable” than rivals while remaining useful, and the company says its models cost one-tenth as much as others.

What is Boson AI building?

Boson AI is building speech-to-speech AI for live conversations, including enterprise voice agents that can talk with customers and respond in real time. The company’s upcoming Higgs RealTime model is meant for spoken interaction rather than text-only chat.

Smola told Fortune that the industry is moving beyond interfaces based only on typed prompts. He said audio and vision can create a richer experience, and he is pitching Boson as a cheaper way for companies to develop and run those systems.

A key part of Boson’s pitch is what Smola described to Fortune as a “full-stack” approach. The company says enterprise customers can operate systems in their own data centers and build custom voice and video models from the ground up, giving clients more direct control over where data sits and how models run in sensitive settings.

Smola has supported open-weight and customizable models, according to Fortune, putting him near Meta and other open-source advocates on that issue. Boson, however, is also building proprietary models because it is selling to business customers.

The company has raised $70 million, Fortune reported. Its backers include Chinese entrepreneur Su Hua and the venture arm of Singapore-based investment firm Temasek. Smola is first aiming at customers in finance, telecommunications, healthcare and insurance.

Customer service and sales are early targets. “A lot of machine interaction for customer support and sales… will become automated with voice agents,” Smola told Fortune, adding that the agents can be “strictly superior” to humans.

Why are OpenAI and Meta racing in voice AI?

Voice AI has become a major contest because full-duplex systems allow people to hold more natural conversations with software. Full-duplex audio lets a user interrupt an AI mid-sentence, change direction and continue speaking without waiting for a rigid turn-taking pattern.

OpenAI Chief Executive Sam Altman recently said he speaks with ChatGPT more than he texts with it and wrote that the company’s new voice model had crossed a threshold, according to Fortune. OpenAI has introduced GPT-Live, a set of full-duplex audio models.

Microsoft recently announced a newer version of its own voice model, while Meta has introduced Muse Spark, which Meta says supports natural conversation, interruptions, topic changes and language switches. Fortune reported that Microsoft is emphasizing workplace productivity, OpenAI is courting app developers, and Meta is linking voice AI to wearable hardware.

Boson still faces the technical limits that make voice systems difficult to ship. Fortune reported that full-duplex audio can raise graphics processing unit costs, and Smola identified latency as a major barrier because a delay that feels acceptable in text can feel disruptive in speech.

Smola said Boson is training its AI to read elements of human speech such as emotional tone, cheerfulness, passive-aggressiveness and rapid talking. He also told Fortune that the longer-term goal is embodied AI: robots that combine reasoning, vision and voice, with practical enterprise uses first, such as sales-call records that preserve context in real time.

This story draws on original reporting from Fortune.