Nuance raises $50 million for full-duplex audiovisual AI

By
CTOL Staff Reporter
1 min read

Nuance, a Seattle startup founded by three former Apple researchers, disclosed a $50 million Series A on September 14 from Lightspeed Venture Partners, Accel, Nvidia's NVentures and others. With a prior $10 million seed round, the company has raised about $60 million before launching a commercial product.

Nuance is developing a model architecture rather than an application for a specific industry. Most real-time voice agents split the interaction into several systems: speech recognition converts audio to text, a language model generates the response, and speech synthesis turns it back into audio. Nuance is building one full-duplex audiovisual model intended to see, hear, reason, speak and react at the same time.

Combining those functions can reduce handoff delays and preserve cues a modular pipeline may discard, including interruption timing, facial expression, eye contact and back-channel responses. Those cues help conversation feel continuous. Nuance describes roughly 500 milliseconds as the hard floor for natural conversational latency. The company has not published a production benchmark demonstrating that design target.

Fewer handoffs may come with higher compute costs

The founders have relevant systems experience. Chief executive Fangchang Ma worked on Apple projects including ARKit depth, Vision Pro depth and iPhone imaging; co-founders Edward Zhang and Karren Yang came from graphics and audiovisual research roles. Nuance says the team has shipped low-latency machine-learning products used by millions of people.

Nuance has not published a head-to-head test against a strong modular voice stack on end-to-end latency, interruption handling, user preference or inference cost. It also has no disclosed revenue, paid-customer cohort, retention data or production latency distribution under load. Those omissions leave its commercial advantage untested.

An integrated model avoids serial handoffs and can keep audiovisual context in one learned system. A modular stack, however, can assign speech recognition, reasoning and synthesis to specialized models, scale those components separately and replace a costly layer without retraining the whole product. A full-duplex model that continuously processes video and audio may deliver a better conversation while consuming more compute per minute.

Nvidia's participation fits that tradeoff: better real-time interaction could create a demanding inference workload. The investment does not establish that the workload is cheaper, that users prefer it or that the product will support software-like gross margins.

The $60 million gives Nuance funding to test whether better interaction can justify continuous audiovisual inference. A research preview can show how the conversation feels. Production latency, usage and serving costs will determine whether that improvement can sustain a business; the public data do not yet answer that.

Sources

You May Also Like

This article is submitted by our user under the News Submission Rules and Guidelines. The cover photo is computer generated art for illustrative purposes only; not indicative of factual content. If you believe this article infringes upon copyright rights, please do not hesitate to report it by sending an email to us. Your vigilance and cooperation are invaluable in helping us maintain a respectful and legally compliant community.

Subscribe to our Newsletter

Get the latest in enterprise business and tech with exclusive peeks at our new offerings

We use cookies on our website to enable certain functions, to provide more relevant information to you and to optimize your experience on our website. Further information can be found in our Privacy Policy and our Terms of Service . Mandatory information can be found in the legal notice