Three former Apple researchers have secured a fifty million dollar Series A to build an audiovisual foundation model that ingests video and audio together and emits both in a single pass, bypassing the stitching of speech-to-text, language model, and text-to-speech that gives today’s avatars their characteristic lag. Lightspeed Venture Partners led the round. Accel, which led the ten million dollar seed last July, returned alongside Nvidia’s NVentures, South Park Commons, and Define Ventures. The company, founded in 2025 and based in Seattle, has eight employees and no product yet.

The architecture bet

Chief executive Fangchang Ma argues that the pipeline approach, transcribe, reason, synthesize, cannot capture the micro-expressions and back-channel cues that make human conversation feel natural. Nuance’s alternative is a unified model with audio and video inputs and audio and video outputs. Ma says the training data cannot be scraped; the startup records consented, dual-camera conversations with separated audio tracks to capture unscripted interaction. That data strategy is expensive and slow, which is why the round exists.

Capital structure and leverage

The Series A brings total raised to sixty million. Lightspeed takes the lead position; Accel’s follow-on signals confidence but also protects its seed stake. Nvidia’s participation through NVentures is the most telling signal: the chipmaker rarely backs pre-product model labs unless the compute profile aligns with its roadmap. No break fees, no earnouts, no disclosed valuation, the term sheet is clean, which at this stage usually means the founders kept control of the research agenda.

Go-to-market comes next

Ma says the next hires will be go-to-market specialists to define pricing for enterprise use cases: AI interviewers, sales agents, language tutors, customer support avatars. A research preview is promised before year end. Until then, the company is a compute-intensive research project with a venture budget and a hypothesis that the market will pay a premium for latency that feels human. The money buys time to prove it.