Two Undergrads Built An AI Speech Model to Rival Notebooklm
A Pair of Undergrads, Neiter with Extensive AI Expertise, Say that they've Created an Openly Avaılable Hunting Model That
Two Undergrads Built An AI Speech Model to Rival Notebooklm
A Pair of Undergrads, Neiter with Extensive AI Expertise, Say that they’ve Created an Openly Avaılable Hunting Model That Can Generarate Podcast-Style Clips Similar to Google’s Notebooklm.
The Market for synthetic Speech Tools is vast and Growing. Elevenlabs is one of the Largest Players, But The No Shortage of Challengers (See Playai, Sesame, and So On). Investors Believe that These Tools have immense potental. According to Pitchbook, Startups Developing Voice AI TECHISED OVER $ 398 Million in vc Funding Last Year.
Toby Kim, One of the Korea-Based Co-Founders of Nari Labs, The Group Behind The Newly Releaced Model, Said That He and His Fellow Co-Founder Started Learning About Speech AI Three Moths ago. Inspired by notebooklm, they wated to create a model that of the Public Oven Generated and “Freedom in the script.”
WHO USED Google’s TPU Research Cloud program, which provides researchers with the Company’s TPU AI Chips, to Train Nari’s Model, DIA. Weighing in at 1.6 Billion Parameters, Dia Can Generate Dialogue from A Script, Letting Users Customize Speakers’ Tones and Insert Disfluencies, Coughs, Laughs, and Oh Nonverbal Cues.
Parameters are the internal variables models. Generally, Models with More Parameters Perform Better.
Availblet from the AI Giant Platform Huging Face and Github, DIA Can Run on Most Contemporary Pcs with At Least 10GB of Vram. It Generates A Random Voice Unless Prompted with a Description of An Inteneded Style, BUT I Also Clone A Person’s Voice.
In TechCrunch’s Broup Testing of Dia Through Nari’s Web Demo, Dia Worked Quite Well, Uncomplaingly Generationing Two-Way Chats About Any Subject. The Quality of the Voices See Competitive With the Tools, and the Voice Cloning Function is the Easiut this Reporter Has Tried.
Here’s A sample:
Like Many Voice Generators, Dia Offers Little in the Way of Safavuards, Howver. It’d be Trivially Easy to Craft Disinformation or A SCAMMY Recording. On Dia’s Project Pages, Nari Discourage Abuse of the Model to impersonate, Deceive, Or Ophise Engage in Illicit Campaigns, But The Group Says it “Isn’t Responsible” for Misuse.
Nari Also Hasn’t Discloseed Who Data it Scraped to Train DIA. Its Possible Dia Was Developed Using Copyrighted Content Training Models on Copyright. SOME AI Companies Claim That Fair Use Shields Them From Liability, Who Rights Holders Holders Holders Holders Assert That Fair Use Doesn Apply to Training.
In Any Event, Kim Says Nari’s Plan is to Create a SyntHetic Voice Platform Nari Also Intends to Release a Technical Report for Dia, and to Expand The Model’s Support to Languages Beyond Englishhh.