Meta presents Voicebox, the synthetic voice revolution that will not be at your disposal


Meta Platforms, the parent company of Facebook, has presented Voicebox, its generative model of Artificial Intelligence for convert text to speech. Although still a research project, Meta says Voicebox can generate speech in six languages ​​from samples of just two seconds, and could in the future be used for “natural and authentic” translations, among other applications.

Voicebox is a significant advance in the generation of speech from text, since it can perform tasks without the need for specific training for each one. However, due to possible misuse risksMeta is not making the Voicebox model or code publicly available at this time.

Voicebox: the synthetic voice revolution that Meta doesn’t want you to use

Meta recently unveiled LLaMA, the new ChatGPT. Now, he has shared an overview of his revolutionary Artificial Intelligence system called “Voicebox”, which gives users the ability to convert text to audio with a wide range of styles and voices. This innovative technology will translate written words into captivating soundsoffering a unique and personalized hearing experience.

The model is capable of producing high-quality audio clips and editing pre-recorded audio, removing unwanted noise such as car horns or barking dogs, while maintaining the content and style of the audio. Also, Voicebox is multilingual and offers exceptional results compared to other similar models.


Although the generated voices are not perfect and they sound more like reading a book than casual conversation, are very realistic and promise great potential in various areas. The announcement of Voicebox by Meta researchers highlights its unique capability and superiority in this area. Meta hopes that the model will evolve over time and can acquire a more natural intonation.

Meta says its new AI model is too dangerous for public release

Although Meta has shared information about Voicebox, in order to maintain transparency in the advances in this field, does not plan to make it open source. The company recognizes the importance of further research in a responsible manner and prefers not to make the model publicly available at this time, given concerns regarding potential misuse.

In the future, this technology promises beneficial applicationssuch as allowing visually impaired people to hear messages written in their own voices. However, Meta is aware of the concerns about the creation of deepfakes and has made important decisions to address these risks. She maintains a responsible approach in her research in Artificial Intelligence.

While some people have used similar tools online to create synthesized voice clips of their favorite characters as a fun form of entertainment, others have abused them online. harassment campaigns aimed at dubbing actors. Therefore, Meta might be taking precautions to avoid potential harm in this regard.

Currently, imitating the voice using Artificial Intelligence is already possible. Now, this new tool to convert text to speech could generate exceptional results In the not too distant future. However, ethical concerns and the risks of misuse are latent.


Related News

Alienware Aurora R15 現在は Nvidia RTX 4090、Intel 第 13 世代。 が付属しています

Alienware は、一部の PC ゲーマーを確実に喜ばせるゲームに焦点を当てた一連のデバイスを発表しています。 Aurora R15 は、以下を含む強力なゲーミング デスクトップです。

Fire TV Stickのスリープモードをオフにして、デバイスを常にオンにしておく方法

Amazon の Fire TV Stick デバイスは、テレビのスマート TV 機能を拡張する優れた代替手段です。 彼らと一緒に、私たちはまったく異なるテレビを提供します

Nintendo Directで発表されたピクミン4

Nintendo Direct の期間中、非常に大規模な予告編が連続して見られました。 これらの 4 つは、2023 年中に Nintendo Switch に登場するピクミン XNUMX を示していました。その後

上位 XNUMX つの SmartStart Young Innovator チームがインキュベーション段階に進みます

ブートキャンプ、「Hatch」と「Digithon」のチャレンジ アクティビティを経て、次のフェーズに進む XNUMX つのチームが選ばれました。