Meta presents Voicebox, the synthetic voice revolution that will not be at your disposal


Meta Platforms, the parent company of Facebook, has presented Voicebox, its generative model of Artificial Intelligence for convert text to speech. Although still a research project, Meta says Voicebox can generate speech in six languages ​​from samples of just two seconds, and could in the future be used for “natural and authentic” translations, among other applications.

Voicebox is a significant advance in the generation of speech from text, since it can perform tasks without the need for specific training for each one. However, due to possible misuse risksMeta is not making the Voicebox model or code publicly available at this time.

Voicebox: the synthetic voice revolution that Meta doesn’t want you to use

Meta recently unveiled LLaMA, the new ChatGPT. Now, he has shared an overview of his revolutionary Artificial Intelligence system called “Voicebox”, which gives users the ability to convert text to audio with a wide range of styles and voices. This innovative technology will translate written words into captivating soundsoffering a unique and personalized hearing experience.

The model is capable of producing high-quality audio clips and editing pre-recorded audio, removing unwanted noise such as car horns or barking dogs, while maintaining the content and style of the audio. Also, Voicebox is multilingual and offers exceptional results compared to other similar models.


Although the generated voices are not perfect and they sound more like reading a book than casual conversation, are very realistic and promise great potential in various areas. The announcement of Voicebox by Meta researchers highlights its unique capability and superiority in this area. Meta hopes that the model will evolve over time and can acquire a more natural intonation.

Meta says its new AI model is too dangerous for public release

Although Meta has shared information about Voicebox, in order to maintain transparency in the advances in this field, does not plan to make it open source. The company recognizes the importance of further research in a responsible manner and prefers not to make the model publicly available at this time, given concerns regarding potential misuse.

In the future, this technology promises beneficial applicationssuch as allowing visually impaired people to hear messages written in their own voices. However, Meta is aware of the concerns about the creation of deepfakes and has made important decisions to address these risks. She maintains a responsible approach in her research in Artificial Intelligence.

While some people have used similar tools online to create synthesized voice clips of their favorite characters as a fun form of entertainment, others have abused them online. harassment campaigns aimed at dubbing actors. Therefore, Meta might be taking precautions to avoid potential harm in this regard.

Currently, imitating the voice using Artificial Intelligence is already possible. Now, this new tool to convert text to speech could generate exceptional results In the not too distant future. However, ethical concerns and the risks of misuse are latent.


Related News

Xiaomi Memory extension: What is it and why should you have it activated or deactivated on your mobile

With the increase in users every year using smartphones for all kinds of tasks, manufacturers have had to improve their features. With this increase, they

Interview with Atsushi Ohkubo at Lucca Comics and Games 2022

Atsushi Ohkubo is among the guests of Lucca Comics and Games 2022 that we had the opportunity to interview during the kermesse, telling us something more

How to play the new Halloween-themed Google doodle

As expected, Google has recently launched its new doodle, this time in version Halloween 2022so that its users can enjoy a fun game with a "Halloween Day"

Moto G Play 2022 is almost here as new leak sheds light on its design and hardware

The successor to last year's Moto G Play will have some upgrade options

What happened to mobile phones with a 3D screen?

At the beginning of the last decade there was a boom for 3D in films caused at the end of 2009 with the premiere of 'Avatar', whose sequel, 'The Sense of

Nothing Ear Stick vs Nothing Ear (1): More Than Just a Different Charging Case

With little time on the market, the brand Nothing has managed to generate an impact with the launch of each of its products. The Nothing Ear (1) arrived with