Skip to main content
  1. Home
  2. Emerging Tech
  3. News

This humanlike synthesized speech could be the future of audiobooks

Add as a preferred source on Google

Synthesized voices like those used by Siri and Alexa are fine for telling us the day’s weather forecast or how many minutes remain on a cooking timer, but would you really want their flat, monotonous tones reading you audiobooks? Probably not, which is why most of us turn to human-voiced services like Audible to get our audiobook fix. Human voice actors might not get the nod for too much longer, however, due to to the pioneering work of a London-based startup called DeepZen.

Using artificial intelligence algorithms, augmented by the technological firepower of IBM’s Power A.I. and Watson technologies, DeepZen has developed text-to-speech tools that not only sound human at first listen, but can also pick up on the emotional cues needed for reading text in a compelling manner. In doing so, the company claims that it could reduce the time and cost to produce audiobooks by up to 90%.

Recommended Videos

“Our system is truly revolutionary,” Taylan Kamis, CEO and co-founder of DeepZen, told Digital Trends. “It works using deep learning and neural networks to understand how a human talks and reads. We then train the system so it can recognize where to apply the right emotions and intonation when reading a piece of text. The result is humanlike speech very closely resembling the real thing.”

Inevitably, work like this can be cast as yet another example of cutting-edge A.I. tools threatening a human profession. In this case, that profession involves actors who, despite what a few high-profile figures are able to achieve, don’t have the most steady, stable careers as it is. It would be naive to think that software such as this won’t have an impact on the future of voice actors, but, as Kamis points out, there are plenty of scenarios in which tools such as DeepZen’s could be a net positive for humanity.

For example, it could make possible the creation of audiobooks based on works by new and emerging writers, or from publishers who don’t have the luxury of big budgets. It could also be used to help develop superior text-to-speech tools for people who have dyslexia or otherwise have trouble reading.

“As for the future, we are also looking at producing voice-overs for the video production industry, as well as gaming, where there is a need for real-time text-to-speech to enhance the player experience,” Kami said. “We are also looking at other languages.”

You can check out a sample of the system here.

Luke Dormehl
I'm a UK-based tech writer covering Cool Tech at Digital Trends. I've also written for Fast Company, Wired, the Guardian…
Google Flow Music Spaces is becoming an all-in-one AI music studio
The latest update brings song editing, stem splitting, image generation, lyrics, and video creation into a single workspace.
Google Flow Music Featured

The company has announced a major update for Flow Music Spaces, bringing a suite of new AI-powered creative tools directly into collaborative workspaces. Instead of jumping between different apps or sessions, users can now generate and edit songs, split audio stems, create images, write lyrics, and even generate videos without leaving a Space.

What's new in Flow Music Spaces?

Read more
A Florida pastor asked ChatGPT if he was okay. It nearly got him killed.
When "just relax and recover" turns out to be the worst medical advice you could get.
ChatGPT on iPhone

Scott Winters, a pastor from Florida, is suing OpenAI after ChatGPT allegedly gave him what his lawsuit calls "extremely dangerous" medical advice, advice that nearly cost him his life.

What happened to the pastor?

Read more
Scientists made robots curious like toddlers, and it helped them learn language twice as fast
Give a robot curiosity, and it starts learning language (and misbehaving) just like a toddler.
robots and kids photo

Scientists have been trying to figure out how kids pick up language so fast for decades, and a new study out of the Okinawa Institute of Science and Technology (OIST) might have cracked part of the puzzle: curiosity.

Researchers built a virtual robot with a brain-inspired neural network and set it loose in a simulated 3D world full of shapes, colors, and simple commands like "push left magenta dumbbell." 

Read more