Skip to main content
  1. Home
  2. Computing
  3. Emerging Tech
  4. News

Programmer trains artificial intelligence to draw faces from text descriptions

Add as a preferred source on Google
T2F training time lapse

Programmer Animesh Karnewar wanted to know how characters described in books would appear in reality, so he turned to artificial intelligence to see if it could properly render these fictional people. Called T2F, the research project uses a generative adversarial network (GAN) to encode text and synthesize facial images.

Recommended Videos

Simply put, a GAN consists of two neural networks that argue with each other to produce the best results. For example, the job of network No. 1 is to fool network No. 2 into believing a rendered image is a real photograph while network No. 2 sets out to prove the alleged photo is just a rendered image. This back-and-forth process fine-tunes the rendering process until network No. 2 is eventually fooled.

Karnewar started the project using a dataset called Face2Text provided by researchers at the University of Copenhagen, which contains natural language descriptions for 400 random images.

“The descriptions are cleaned to remove reluctant and irrelevant captions provided for the people in the images,” he writes. “Some of the descriptions not only describe the facial features, but also provide some implied information from the pictures.”

While the results stemming from Karnewar’s T2F project aren’t exactly photorealistic, it’s a start. The video embedded above shows a time-lapsed view of how the GAN was trained to render illustrations from text, starting with solid blocks of color and ending with rough but identifiable pixilated renderings.

“I found that the generated samples at higher resolutions (32 x 32 and 64 x 64) has more background noise compared to the samples generated at lower resolutions,” Karnewar explains. “I perceive it due to the insufficient amount of data (only 400 images).”

The technique used to train the adversarial networks is called “Progressive Growing of GANs,” which improves quality and stability over time. As the video shows, the image generator starts from an extremely low resolution. New layers are slowly introduced into the model, increasing the details as the training progresses over time.

“The Progressive Growing of GANs is a phenomenal technique for training GANs faster and in a more stable manner,” he adds. “This can be coupled with various novel contributions from other papers.”

Image used with permission by copyright holder

In a provided example, the text description illustrates a woman in her late 20s with long brown hair swiped over to one side, gentle facial features and no make-up. She’s “casual” and “relaxed.” Another description illustrates a man in his 40s with an elongated face, a prominent nose, brown eyes, a receding hairline and a short mustache. Although the end results are extremely pixelated, the final renders show great progress in how A.I. can generate faces from scratch.

Karnewar says he plans to scale out the project to integrate additional datasets such as Flicker8K and Coco captions. Eventually, T2F could be used in the law enforcement field to identify victims and/or criminals based on text descriptions, among other applications. He’s open to suggestions and contributions to the project.

To access the code and contribute, head to Karnewar’s repository on Github here.

Kevin Parrish
Kevin started taking PCs apart in the 90s when Quake was on the way and his PC lacked the required components. Since then…
macOS 27 is finally here with Apple’s biggest AI upgrades yet
Apple's latest Mac update brings a smarter Siri, deeper Apple Intelligence integration, and a host of useful improvements to everyday macOS tasks.
macOS 27 on a Mac on black background

Apple has officially released macOS 27 Golden Gate, bringing the latest generation of Apple Intelligence to the Mac alongside a collection of performance, search, connectivity, and usability improvements.

The biggest addition in today's release is Siri AI, which Apple describes as an entirely new version of the assistant. On Mac, Siri can understand personal context, answer questions about what's on your screen, take systemwide actions, and pull information from the web. It can also work with information from messages, emails, photos, and other personal content to provide more relevant answers.

Read more
Microsoft’s new AI rulebook puts human control above everything else
Microsoft's drat Humanist AI Code of Conduct lays out strict limits for its MAI models, from resisting shutdowns to claiming consciousness, as the company calls for a safer approach to frontier AI.
Person, Walking, Adult

Microsoft today published a draft "Humanist AI Code of Conduct," a rulebook meant to guide how its MAI models behave, and it opens with a blunt promise: People matter more than AI. The document lands just days after Anthropic CEO Dario Amodei called on the industry to slow down frontier AI development.

What the code requires

Read more
I didn’t know I needed an AI voice recorder until I tried the Comulytic Note Pro
I finally tried an AI voice recorder, and the Comulytic Note Pro makes a strong case
Comulytic Note Pro

I’ll admit that I was skeptical when I first picked up the Comulytic Note Pro. This is my first time using a dedicated AI note-taking device, so I came into it with a fairly simple question: why should I carry another gadget when my phone can already record audio and an AI chatbot can summarize it?

After spending time understanding how the Note Pro fits into that workflow, I think the answer is convenience. It takes the recording part out of my phone, gives me a physical button to start capturing audio, and then turns those recordings into transcripts, summaries, highlights,s and action points through its app. That sounds like a small change, but it addresses the annoying part of note-taking: trying to listen, remember, type, and participate in a conversation at the same time.

Read more