Skip to main content
  1. Home
  2. Computing
  3. News

AI voice chats still feel awkward because assistants don’t know when to talk

Thinking Machines Lab is testing faster full duplex AI that can listen and respond at the same time

Add as a preferred source on Google
Electronics, Mobile Phone, Phone
OpenAI

Thinking Machines Lab says it’s building full duplex AI, which means an AI system can take in what someone is saying while generating a response. In plain English, it’s closer to a phone call than a walkie-talkie.

The startup, founded last year by former OpenAI CTO Mira Murati, announced interaction models, starting with TML-Interaction-Small. It says the system can respond in 0.40 seconds, a pace that puts it near ordinary human back-and-forth.

Recommended Videos

There’s a catch for anyone hoping to try it today. This remains a research preview, with limited access planned in the next few months and a broader release expected later this year.

A faster kind of AI exchange

The core idea is easy to understand, and the change is meaningful. Instead of waiting for someone to finish speaking before working on an answer, the model processes incoming speech while preparing its response.

That delay matters because pauses make AI assistants sound artificial. Thinking Machines Lab frames TML-Interaction-Small’s 0.40-second response time as close to natural conversation speed, which would be a noticeable shift for voice tools.

It also claims that pace is faster than comparable models from OpenAI and Google. The benchmark gives the announcement weight, but outside users still need to test whether the experience works as smoothly as the number suggests.

When speed becomes behavior

An assistant that answers while it’s still taking in information changes what users expect from a voice chat. The conversation can move faster, but the system also has to manage timing with much more care.

That tradeoff matters when someone wants quick clarification instead of a long generated reply. Faster responses won’t help much if the assistant jumps in too early, misunderstands the speaker, or breaks the flow it’s supposed to improve.

For now, the architecture is the news. The real product test is whether the interaction model can make better timing feel automatic.

What to watch before launch

The release timeline is the key detail now. Thinking Machines Lab says a limited research preview is coming in the next few months, followed by broader access later this year.

Availability, pricing, supported platforms, and performance outside controlled testing are still unclear. Those missing pieces matter because a faster model only helps if people can use it in everyday voice tools.

For anyone who uses AI voice assistants, the practical move is to watch the preview closely. Full duplex AI has promise, but hands-on testing should show whether faster responses actually make daily AI conversations easier.

Paulo Vargas
Paulo Vargas is an English major turned reporter turned technical writer, with a career that has always circled back to…
The Mac Pro nearly received an M3 Extreme chip twice as powerful as M3 Ultra
High production costs likely killed Apple’s M3 Extreme plans
Apple's Mac Pro on a table at a press event.

Apple discontinued the Mac Pro earlier this year, ending a 20-year run for a computer that once represented the very best of the company’s desktop lineup. However, Apple reportedly had much bigger plans for the machine before ultimately replacing it with the Mac Studio.

According to Bloomberg’s Mark Gurman, Apple developed an M3 Extreme chip that could have offered twice as many CPU and GPU cores as the M3 Ultra. The processor was intended to sit above the Ultra tier and could have finally given the Mac Pro the performance advantage it badly needed. Apple eventually abandoned the chip due to concerns over production costs and limited demand for such an expensive machine.

Read more
Hidden prompts can secretly rewrite an AI’s memory, and researchers say that’s a serious problem
Researchers discover AI attack that rewrites an assistant's long-term memory
Chatbot on a smartphone.

Large language models are getting better at remembering us. Whether it's your preferred writing style, recurring tasks, shopping habits or project deadlines, AI assistants are increasingly storing long-term memories to make future conversations feel more personal and useful. But according to new research, that same feature could become one of AI's biggest security vulnerabilities.

Researchers from New Mexico State University have demonstrated a new attack called GhostWriter, capable of secretly planting false memories inside AI agents. Rather than stealing information outright, the attack manipulates what an AI remembers, potentially causing it to make dangerous decisions long after the original attack has taken place.

Read more
This experiment shows how easy it is to poison an open-weight AI model for under $100
This research raises new doubts about trusting open weight AI models.
Computer, Electronics, Laptop

Open-weight AI models have been having a moment lately. Just this month, Moonshot's massive Kimi K3 model landed close behind Claude Fable 5 and GPT 5.6 Sol in several benchmarks, all while remaining fully open-weight and downloadable by anyone.

However, Katie Paxton-Fear, a cybersecurity lecturer at Manchester Metropolitan University and staff security advocate at Semgrep, managed to poison an open-weight model and proved how easily that openness can be turned against you (via The Register).

Read more