Why Aren't We Using Her?

An exploration of consumer voice AI in mid-2026

By Charles Niu | July 21, 2026

Theodore from Her scratching his head beside the movie's lowercase logo: her?

In 2026, you can’t throw a rock in the Valley without hitting five startups claiming they’re building Her. It’s become the “holy grail” of voice AI.

Crazy thing is? We kinda do have Her now.

Sam Altman's one-word tweet: "her" - 21.7M views
Remember this tweet? This was two years ago.

Over the past few years, voice AI has become astonishingly good. GPT-Live-1, Gemini Live, Sesame, and Thinking Machines’ upcoming interaction model all promise some combination of human-like prosody, emotional expressiveness, minimal latency, multimodal understanding, and tool use.

For just $20/month (less than a day’s DoorDash order) you get a voice AI you can confide in, flirt with, and even use to “manage your emails.”

And yet, for all this progress, how many people do you know who actually talk to an AI every day?

I’m a Silicon Valley bubblite through and through and even I can’t think of a single person who actually uses voice AI. I myself might use it on the rare occasion when I want to ask a question while I’m commuting, but…that’s basically it.

This creates a strange disconnect. The technology keeps improving dramatically, but consumer behavior barely seems to move.

So what’s happening?

  • Are consumers simply slow to adopt?
  • Is the tech still missing something?
  • Or have we accidentally mistaken good sci-fi for something actually useful?

We have Her. Why aren’t we using Her?

1Consumers Just Need Time to Catch Up

Let’s begin with the most optimistic case: the tech is here, but consumer behavior just takes a while to adjust.

The iPhone was a transformative breakthrough, but even its adoption was hardly instantaneous. Apple sold roughly 1.4 million iPhones in its first year, but it took two years before exponential growth really kicked off. For voice AI, the really good stuff has only just started to emerge this year, so we may still be in that initial lull.

Chart of worldwide iPhone unit sales from 2007 to 2012, with an annotation on the flat early portion of the curve reading 'voice ai might still be here'
When the iPhone first came out, it was more of a novelty or status symbol. Now it's the norm

If we’re looking at a Her-like experience more broadly, however, there may already be signs of people “living in the future.”

AI Companions

AI companion apps like Character AI (and its successors) have already attracted millions of users wanting to chat with their favorite anime characters. Some users spend hours a day with them and develop surprisingly deep emotional attachments.

Concept showing Character AI voice calls on three smartphones
It’s been 2 years since Character AI added voice chat, what would it look like if it released again today?

But there is an important distinction: voice was not necessary to make these products successful.

In fact, the rollout of voice features in these companion apps came relatively uncelebrated, and most AI companion power users still prefer text to this day.

The argument could be made that voice features for these companion apps are “tacked on.” Most AI companion apps aren’t “dialogue” experiences, but rather “roleplay” experiences with descriptive narratives that sound awkward when voiced aloud. So what about an AI companion that’s voice-first?

Enter Sesame

Founded by ex-Oculus leads, Sesame made headlines last year for showing off their unusually natural-sounding voice agent. Rather than relying on a radically larger foundation model, the team focused on careful engineering, high-quality conversational data, speech fillers, and expressive prosody.

The result felt uncannily close to speaking with Her (or Him - their male persona Miles is arguably more popular).

A hand holding a phone during a voice call with Sesame's Maya persona
The number 1 request from the community is, unsurprisingly, uncensored models.

More recently, Sesame released iOS and Android apps with additional personalities, tool use, research capabilities, and other assistant features. The experience is shockingly good (to the point where I was very confused why the app release didn’t go immediately viral). The model laughs realistically, hesitates, and even sounds like it’s genuinely thinking.

Sesame: Personal Agents App Store listing - 'Crafted for conversation,' with screenshots of the Maya and Simone personas
Sesame is worth checking out if you have any interest in voice AI.

And I still stopped using it after just three days.

As impressive as the tech is, it only took a few calls before the conversation patterns started feeling repetitive. The speech fillers became predictable, the affectations got annoying, and the illusion fell apart - at which point it felt like I was talking to a strictly less efficient ChatGPT.

Then Came GPT-Live-1

OpenAI’s latest flagship voice model, released shortly afterward (less than two weeks ago at the time of writing), pushed the technical frontier even further. GPT-Live-1 brings together capabilities like:

  • Natural interruption and backchanneling
  • Multimodal inputs
  • Real-time translation
  • Strong general intelligence
  • Low conversational latency

But perhaps most impressively, OpenAI rolled it out broadly across all of ChatGPT’s paid tiers rather than treating it as a niche research preview (free users received a lighter-weight version). This made duplex voice available to an enormous consumer audience of over 50 million people without requiring a separate product or subscription.

Three grandmothers in the GPT-Live-1 launch video beneath the words 'The all-new ChatGPT Voice'
The launch video of GPT-Live-1 prominently featured grandmas, an indication of their target consumer profile.

This rollout creates an interesting experiment: if accessibility and technical quality were the primary barriers, voice usage should now begin to grow dramatically.

It may be too early to tell just yet, but in my own case, I stopped using GPT-Live-1 after just one day (even faster than Sesame!). As much as I admired the engineering, I simply could not find enough moments in my life when I wanted to start talking with an AI.

So what if having a conversation with Her wasn’t an explicit decision, but rather a passive, always-on experience?

The Wearable Bet

Smartphones add a surprising amount of friction to voice AI (just look at how Siri’s doing). Starting a conversation is often more cumbersome than just typing. You have to take out your phone, open the app, enter voice mode, and then continue holding the device while you talk.

Over the past two years, new waves of AI wearables have attempted to remove this friction by making the assistant continuously available.

Humane’s AI Pin was one of the earliest large-scale players in this space. Their device was designed as a voice-first alternative to the smartphone, allowing users to make calls, send messages, ask questions, and even project a small interface onto their hand. But it was slow, expensive, unreliable, and worst of all - extraneous. It felt like it was solving a problem people didn’t really have. By February 2025, Humane’s assets were acquired by HP, and its users were left with a $700 brick.

The clearest commercial success in this category so far is the Meta Ray-Bans. There are a few reasons Meta’s approach appears to be working:

  1. The product uses a familiar accessory and a comfortable, socially acceptable form factor.
  2. Meta has positioned their glasses as a fashion product, supported by Ray-Ban’s brand as well as celebrity ambassadors.
  3. Even without AI, the glasses are already useful for recording things and listening to music.

Voice AI is only one layer of the Meta Ray-Ban experience. Users can ask Meta AI questions about what they are seeing or issue simple commands, but the product does not depend on consumers wanting to maintain a prolonged relationship with an AI companion.

Kylie Jenner wearing Meta Starfire glasses beside a product image of the glasses
The recent release of Starfire features uses Kylie Jenner’s voice for Meta AI.

In other words, Meta’s strategy appears to be hardware first, voice second.

Interestingly, Sesame is also investing in building their own glasses, but they seem to be making the opposite bet: build the voice AI first, then eventually give it access to glasses, cameras, and the rest of the user’s environment.

So far, neither approach has demonstrated that consumers broadly want to spend their days in open-ended conversation with a Her. So is it simply a matter of time? Or is there still something fundamentally missing on the tech side?

2We Don’t Actually Have Her Yet

The second perspective we’ll look at is more skeptical: not only do we not have Her, but depending on who you ask, we might not even be close.

Samantha (the AI in the movie) is not just a pretty voice with good prosody. She knows when to speak and when to stay quiet. She adapts to subtle emotional shifts, and she understands what Theodore (the movie’s protagonist) is physically doing.

Stills from Her (2013): Samantha's handheld device showing a call from Samantha, and Theodore's desk with her desktop screen
Samantha appears in the movie as a desktop and phone-like OS

Current voice products may be able to reproduce parts of Her - at least on a surface level - but several deeper capabilities remain incomplete. Three gaps stand out in particular:

i. Full-duplex conversation

Most voice agents have historically relied on discrete turn-taking:

You speak → The system detects that you stopped → It thinks → Then it replies.

The interaction takes place over one conversational channel, like a walkie-talkie (what we call half-duplex).

Human conversation is messier: we interrupt, we overlap, we say “mhm,” “right,” or “I see” all while someone else is speaking. In other words, we can both listen and think while also talking. Unlike a walkie-talkie, full-duplex is more like a conversation you’d have over the phone.

In 2024, Kyutai’s Moshi demonstrated one of the first real-time, full-duplex speech models and released it openly (our team did an in-depth breakdown of the paper in our previous blog post). Since then, nearly every major voice lab has begun exploring architectures that can listen and speak at the same time.

Thinking Machines Lab introducing interaction models in a full-duplex voice demonstration
Thinking Machines made headlines a couple months ago showing off their impressive full-duplex interaction models.

OpenAI’s GPT-Live-1 is perhaps the first full-duplex voice model by a major foundation lab out the gate, but it’s not going to be the only one. Google, Amazon, Thinking Machines, and others are pursuing similar directions. 2026 may be the year full-duplex voice finally becomes widely available.

Despite these exciting advancements, it’s worth calling out that full-duplex isn’t a silver bullet.

It currently comes with trade-offs in intelligence, tool-calling accuracy, and increased inference cost. The question on everyone’s mind right now is: are the benefits from full-duplex truly worth the cost? Or are there other directions that are more worth pushing on?

ii. Understanding emotion

Many voice-generation services today (i.e. ElevenLabs) offer “emotion tags” that instruct a text-to-speech (TTS) model to sound happy, sad, excited, frustrated, etc. Sounding emotional, however, is not the same as understanding emotion.

This field is called speech emotion recognition (SER). For much of its history, researchers treated it as a classification task. A model would listen to a recording and assign it one of a handful of labels: happy, sad, angry, fearful, etc.

But human emotion does not always fit neatly into discrete categories. A person can sound hesitant without being afraid, sarcastic without being angry, etc.

More recent models such as emotion2vec and emotion2vec+ (2024) therefore learn continuous latent representations from audio. Rather than immediately reducing speech to a single label, they convert it into an embedding that preserves emotion-relevant acoustic information such as pitch, rhythm, speaking rate, emphasis, hesitation, and vocal tension.

Still, tone alone is not enough. Sarcasm, for instance, may depend on the contrast between someone’s words and how they say them. A raised voice could signal anger, or it could just be a noisy room - it’s hard to tell without taking in additional context clues.

The frontier is therefore moving toward systems that combine:

  1. What was said (language)
  2. How it was said (tone)
  3. The surrounding conversation (context)

Recent systems such as C²SER (2025), for example, combine a speech encoder that captures linguistic meaning with an emotion-specialized encoder that captures vocal delivery. This allows a language model to reason over both signals together simultaneously.

More general voice models take a similar approach. Systems such as Ultravox pass audio through a speech encoder and project the resulting representations into the language model’s embedding space. In principle, this gives the LLM access to both the words and the acoustic information that a transcript would have otherwise discarded.

In practice, the problem remains unsolved. Current speech-language models can still over-rely on specific words, miss contradictory vocal cues, or confuse emotion with accent, speaker identity, and recording conditions.

Until voice agents can consistently combine language, tone, and context, they may sound like an uncanny simulation of empathy while remaining unable to understand how the user actually feels.

iii. World context

The first two gaps concern the conversation itself. But to truly achieve Her, an AI must also understand the world in which the conversation is happening.

After all, in the movie, Samantha is more than just a voice - she’s the entire operating system. She’s able to see through Theodore’s cameras, read through his emails, and even join him in games. Their relationship feels more compelling because they are able to share the same experiences.

This involves two related capabilities:

  • World state: What’s happening right now?
  • World model: How does the environment work, and what’s going to happen next?

If we take the example of a game of Valorant, the world state would be what tells the agent that the player’s in a 1 v 5 situation, whereas the world model would help it understand what this actually means and what to say next (“gg go next” or “I’ve seen you win these” - or just shut up and let the player focus).

Research such as Yann LeCun’s Joint Embedding Predictive Architecture, or JEPA, points toward models that predict abstract representations of the future rather than every individual pixel. This may be relevant for a gaming companion, as the important fact is not the exact appearance of the next frame, but rather what’s important to the player - are they in danger, what’s happening around them, etc.

As such, games may be a natural place to create Her.

Unlike the physical world, which can only be understood by AI if reconstructed through camera, microphone, and sensor data (and even then the data will never be fully complete), games are constrained virtual worlds whose internal state is already fully available. A model in a game can directly know where the player is, which characters are nearby, what happened earlier, and what actions are possible.

We’re not the only ones who believe in games. General Intuition recently released MIRA, a world model that simulates an entire multiplayer game of Rocket League. There is no hand-written physics engine or game logic - the model learns the dynamics of the world directly from thousands of hours of play data. It’s by far the most compelling gameplay we’ve seen from a world model to date.

MIRA world model rendering a multiplayer Rocket League match from four players' perspectives
MIRA proved that you can even make a multiplayer game with world models.

MIRA does not solve conversational AI, but it does demonstrate one important ingredient: a model that can learn how a virtual world evolves.

Now imagine combining the two sides:

  • A voice model that gives the AI the ability to communicate naturally (duplex and emotion).
  • A world model that gives it the ability to understand what’s actually happening.

Instead of a chatbot, we’d have a kind of co-presence. That may be the threshold we need to cross before we can truly say we have Her.

3Voice Calling Isn’t Actually the Right Interface

The third and final case is the most bearish - what if a Her-like “voice call” is fundamentally the wrong way to think about the consumer experience?

The failure of VR offers a useful warning.

Virtual reality looks extraordinary in movies. It sounds inevitable in science fiction. And the first time you demo it, it’s genuinely mind-blowing. When I first started working in VR, I was convinced it was the future of computing.

$200 billion dollars and a decade later however, it turns out that the activation energy required for VR just makes it too cumbersome to use.

You have to put on a headset. You lose access to your phone. You need to clear a space. The device has to be charged. Your hair gets messed up. Sometimes you get motion sickness. For most things, it’s easier to just use a laptop.

Mark Zuckerberg walking past a large audience wearing virtual reality headsets
VR is amazing, but it’s also hard not to feel a little dystopian about it.

Her may have a similar problem - it looks great in a movie, but in real life voice isn’t actually a good UX.

You can’t comfortably speak to an AI about sensitive stuff in the office. You can’t ask personal questions when you’re on the subway. You can’t even skim a long response. Voice as a UI is temporally linear - you can’t compare multiple things side by side, jump forward or backward, or even just take in the information at your own pace.

It’s also worth noting that Her was just a movie trying to tell a story.

Spike Jonze (the director) wasn’t trying to lay out some grand vision for the future. Rather, narratively, it was a commentary on the present: in 2013 people were already beginning to "fall in love" with their phones through para-social relationships and social media (and arguably more true now in 2026). Samantha was never meant to be a product design proposal.

So where does that leave voice? If “a phone call with an AI” is not the endgame, when does speaking actually become useful?

Productivity Apps

We may already be seeing some answers.

  • Wispr Flow treats speech as a faster and more natural keyboard.
  • Granola listens passively and creates meeting notes in the background.

While these products are a far cry from the romantic vision presented in Her, together they represent an industry already worth roughly $20 billion. Their success points toward a really simple, grounded principle: use voice where typing is inconvenient.

Granola AI Notepad turning a live meeting transcript into organized notes
Typing notes is inconvenient, so using ASR to do it for you is a natural solution.

Car Assistants

We are already seeing this principle play out in cars.

Earlier, I mentioned using voice AI during a commute. Tesla has now begun integrating Grok Voice into its cars, while Mercedes-Benz, Rivian, and other automakers are pursuing conversational assistants of their own.

The appeal is straightforward. Texting while driving is dangerous as hell - and it’s just frankly hard to do. Voice AI as a hands-free experience therefore becomes a natural user interface here.

Gaming Companions

Gaming feels like another natural fit for voice AI.

After all, a gamer’s hands are always occupied playing the game, but they still want to coordinate and socialize - it’s no wonder Discord found such meteoric success building a voice chat platform for gamers.

What’s weird is that despite the apparent fit, we’ve yet to see a voice AI product succeed in gaming.

Backseat.GG was one early attempt. It offered voice-based coaching through AI personalities modeled after popular creators such as Tyler1 and Emiru. The service launched in 2024 and shut down later that year.

Backseat AI creator selection page featuring Jankos and Emiru alongside Tyler1 branding
Backed by Tyler1, the prospect of inviting your favorite creators to backseat you was a great meme.

Having tried it myself, the underlying concept of a personality-based coach is sound, but in practice it just kept getting in the way.

It clashed with in-game audio, was annoying when it spoke up out of turn, and most critically, it made it difficult to talk to the friends I was already in a Discord call with. Instead of complementing the social experience, it just became another source of noise.

That’s an important constraint worth calling out - voice may free up your hands, but it still consumes auditory attention.

A voice agent cannot simply speak whenever it has something to say. It must understand what is happening, recognize when the user is occupied, and decide whether its contribution is worth the interruption.

Discord’s inability to expand beyond gaming reinforces the same point. Voice is not actually the universally preferable UX. It dominates in a context where it solves a clear coordination problem and where users already accept the costs of synchronous conversation.

During my time working in VR, we spent billions of dollars trying to recreate that feeling of another person’s presence through the development of “telepresence” technology. Despite extraordinary progress, the experience remains far from sharing a physical room with someone.

That history makes me skeptical of the idea that a more natural artificial phone call is, by itself, the destination.

Perhaps we have framed consumer voice AI incorrectly from the beginning.

Voice may not replace every interface. It may instead become the best interface for a particular class of moments: driving, walking, cooking, exercising, gaming, or doing anything else that leaves your hands and eyes occupied.

Conclusion

When I started writing this post, I thought I was trying to answer a simple question:

If voice AI has gotten so good, why isn’t everyone using it?

By the end, I no longer think there is a single explanation.

Adoption may take time. The technology still has meaningful gaps. But none of these answers fully explain the problem.

Earlier, I compared voice AI to the early days of the iPhone. Looking back, what made smartphones indispensable was not the touchscreen, industrial design, or falling price point - it was the App Store.

Suddenly, the iPhone stopped being just an impressive piece of hardware and instead became a platform where people could do things they couldn’t do before. Sharing photos, exploring maps, watching videos - whatever it was, there was an app for that.

So while thinking through the Her problem, I keep thinking back to the question my own team at Frisson Labs has heard over and over again while testing our voice AI characters with users:

“What does it do?”

Across consumer voice AI, it seems, this question remains largely unanswered.

When it comes to bringing voice AI to the mainstream, I don’t think the inflection point will come from shaving another 50 milliseconds off latency, or improving expressiveness by 10% (though they are important as well).

Instead, it’ll happen when people discover something valuable that’s only possible because of voice.

  • Maybe it’s a companion who inhabits your favorite game alongside you.
  • Maybe it’s an assistant that follows you during your commute and daily tasks.
  • Maybe it is an AI that can join your meetings and participate without constantly demanding your attention.

Or maybe it is something none of us have even imagined.

Theodore sitting at his desk beside Samantha's desktop operating system in Her
What would an “app store” for Her even look like?

Why aren’t we using Her? Because we don’t know what she does for us yet.

That is the problem we are interested in at Frisson Labs.

We believe combining natural voice interaction with world understanding unlocks the ability for an AI to actually participate in an experience, not just talk about it.

Games are one of the first places this will become possible. These are worlds people already care about, environments whose state we can fully observe (without additional cameras and sensors), and spaces where voice is already a natural interface.

We want to build AI that can be present inside those worlds with you, and find out what kinds of human capabilities this would unlock.

If this article resonated with you - or if you think I’m completely off my knocker - I’d genuinely love to hear your thoughts.

Hit me up on X or LinkedIn and let’s chat.

Follow @coatol5