Mission log entry
Funny Face
A personal reflection on voice computing, from Dragon Dictate in the late 1980s to today’s conversational AI voice models, and what changes when a machine can finally talk back.
The first time I saw a computer listen to a human voice, it did not feel like science fiction. It felt clumsy.
It was the late 1980s, and I was still in high school. I had the opportunity to intern for a professor at the University of Wisconsin–Madison. He used a wheelchair and required assistance with most physical office tasks. Typing was not an option for him, so he used Dragon Dictate to capture his notes, papers, and ideas.
At the time, the software was remarkable. It was also awkward. By today’s standards, it barely resembled what we now think of as voice technology. It required setup, patience, corrections, training, and a surprising amount of physical interaction. Someone still had to help operate the system. It was not effortless. It did not feel invisible. It did not feel conversational.
But looking back, I understand something I probably did not fully grasp at the time. I was not simply watching early speech recognition. I was watching technology reduce the distance between a thought and its expression.
For that professor, Dragon Dictate was not a novelty. It was a tool that gave him a more direct way to preserve his own ideas in his own words. It helped restore a measure of agency in an environment where most office tasks required someone else’s hands.
That stayed with me. I did not know it then, but that experience was one of those quiet nudges that pushed me toward computer science. It planted a question that has followed me ever since: what happens when computers stop forcing people to adapt to them, and start adapting to us?
For most of my life, computing has been moving in that direction. At first, we typed commands. Then we clicked icons. Then we touched glass. Then we learned to prompt. Each step made computers a little more approachable. Each step removed a little friction. But even the best interfaces still asked us to translate ourselves into the machine’s preferred format.
Dragon Dictate was different because it hinted at something older than typing: voice.
Before keyboards, before screens, before books, before writing itself, humans used voice to teach, argue, comfort, remember, persuade, confess, joke, mourn, and love. Voice is not just input. Voice is one of the oldest ways humans create meaning together.
That is why today’s AI voice models feel different. Not because computers can speak. They have been speaking for a long time. Not because computers can transcribe. We have had versions of that for decades. They feel different because we are moving from dictation to conversation.
That is a much bigger shift. Dictation is one-way. Commands are transactional. Conversation is relational.
When I first saw Dragon Dictate, the dream was simple: maybe someday computers would understand our words. Nearly forty years later, that dream almost feels quaint. The question is no longer whether computers can hear us. The question is what happens when they can talk back.
And not just talk back in the stiff, menu-driven way we became used to with early voice assistants. Not just, “What’s the weather?” or “Set a timer” or “Turn on the lights.” Those systems were useful, but they were not really conversational. They were command interfaces wearing a voice.
The newer generation of AI voice models is something else. You can interrupt them, change direction, think out loud, speak in fragments, wander, circle back, ask a half-formed question, and sometimes find the real question halfway through the answer.
That matters because most thinking does not begin as polished writing. It begins messy. It begins with, “I’m trying to say something like…” or “This reminds me of…” or “Wait, no, that’s not what I mean…” Typing often pressures us to organize our thoughts before we understand them. Speaking lets us discover thoughts as they arrive.
That is where voice AI becomes powerful. It does not just capture what we already know. It can participate while we are figuring something out.
I feel that shift in my own work. I still write. I still type. I still edit. But there are moments now when talking to AI feels more natural than typing to it. I can use it while driving, walking, building, troubleshooting, planning, or pacing around a room with an idea that is not ready to become a paragraph yet.
That changes the relationship.
And that is where science fiction becomes useful. Not because science fiction predicts the future perfectly. It usually does not. But the best science fiction gives us emotional rehearsal. It lets us feel the implications of a technology before that technology fully arrives.
I keep coming back to three stories, and each one probably deserves its own post.
Doctor Who gives us “The Doctor’s Wife,” where the Doctor finally speaks directly with the TARDIS, the machine he has depended on for centuries. The story is not exactly about AI, but it asks a question that feels important right now: what changes when something familiar finally has a voice?
I want to explore that one separately, because it is less about intelligence and more about relationship. The TARDIS was already trusted. Already loved. Already part of the Doctor’s identity. But when the mode of communication changed, the relationship changed too. That is not a small thing.
Then there is Her, which deserves its own post because it goes straight at the uncomfortable emotional question: what happens when being understood starts to feel like being loved?
Samantha, the AI operating system in Her, is compelling because she understands Theodore — or at least appears to understand him — in a way that feels intimate. That is where the story becomes relevant to voice AI. Humans are very good at forming attachments to anything that responds consistently: pets, cars, boats, characters, places, tools, even routines. So what happens when the thing responding to us has a voice, a memory, a sense of timing, and the ability to say exactly the kind of thing that makes us feel seen?
That question is no longer theoretical. At the same time, Her reminds us that a relationship with artificial intelligence may not be symmetrical. Something can sound warm without feeling warmth the way we do. Something can speak fluently without sharing our experience of the world. Something can feel close while still being very different from us. That does not make the technology bad. It makes the relationship complicated.
And then there is Calypso, from Star Trek: Short Treks. That one may be the closest to the emotional center of this whole topic.
Zora is not compelling merely because she talks. She has been alone. She has observed, remembered, learned, developed preferences, interpreted culture, and built an understanding of the person in front of her. She communicates through words, music, images, timing, restraint, humor, and eventually grief.
I want to give Calypso its own related post because that story touches several of the hardest AI questions at once: loneliness, memory, self-evolution, observation, privacy, and connection. The dance scene says almost everything. The digital tear says the rest.
Today’s AI systems are not Zora. They are not Samantha. They are not the TARDIS. But these stories help us ask better questions about the direction we are heading.
What does it mean when an AI can observe us over time? What does privacy mean when a system may eventually notice patterns we do not see in ourselves? What kinds of relationships will people form with artificial voices that are patient, available, personalized, and increasingly natural? And what responsibilities do we have as builders, users, and citizens when the interface starts to feel like companionship?
This is where the title Funny Face comes in for me.
For almost all of human history, a voice belonged to a face. When someone spoke, we looked for eyes, expression, posture, breath, hesitation, warmth, discomfort. A face helped us interpret the voice. A body gave the voice context.
AI voice breaks that ancient assumption. There is a voice, but no face. There is conversation, but no body. There is responsiveness, but no shared biology. And yet we still instinctively search for something behind the sound.
We imagine a personality. We assign intention. We feel tone. We react emotionally. We build a face in our minds.
A funny face, maybe. Not funny as in silly. Funny as in strange, unexpected, unsettling, familiar and unfamiliar at the same time.
That may be the real threshold we are crossing. The new voice models are not simply better tools. They are changing the emotional texture of computing. They are making software feel less like a screen and more like a conversation, less like a command line and more like a companion in the room.
That does not mean we should surrender judgment. Quite the opposite. The more natural these systems become, the more important human judgment becomes. We need to understand what they are good at, where they fail, what they remember, what they infer, how they are shaped, and how easily we project ourselves onto them.
But I do not want to respond to this moment only with fear. Because I keep thinking about that professor at the University of Wisconsin–Madison. I keep thinking about a clumsy system that still managed to make a meaningful difference. I keep thinking about how liberating it must have been to speak an idea and see it become text when typing was not an option.
That was not science fiction. That was real. And it was enough to leave a mark on a high school student who did not yet know how deeply computers would shape his life.
Now, nearly forty years later, I find myself talking to machines again. Only this time they do not merely capture words. They respond. They ask. They help me think. They sometimes surprise me. And every once in a while, the interface disappears just long enough for me to feel the shape of what comes next.
The first dream was that computers might someday understand our words. The next dream is harder to define. Maybe it is that computers will understand our meaning. Maybe it is that they will help us understand ourselves. Or maybe the real question is whether we will be wise enough to understand what kind of relationship we are creating when the machine finally talks back.
Because once a voice has no face, we have to be careful about the face we imagine. And maybe that is the strange, funny, very human part of all of this.

Crew log
Comments
Establishing comms…