BRIDGES

Think, Don’t Speak

A mixed-media collage of a vintage typewriter rendered in black and white, with two hands reaching in from the left and right sides to type on its keys. Five torn strips of colored paper in blue, red, teal, pink, and yellow rise from the typewriter's carriage like pages being produced. The background is white with scattered black ink speckles.

Translated from Chinese version with assistant of LLM. Translated text is reviewed by human.

Following the demise of TNT1, voice input has recently sparked another round of discussion. This time, it started with voice input tools led by Typeless. They claim to enable users to achieve perfect text output through speech. Typeless even includes a “manifesto” page on its website, confidently asserting that “Voice is our default,” thereby highlighting Typeless’s revolutionary nature.

A key reason traditional voice input has been unsatisfactory is that algorithms typically faithfully transcribe every word spoken. For example, when we speak, there may be pauses or instances where the mouth moves faster than the brain, and the algorithm outputs these without hesitation. In contrast, new voice input tools like Typeless and Wispr Flow pass the transcription results through a large language model for a round of processing. These apps are not only removing filler words, like “um,” “ah,” “yeah,” and “uh,” but also understanding commands for replacements like “oh no, what I meant was...”, and then polishing the output before finally delivering it to the input field.

The first time I saw this demo was on a friend’s phone. He showed me how, in the ChatGPT app, he used natural language to compose a very long prompt, essentially saying whatever came to mind. On the surface, the final transcription result was very clean and crisp: Typeless not only removed all filler words but also slightly polished the output. Over the past few months of group chats, many members have had nothing but praise for Typeless.

It all sounds great, but when I tried using Typeless myself, I found the actual experience to be completely at odds with my expectations. First, there’s the issue of “intrusion.” For me, the best scenario to use Typeless is when chatting with others, but in such situations, I might physically be with other people. In the presence of unrelated individuals, voice chatting makes me seem a bit rude. (I have heard Typeless has a whisper mode, but I haven’t actually tried it, so I can’t really comment on that.)

“Intrusion” is, after all, a defect that can be considered as a surmountable defect: if you absolutely must use it in public, you can at most move to a secluded spot or cover your mouth with your hand. What truly makes me keep my distance from Typeless is that, I feel voice input tools like Typeless are stripping away my thought process.

That said, if your input is heavily focused on “quick replies,” Typeless can indeed boost your efficiency significantly. A lawyer friend in one of group chats I joined is quite fond of Typeless, because he needs to respond to clients’ questions quickly, and since it involves professional legal opinions, each client’s question requires a very detailed reply. In such cases, Typeless can produce reasonable and professional responses in an extremely short time, making it a highly sensible choice.

But I’m different. Aside from chatting, a large portion of my energy at this stage goes into content creation. Since the last time I tried letting an LLM take the lead in writing (and had an unpleasant experience), I have insisted on handling the writing and output myself. I might use an LLM occasionally, but at most to help me check for awkward phrasing, typos, or to organize my thoughts.

Whether it’s Typeless or Wispr Flow, their marketing wants you to blindly believe that “because speaking is faster than typing, speaking is a better input method.” Wispr Flow even launched a typing competition: if your typing speed exceeds your speaking speed, you can win a Porsche. The typical process of such a typing competition involves testing keyboard input and voice input speed using the same fixed text.

The problem is, at least as I write the sentences you are reading, every piece of content requires careful “推敲” (deliberation) in its composition and creation. For me, a typical writing process goes something like this: first, I need to think about “how to phrase this sentence,” then type it out. I notice some word choices aren’t quite accurate, use the trackpad to select the inappropriate words, delete them, and consider which word would be more appropriate. Sometimes, halfway through writing, I go back to review what I've already written: “Could I summarise this with an idiom?” “Is this sentence grammatically incorrect?” “Is this expression not quite accurate?” These thoughts are endless while I am writing. After finishing a first draft, I still need to spend time reading through it from beginning to end, checking for any imperfections, and then making individual revisions.

So, the “revolutionary” voice input apps’ marketing claim that “speaking is faster than typing” simply doesn’t hold up (at least in the realm of text creation), because the normal pace of creating written content is significantly slower than typing a fixed text in a competition. In other words, creation is not inherently a “speed-first” activity. Typing, in essence, is a process of thinking: I input these words and phrases, read through them once, and see if they truly fit; rather than using a linear, uneditable expression process that forces me to interrupt my thoughts and correct errors with phrases like “no, it should be...”

A couple of days ago, I wanted to share a passage on Telegram and add my own notes. Due to a bug in the Telegram macOS client that causes garbled text when using a Chinese input method with input box buffer mode in the pop-up input box, I panicked and had to resort to Typeless input. Then I fell into 30 seconds of hell: speaking into Typeless, I not only repeatedly uttered filler words like “uh” and “um,” but also made frequent errors. In the end, Typeless was nearly impossible to polish into coherent phrases from my chaotic speech. I had to type my notes into another text editing app and then paste them in.

At least for me and my creative process, voice input is not an option: creation is a mental labor, and voice input is attempting to strip away my thought process, forcing me to outsource it. Naturally, I don’t see voice input as a better alternative to keyboard input. From this perspective, whether it’s Typeless’s marketing material grandiosely labeled as a “manifesto” or Wispr Flow’s typing competition, they are all castles in the air that ignore the true essence of creation.

A friend of mine shared with me an excerpt from Six Memos for the Next Millennium. I haven’t read the book, but I find this passage perfectly sums up part of my creative philosophy: “I think that my first impulse arises from a hypersensitivity or allergy It seems to me that language is always used in a random, approximate, careless manner, and this distresses me unbearably. Please don’t think that my reaction is the result of intolerance toward my neighbor: the worst discomfort of all comes from hearing myself speak. That’s why I try to talk as little as possible. If I prefer writing, it is because I can revise each sentence until I reach the point where—if not exactly satisfied with my words—I am able at least to eliminate those reasons for dissatisfaction that I can put a finger on. Literature — and I mean the literature that matches up to these requirements — is the Promised Land in which language becomes what it really ought to be.”

Feature image: Getty Images / Unsplash+

  1. Means “Touch and Talk”, a device made by the Chinese brand “Smartisan”. Encourage user to interact device without keyboard.

#Content Creating #Voice Input #Writing