All posts
Product Thinking

The awkward pause is part of the sentence

Voice products should not treat natural speech as bad input.

Kshitija KatareProduct Experience · Behaviour Design · Communication Systems · 6 min read

I have a problem with “speak clearly”

I have a problem with the phrase “speak clearly.”

Not because clarity is bad. Obviously.

But because the moment a product tells me to “speak clearly”, I immediately become conscious of how I am speaking.

Suddenly I am performing for the microphone. I slow down. I try to form a perfect sentence before saying it. I avoid changing my mind halfway through.

And at that point, we have taken one of the most natural things humans do, talking, and turned it into another interface we have to learn.

That is exactly what I do not want from Lungoor.

Real thinking is surprisingly messy

Listen to yourself the next time you are explaining something to a colleague. You probably do some version of this:

“Let’s move the review to Thursday... actually Friday might be better because Rohan still has to send that thing... wait, no, Thursday is okay if we get it before lunch.”

That is not poor communication. That is thought happening live.

The important part is that by the end, your intent is reasonably clear. Traditional interfaces often expect intent to arrive neatly packaged. Humans rarely work like that.

We pause. We repair sentences. We remember context halfway through. We say “that thing” because both people in the room know exactly what that thing is.

In India, we also switch languages without holding a committee meeting about it first. A sentence can begin in English, pick up two Hindi words, borrow a work term that has no useful translation, then return to English.

Nobody thinks this is unusual. It is simply conversation.

Do not improve the user before you improve the interface.

The product should do more work, not the person

This is one of the behavioural principles I keep coming back to while working on Lungoor: do not improve the user before you improve the interface.

If a person has to learn to remove every “umm”, speak punctuation, avoid restarting a sentence, use one language at a time and remember a list of voice commands, then we have technically built voice software. We have not necessarily built a better way to work.

The better question is: how much of normal human behaviour can the product absorb gracefully?

Can you hesitate? Can you correct yourself? Can you say something badly first? Can you give the important instruction at the end? Can you mix languages? Can the system understand that “actually...” often means “please update what I just told you”?

Those details look tiny on a feature list. In use, they are the difference between talking naturally and operating a machine.

There is a line we should not cross

The interesting challenge is that cleanup can very easily become rewriting.

Suppose I say: “Tell them I don’t think we should launch this Friday because the payment issue is still risky.”

A useful system might remove filler, fix punctuation and make the sentence easier to read. It should not casually turn that into: “We have decided to postpone the launch due to unresolved payment concerns.”

That sounds cleaner. It also changed me. “I don’t think we should” became “we have decided.” Risk became certainty.

This is why I do not think the goal of voice AI is to make everybody sound impressive. Sometimes the slightly unsure sentence is the honest sentence.

Our job is to help people communicate their thought more clearly, not quietly replace the thought with one the model prefers.

Good behaviour design often looks like nothing happened

There is a funny thing about product design. When behavioural design is bad, everyone notices it. There are prompts everywhere. Nudges everywhere. Tutorials. Warnings. Confirmations.

When it is good, the person usually says: “Oh nice.” Then continues working.

That is the experience I want from voice. You should not have to admire the AI. You should be able to forget about it.

Say something. See that it understood you. Make a small correction if needed. Move on.

The awkward pause can stay awkward. The product can handle it.

What we are really designing

We talk a lot about voice interfaces as if voice is simply a replacement keyboard. I think that undersells the interesting part.

A keyboard receives fairly structured input. Voice receives thought while it is still becoming thought.

That means the design problem is not only speech recognition. It is patience. Context. Correction. Ambiguity. Tone. And knowing when not to improve something.

Those are very human problems. Which is probably why I find them so interesting.

The best voice product will not teach people how to speak to AI. It will get very good at understanding how people already speak. Including the pauses. Especially the pauses.

Try Lungoor Voice with a sentence you would never bother typing perfectly.

Try Lungoor Voice